{"doi":"10.1002/pro.4239","title":"The importance of residue‐level filtering and the Top2018 best‐parts dataset of high‐quality protein residues","abstract":"We have curated a high-quality, \"best-parts\" reference dataset of about 3 million protein residues in about 15,000 PDB-format coordinate files, each containing only residues with good electron density support for a physically acceptable model conformation. The resulting prefiltered data typically contain the entire core of each chain, in quite long continuous fragments. Each reference file is a single protein chain, and the total set of files were selected for low redundancy, high resolution, good MolProbity score, and other chain-level criteria. Then each residue was critically tested for adequate local map quality to firmly support its conformation, which must also be free of serious clashes or covalent-geometry outliers. The resulting Top2018 prefiltered datasets have been released on the Zenodo online web service and are freely available for all uses under a Creative Commons license. Currently, one dataset is residue filtered on main chain plus Cβ atoms, and a second dataset is full-residue filtered; each is available at four different sequence-identity levels. Here, we illustrate both statistics and examples that show the beneficial consequences of residue-level filtering. That process is necessary because even the best of structures contain a few highly disordered local regions with poor density and low-confidence conformations that should not be included in reference data. Therefore, the open distribution of these very large, prefiltered reference datasets constitutes a notable advance for structural bioinformatics and the fields that depend upon it.","journal":"Protein Science","year":2021,"id":179898,"datarank":0.7949220707139988,"base_score":2.9444389791664403,"endowment":2.9444389791664403,"self_citation_contribution":0.44166584687496613,"citation_network_contribution":0.3532562238390327,"self_endowment_contribution":0.44166584687496613,"citer_contribution":0.3532562238390327,"corpus_percentile":74.71184342848302,"corpus_rank":3270,"citation_count":18,"citer_count":14,"citers_with_citation_signal":11,"citers_with_endowment":11,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9449,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":55086,"name":"David Richardson","orcid":"0000-0001-5069-343X","position":1,"is_corresponding":false},{"id":55082,"name":"Jane S. Richardson","orcid":"0000-0002-3311-2944","position":2,"is_corresponding":false},{"id":551613,"name":"Christopher J. Williams","orcid":"0000-0002-5808-8768","position":0,"is_corresponding":true}],"reference_count":33,"raw_metadata":null,"created_at":"2026-07-18T23:47:53.089578Z","pmid":"34779043","pmcid":"PMC8740842","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}