{"doi":"10.1002/pro.4271","title":"Group depositions to the Protein Data Bank need adequate presentation and different archiving protocol","abstract":"Accurate experimentally determined structure models of biological macromolecules are used by a large and diverse community of researchers. It is an established practice to base the assessment of the structure model quality on both, expectations of correct stereochemistry and, most importantly, on examination of the model's fit to the primary experimental evidence. In the case of X-ray crystallography, the primary evidence is provided by the electron density map. The worldwide Protein Data Bank (wwPDB1) is a global repository of macromolecular models and the accompanying experimental data that allow to examine agreement between the electron density and structural model using programs such as Coot,2 Chimera,3 Pymol,4 or Molstack.5 Throughout its 50-year history, the PDB has accumulated over 180,000 macromolecular structures, and gained the reputation of the gold standard in structural biology and of the most reliable data resource in biomedical research in general.6 Recently, the PDB has seen an influx of many depositions from large-scale crystallographic fragment screening projects using a complex computational procedure called Pan-Dataset Density Analysis (PanDDA7). Based on a sophisticated multi-data-set analysis of reference models and potential ligand-complex crystals,7 such group depositions serve the purpose of identifying very low-occupancy small molecule ligands in macromolecular complexes. In brief, the idea of PanDDA consists of partial background solvent inclusion and subtraction of a virtual “multi-crystal ground state,” to produce so-called “event map,” revealing the supposed ligands in each of a multitude of different ligand data sets. As a result, the PDB accumulates large numbers of “group depositions” of many putative ligand complexes of the same protein target that presently do not conform to the primary objective of the PDB as a repository of high-quality structure models and data. In particular, the maps that can be retrieved for such deposits have dubious agreement with the models (Figure 1). While suited for fragment screening and lead discovery,8 the group deposition models dumped en masse into the PDB do not conform to the quality standards9, 10 expected of PDB entries. In particular, PanDDA deposits confuse most biomedical researchers as their data structure is different than that of other PDB deposits and the quality of the structure models, despite occasional high nominal resolution, is often questionable. The presence of group deposits that do not conform to PDB standards of data retrieval and model quality, but nevertheless are presented on a par with conventional entries, degrades the PDB integrity. For standard entries, the PDB-provided map coefficients or density maps are calculated from the supplied experimental and model data, allowing validation of the model against experiment. This procedure is not possible for PanDDA entries, as the “event maps”7 have a completely different nature and purpose. Consequently, group deposition entries are difficult or even impossible to individually validate and assess. They can mislead PDB users when selected as models to underpin further studies and may also mislead systems like AlphaFold211 during selection of optimal templates for structure prediction. Presence of nonconforming entries is particularly problematic for automated data-mining projects, including applications of Artificial Intelligence, as the presence of such data adds unexpected levels of noise during the learning and testing stages. The gold standard of structural biology, that is, the agreement of a structure model with underlying electron density, fails for group deposition models that severely disagree with the user-accessible electron density maps (Figure 1). If group deposition entries violating accepted quality metrics9, 10 become primary references for their protein families due to their reported very high resolution, the effect could be disastrous (Figure 1b). The presence of such dep","journal":"Protein Science","year":2022,"id":268738,"datarank":1.0495010440257428,"base_score":2.995732273553991,"endowment":2.995732273553991,"self_citation_contribution":0.4493598410330987,"citation_network_contribution":0.6001412029926442,"self_endowment_contribution":0.4493598410330987,"citer_contribution":0.6001412029926442,"corpus_percentile":80.69931151852711,"corpus_rank":2496,"citation_count":19,"citer_count":16,"citers_with_citation_signal":14,"citers_with_endowment":14,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.5538,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2022-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":376430,"name":"Alexander Wlodawer","orcid":"0000-0002-5510-9703","position":1,"is_corresponding":false},{"id":318134,"name":"Zbigniew Dauter","orcid":"0000-0002-8806-9066","position":2,"is_corresponding":false},{"id":17727,"name":"W. Minor","orcid":"0000-0001-7075-7090","position":3,"is_corresponding":false},{"id":376433,"name":"Bernhard Rupp","orcid":"0000-0002-3300-6965","position":4,"is_corresponding":false},{"id":376434,"name":"Mariusz Jaskólski","orcid":"0000-0003-1587-6489","position":0,"is_corresponding":true}],"reference_count":13,"raw_metadata":null,"created_at":"2026-07-19T00:27:14.088663Z","pmid":"35000237","pmcid":"PMC8927872","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}