{"doi":"10.1016/j.jbc.2021.100747","title":"Structural genomics and the Protein Data Bank","abstract":"The field of Structural Genomics arose over the last 3 decades to address a large and rapidly growing divergence between microbial genomic, functional, and structural data. Several international programs took advantage of the vast genomic sequence information and evaluated the feasibility of structure determination for expanded and newly discovered protein families. As a consequence, structural genomics has developed structure-determination pipelines and applied them to a wide range of novel, uncharacterized proteins, often from “microbial dark matter,” and later to proteins from human pathogens. Advances were especially needed in protein production and rapid de novo structure solution. The experimental three-dimensional models were promptly made public, facilitating structure determination of other members of the family and helping to understand their molecular and biochemical functions. Improvements in experimental methods and databases resulted in fast progress in molecular and structural biology. The Protein Data Bank structure repository played a central role in the coordination of structural genomics efforts and the structural biology community as a whole. It facilitated development of standards and validation tools essential for maintaining high quality of deposited structural data. The field of Structural Genomics arose over the last 3 decades to address a large and rapidly growing divergence between microbial genomic, functional, and structural data. Several international programs took advantage of the vast genomic sequence information and evaluated the feasibility of structure determination for expanded and newly discovered protein families. As a consequence, structural genomics has developed structure-determination pipelines and applied them to a wide range of novel, uncharacterized proteins, often from “microbial dark matter,” and later to proteins from human pathogens. Advances were especially needed in protein production and rapid de novo structure solution. The experimental three-dimensional models were promptly made public, facilitating structure determination of other members of the family and helping to understand their molecular and biochemical functions. Improvements in experimental methods and databases resulted in fast progress in molecular and structural biology. The Protein Data Bank structure repository played a central role in the coordination of structural genomics efforts and the structural biology community as a whole. It facilitated development of standards and validation tools essential for maintaining high quality of deposited structural data. The concept of Structural Genomics (SG) was born as a result of exponential progress in genome sequencing. The fast growth of DNA sequence information in the 1990s led to the generation of huge amounts of genomic data, which was accompanied by significant knowledge gaps in our understanding of biological roles and biochemical functions encoded in the genomes. Of importance, the sequence information bore little insights about the proteins (often called hypothetical) these newly discovered genes programmed, hampering progress toward functional interpretation. Massive accumulation of genomic and metagenomic sequences posed many questions that could not simply be neglected or ignored. To address these new challenges, the National Institutes of Health, Department of Energy, RIKEN, Gates Foundation, Wellcome Trust, and other numerous government and private agencies around the world funded structural genomics programs as early as 1997 to 2000. Table 1 summarizes the contribution of larger SG programs to determination of protein structures.Table 1Top 20 structural genomics programsCenterNumber of PDB depositsOrigin and fundingTechniques usedRIKEN Structural Genomics/Proteomics Initiative2746Japan, government, National Project on Protein Structural and Functional AnalysesNMR, X-rayMidwest Center for Structural Genomics1955USA, PSI/NIH/NIGMSX-ray, NMRStructural Geno","journal":"Journal of Biological Chemistry","year":2021,"id":178373,"datarank":1.1107498055690046,"base_score":2.995732273553991,"endowment":2.995732273553991,"self_citation_contribution":0.4493598410330987,"citation_network_contribution":0.6613899645359058,"self_endowment_contribution":0.4493598410330987,"citer_contribution":0.6613899645359058,"corpus_percentile":81.66627987932235,"corpus_rank":2371,"citation_count":19,"citer_count":19,"citers_with_citation_signal":14,"citers_with_endowment":14,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.5107,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":91662,"name":"A. Joachimiak","orcid":"0000-0003-2535-6209","position":1,"is_corresponding":false},{"id":231325,"name":"K. Michalska","orcid":"0000-0001-7140-3649","position":0,"is_corresponding":true}],"reference_count":71,"raw_metadata":null,"created_at":"2026-07-18T23:47:40.592580Z","pmid":"33957120","pmcid":"PMC8166929","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}