{"doi":"10.1110/ps.073037907","title":"The challenge of protein structure determination—lessons from structural genomics","abstract":"<jats:title>Abstract</jats:title><jats:p>The process of experimental determination of protein structure is marred with a high ratio of failures at many stages. With availability of large quantities of data from high‐throughput structure determination in structural genomics centers, we can now learn to recognize protein features correlated with failures; thus, we can recognize proteins more likely to succeed and eventually learn how to modify those that are less likely to succeed. Here, we identify several protein features that correlate strongly with successful protein production and crystallization and combine them into a single score that assesses “crystallization feasibility.” The formula derived here was tested with a jackknife procedure and validated on independent benchmark sets. The “crystallization feasibility” score described here is being applied to target selection in the Joint Center for Structural Genomics, and is now contributing to increasing the success rate, lowering the costs, and shortening the time for protein structure determination. Analyses of PDB depositions suggest that very similar features also play a role in non‐high‐throughput structure determination, suggesting that this crystallization feasibility score would also be of significant interest to structural biology, as well as to molecular and biochemistry laboratories.</jats:p>","journal":"Protein Science","year":2007,"id":607899,"datarank":0.7649799641736299,"base_score":5.099866427824199,"endowment":5.099866427824199,"self_citation_contribution":0.7649799641736299,"citation_network_contribution":0.0,"self_endowment_contribution":0.7649799641736299,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":163,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":239254,"name":"Lukasz Jaroszewski","orcid":null,"position":1,"is_corresponding":false},{"id":1561041,"name":"Ana P.C. Rodrigues","orcid":null,"position":2,"is_corresponding":false},{"id":116644,"name":"Leszek Rychlewski","orcid":null,"position":3,"is_corresponding":false},{"id":105430,"name":"Ian A. Wilson","orcid":"0000-0002-6469-2419","position":4,"is_corresponding":false},{"id":264648,"name":"Scott A. Lesley","orcid":"0000-0002-7905-8840","position":5,"is_corresponding":false},{"id":12333,"name":"Adam Godzik","orcid":"0000-0002-2425-852X","position":6,"is_corresponding":false},{"id":1561040,"name":"Lukasz Slabinski","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"The challenge of protein structure determination—lessons from structural genomics","abstract":"<jats:title>Abstract</jats:title><jats:p>The process of experimental determination of protein structure is marred with a high ratio of failures at many stages. With availability of large quantities of data from high‐throughput structure determination in structural genomics centers, we can now learn to recognize protein features correlated with failures; thus, we can recognize proteins more likely to succeed and eventually learn how to modify those that are less likely to succeed. Here, we identify several protein features that correlate strongly with successful protein production and crystallization and combine them into a single score that assesses “crystallization feasibility.” The formula derived here was tested with a jackknife procedure and validated on independent benchmark sets. The “crystallization feasibility” score described here is being applied to target selection in the Joint Center for Structural Genomics, and is now contributing to increasing the success rate, lowering the costs, and shortening the time for protein structure determination. Analyses of PDB depositions suggest that very similar features also play a role in non‐high‐throughput structure determination, suggesting that this crystallization feasibility score would also be of significant interest to structural biology, as well as to molecular and biochemistry laboratories.</jats:p>","is_dataset_classified":null,"base_score":5.099866427824199,"endowment":5.099866427824199,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"17962404","pmcid":"PMC2211687","openalex_id":"https://openalex.org/W2006431107","authors":[],"funders":[{"funder_name":"NIGMS NIH HHS","grant_id":"P20 GM076221","title":null},{"funder_name":"NIGMS NIH HHS","grant_id":"U54 GM074898","title":null}],"total_grants":2,"fwci":5.5617,"citation_percentile":0.97204199,"influential_citations":0,"citation_trend":[{"year":2012,"count":14},{"year":2013,"count":7},{"year":2014,"count":9},{"year":2015,"count":7},{"year":2016,"count":7},{"year":2017,"count":8},{"year":2018,"count":7},{"year":2019,"count":5},{"year":2020,"count":11},{"year":2021,"count":9},{"year":2022,"count":8},{"year":2023,"count":18},{"year":2024,"count":6},{"year":2025,"count":4},{"year":2026,"count":3}],"oa_status":"bronze","license":"http://onlinelibrary.wiley.com/termsAndConditions#vor","oa_locations":[{"url":"https://onlinelibrary.wiley.com/doi/pdfdirect/10.1110/ps.073037907","host_type":"journal"},{"url":"https://onlinelibrary.wiley.com/doi/pdfdirect/10.1110/ps.073037907","host_type":"publisher"},{"url":"https://api.wiley.com/onlinelibrary/tdm/v1/articles/10.1110%2Fps.073037907","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/pdf/10.1110/ps.073037907","host_type":"publisher"},{"url":"https://doi.org/10.1110/ps.073037907","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/17962404","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/2211687","host_type":"repository"}],"fields_of_study":["Enzyme Structure and Function","Protein Structure and Dynamics","Advanced Proteomics Techniques and Applications","Computational Biology","Crystallization","Crystallography, X-Ray","Databases, Protein","Genomics","Isoelectric Focusing","Magnetic Resonance Spectroscopy","Probability","Protein Conformation","Protein Structure, Secondary","Proteins","Proteomics","Sequence Analysis, Protein"],"mesh_terms":["Crystallization","Isoelectric Focusing","Magnetic Resonance Spectroscopy","Probability","Protein Conformation","Proteins","Protein Structure, Secondary","Crystallography, X-Ray","Computational Biology","Sequence Analysis, Protein","Genomics","Databases, Protein","Proteomics"],"keywords":["Structural genomics","Genomics","Structural biology","Protein crystallization","Protein Data Bank (RCSB PDB)","Crystallization","Benchmark (surveying)","Computational biology","Protein structure","Protein Data Bank","Computer science","Selection (genetic algorithm)","Throughput","Jackknife resampling","Biology","Chemistry","Genome","Mathematics","Machine learning","Genetics","Biochemistry","Gene","Statistics"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-30T07:11:45.231110Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}