{"doi":"10.1002/prot.24307","title":"Capturing protein sequence–structure specificity using computational sequence design","abstract":"<jats:title>ABSTRACT</jats:title><jats:p>It is well known that protein fold recognition can be greatly improved if models for the underlying evolution history of the folds are taken into account. The improvement, however, exists only if such evolutionary information is available. To circumvent this limitation for protein families that only have a small number of representatives in current sequence databases, we follow an alternate approach in which the benefits of including evolutionary information can be recreated by using sequences generated by computational protein design algorithms. We explore this strategy on a large database of protein templates with 1747 members from different protein families. An automated method is used to design sequences for these templates. We use the backbones from the experimental structures as fixed templates, thread sequences on these backbones using a self‐consistent mean field approach, and score the fitness of the corresponding models using a semi‐empirical physical potential. Sequences designed for one template are translated into a hidden Markov model‐based profile. We describe the implementation of this method, the optimization of its parameters, and its performance. When the native sequences of the protein templates were tested against the library of these profiles, the class, fold, and family memberships of a large majority (&gt;90%) of these sequences were correctly recognized for an <jats:italic>E</jats:italic>‐value threshold of 1. In contrast, when homologous sequences were tested against the same library, a much smaller fraction (35%) of sequences were recognized; The structural classification of protein families corresponding to these sequences, however, are correctly recognized (with an accuracy of &gt;88%). Proteins 2013; © 2013 Wiley Periodicals, Inc.</jats:p>","journal":"Proteins: Structure, Function, and Bioinformatics","year":2013,"id":15272,"datarank":0.5142754853385527,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"self_citation_contribution":0.29188652235829704,"citation_network_contribution":0.22238896298025573,"self_endowment_contribution":0.29188652235829704,"citer_contribution":0.22238896298025573,"corpus_percentile":null,"corpus_rank":null,"citation_count":6,"citer_count":6,"citers_with_citation_signal":6,"citers_with_endowment":6,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":114980,"name":"Patrice Koehl","orcid":null,"position":1,"is_corresponding":false},{"id":117215,"name":"Paul Mach","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"23609941","pmcid":null,"openalex_id":"https://openalex.org/W1556505944","authors":[],"funders":[{"funder_name":"The National Institute of Health","grant_id":"","title":null}],"total_grants":1,"fwci":0.4487,"citation_percentile":0.64311877,"influential_citations":0,"citation_trend":[{"year":2015,"count":1},{"year":2016,"count":2},{"year":2017,"count":1},{"year":2019,"count":1},{"year":2022,"count":1}],"oa_status":"closed","license":"http://onlinelibrary.wiley.com/termsAndConditions#vor","oa_locations":[{"url":"https://api.wiley.com/onlinelibrary/tdm/v1/articles/10.1002%2Fprot.24307","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/pdf/10.1002/prot.24307","host_type":"publisher"},{"url":"https://doi.org/10.1002/prot.24307","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/23609941","host_type":"repository"}],"fields_of_study":["Protein Structure and Dynamics","Machine Learning in Bioinformatics","RNA and protein synthesis mechanisms","Medicine","Biology","Computer Science","Algorithms","Amino Acid Sequence","Computational Biology","Databases, Protein","Markov Chains","Protein Folding","Proteins","Sequence Alignment","Sequence Analysis, Protein","Thermodynamics"],"mesh_terms":["Algorithms","Amino Acid Sequence","Markov Chains","Proteins","Thermodynamics","Sequence Alignment","Protein Folding","Computational Biology","Sequence Analysis, Protein","Databases, Protein"],"keywords":["Template","Hidden Markov model","Protein structure prediction","Protein structure database","Computer science","Protein family","Protein design","Protein sequencing","Sequence (biology)","Protein superfamily","Sequence alignment","Multiple sequence alignment","Structural Classification of Proteins database","Protein structure","Threading (protein sequence)","Computational biology","Biology","Artificial intelligence","Peptide sequence","Sequence database","Genetics","Gene","Hidden Markov models","protein fold recognition","Computational Protein Sequence Design","Sequence Threading"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-06-01T17:05:25.350951Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}