{"doi":"10.3390/ijms140714892","title":"Structure Prediction of Partial-Length Protein Sequences","abstract":"<jats:p>Protein structure information is essential to understand protein function. Computational methods to accurately predict protein structure from the sequence have primarily been evaluated on protein sequences representing full-length native proteins. Here, we demonstrate that top-performing structure prediction methods can accurately predict the partial structures of proteins encoded by sequences that contain approximately 50% or more of the full-length protein sequence. We hypothesize that structure prediction may be useful for predicting functions of proteins whose corresponding genes are mapped expressed sequence tags (ESTs) that encode partial-length amino acid sequences. Additionally, we identify a confidence score representing the quality of a predicted structure as a useful means of predicting the likelihood that an arbitrary polypeptide sequence represents a portion of a foldable protein sequence (“foldability”). This work has ramifications for the prediction of protein structure with limited or noisy sequence information, as well as genome annotation.</jats:p>","journal":"International Journal of Molecular Sciences","year":2013,"id":31571,"datarank":0.4987501957641495,"base_score":1.791759469228055,"endowment":1.791759469228055,"self_citation_contribution":0.26876392038420827,"citation_network_contribution":0.2299862753799412,"self_endowment_contribution":0.26876392038420827,"citer_contribution":0.2299862753799412,"corpus_percentile":null,"corpus_rank":null,"citation_count":5,"citer_count":4,"citers_with_citation_signal":4,"citers_with_endowment":4,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":168910,"name":"Ling-Hong Hung","orcid":null,"position":1,"is_corresponding":false},{"id":28980,"name":"Ram Samudrala","orcid":"0000-0001-9069-8497","position":2,"is_corresponding":false},{"id":168909,"name":"Adrian Laurenzi","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"base_score":1.791759469228055,"endowment":1.791759469228055,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"23867606","pmcid":"PMC3742278","openalex_id":"https://openalex.org/W2034313881","authors":[],"funders":[{"funder_name":"NLM NIH HHS","grant_id":"DP1 LM011509","title":null},{"funder_name":"NIH HHS","grant_id":"DP1 OD006779","title":null},{"funder_name":"National Institutes of Health","grant_id":"5DP1OD006779-02","title":"NOVEL PARADIGMS FOR DRUG DISCOVERY: COMPUTATIONAL MULTITARGET SCREENING"},{"funder_name":"National Science Foundation","grant_id":"0448502","title":"CAREER: Accurate and Automated Protein Structure Prediction"}],"total_grants":4,"fwci":0.2991,"citation_percentile":0.61752154,"influential_citations":1,"citation_trend":[{"year":2015,"count":1},{"year":2016,"count":1},{"year":2017,"count":1},{"year":2018,"count":1},{"year":2020,"count":1}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://www.mdpi.com/1422-0067/14/7/14892/pdf?version=1403146902","host_type":"journal"},{"url":"https://www.mdpi.com/1422-0067/14/7/14892/pdf?version=1403146902","host_type":"GOLD"},{"url":"https://www.mdpi.com/1422-0067/14/7/14892/pdf?version=1403146902","host_type":"publisher"},{"url":"https://www.mdpi.com/1422-0067/14/7/14892/pdf","host_type":"publisher"},{"url":"https://doi.org/10.3390/ijms140714892","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/23867606","host_type":"repository"},{"url":"https://doaj.org/article/c42c382113df4deeb5d7ab4b8ece7018","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/3742278","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC3742278","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC3742278?pdf=render","host_type":"Europe_PMC"},{"url":"http://dx.doi.org/10.3390/ijms140714892","host_type":""},{"url":"https://dx.doi.org/10.3390/ijms140714892","host_type":""}],"fields_of_study":["Protein Structure and Dynamics","Enzyme Structure and Function","RNA and protein synthesis mechanisms","Biology","Medicine","Computer Science","0301 basic medicine","0303 health sciences","03 medical and health sciences","Databases, Protein","Expressed Sequence Tags","Protein Folding","Protein Structure, Tertiary","Proteins","Software"],"mesh_terms":["Proteins","Software","Protein Structure, Tertiary","Protein Folding","Expressed Sequence Tags","Databases, Protein"],"keywords":["Sequence (biology)","Protein structure prediction","Protein function prediction","Protein sequencing","Computational biology","ENCODE","Protein structure","Gene prediction","Sequence alignment","Sequence analysis","Genome","Biology","Peptide sequence","Genetics","Gene","Protein function","Expressed Sequence Tags","Protein Folding","expressed sequence tag","Proteins","Article","Protein Structure, Tertiary","EST","protein design","Databases, Protein","Software"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-06-09T07:40:08.075430Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}