{"doi":"10.1002/jcc.25534","title":"Single‐sequence‐based prediction of protein secondary structures and solvent accessibility by deep whole‐sequence learning","abstract":"<jats:p>Predicting protein structure from sequence alone is challenging. Thus, the majority of methods for protein structure prediction rely on evolutionary information from multiple sequence alignments. In previous work we showed that Long Short‐Term Bidirectional Recurrent Neural Networks (LSTM‐BRNNs) improved over regular neural networks by better capturing intra‐sequence dependencies. Here we show a single‐sequence‐based prediction method employing LSTM‐BRNNs (SPIDER3‐Single), that consistently achieves Q3 accuracy of 72.5%, and correlation coefficient of 0.67 between predicted and actual solvent accessible surface area. Moreover, it yields reasonably accurate prediction of eight‐state secondary structure, main‐chain angles (backbone <jats:styled-content>\n<jats:italic>ϕ</jats:italic></jats:styled-content> and <jats:styled-content>\n<jats:italic>ψ</jats:italic></jats:styled-content> torsion angles and C<jats:styled-content>\n<jats:italic>α</jats:italic></jats:styled-content>‐atom‐based <jats:styled-content>\n<jats:italic>θ</jats:italic></jats:styled-content> and <jats:styled-content>\n<jats:italic>τ</jats:italic></jats:styled-content> angles), half‐sphere exposure, and contact number. The method is more accurate than the corresponding evolutionary‐based method for proteins with few sequence homologs, and computationally efficient for large‐scale screening of protein‐structural properties. It is available as an option in the SPIDER3 server, and a standalone version for download, at <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" xlink:href=\"http://sparks-lab.org\">http://sparks-lab.org</jats:ext-link>. © 2018 Wiley Periodicals, Inc.</jats:p>","journal":"Journal of Computational Chemistry","year":2018,"id":596192,"datarank":0.7301301675683375,"base_score":4.867534450455582,"endowment":4.867534450455582,"self_citation_contribution":0.7301301675683375,"citation_network_contribution":0.0,"self_endowment_contribution":0.7301301675683375,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":129,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":8,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1526793,"name":"Kuldip Paliwal","orcid":null,"position":1,"is_corresponding":false},{"id":1526794,"name":"James Lyons","orcid":null,"position":2,"is_corresponding":false},{"id":131733,"name":"Jaswinder Singh","orcid":null,"position":3,"is_corresponding":false},{"id":952176,"name":"Yuedong Yang","orcid":"0000-0002-6782-2813","position":4,"is_corresponding":false},{"id":292668,"name":"Yaoqi Zhou","orcid":"0000-0002-9958-5699","position":5,"is_corresponding":false},{"id":1526792,"name":"Rhys Heffernan","orcid":"0000-0002-9946-1995","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Single‐sequence‐based prediction of protein secondary structures and solvent accessibility by deep whole‐sequence learning","abstract":"<jats:p>Predicting protein structure from sequence alone is challenging. Thus, the majority of methods for protein structure prediction rely on evolutionary information from multiple sequence alignments. In previous work we showed that Long Short‐Term Bidirectional Recurrent Neural Networks (LSTM‐BRNNs) improved over regular neural networks by better capturing intra‐sequence dependencies. Here we show a single‐sequence‐based prediction method employing LSTM‐BRNNs (SPIDER3‐Single), that consistently achieves Q3 accuracy of 72.5%, and correlation coefficient of 0.67 between predicted and actual solvent accessible surface area. Moreover, it yields reasonably accurate prediction of eight‐state secondary structure, main‐chain angles (backbone <jats:styled-content>\n<jats:italic>ϕ</jats:italic></jats:styled-content> and <jats:styled-content>\n<jats:italic>ψ</jats:italic></jats:styled-content> torsion angles and C<jats:styled-content>\n<jats:italic>α</jats:italic></jats:styled-content>‐atom‐based <jats:styled-content>\n<jats:italic>θ</jats:italic></jats:styled-content> and <jats:styled-content>\n<jats:italic>τ</jats:italic></jats:styled-content> angles), half‐sphere exposure, and contact number. The method is more accurate than the corresponding evolutionary‐based method for proteins with few sequence homologs, and computationally efficient for large‐scale screening of protein‐structural properties. It is available as an option in the SPIDER3 server, and a standalone version for download, at <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" xlink:href=\"http://sparks-lab.org\">http://sparks-lab.org</jats:ext-link>. © 2018 Wiley Periodicals, Inc.</jats:p>","is_dataset_classified":null,"base_score":4.867534450455582,"endowment":4.867534450455582,"datacite_reuse_total":8,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"30368831","pmcid":null,"openalex_id":"https://openalex.org/W2898320067","authors":[],"funders":[{"funder_name":"National Health and Medical Research Council","grant_id":"1121629","title":"Developing species-specific, structure-targeting peptides as a novel class of antibiotics"},{"funder_name":"Australian Research Council","grant_id":"DP180102060","title":"Discovery Projects - Grant ID: DP180102060"}],"total_grants":2,"fwci":5.1069,"citation_percentile":0.96701066,"influential_citations":0,"citation_trend":[{"year":2018,"count":2},{"year":2019,"count":5},{"year":2020,"count":18},{"year":2021,"count":37},{"year":2022,"count":24},{"year":2023,"count":14},{"year":2024,"count":13},{"year":2025,"count":10},{"year":2026,"count":6}],"oa_status":"closed","license":"Wiley Online Library User Agreement","oa_locations":[{"url":"https://api.wiley.com/onlinelibrary/tdm/v1/articles/10.1002%2Fjcc.25534","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/pdf/10.1002/jcc.25534","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/full-xml/10.1002/jcc.25534","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/am-pdf/10.1002/jcc.25534","host_type":"publisher"},{"url":"https://doi.org/10.1002/jcc.25534","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/30368831","host_type":"repository"},{"url":"http://hdl.handle.net/10072/382158","host_type":"repository"},{"url":"https://rss.onlinelibrary.wiley.com/doi/am-pdf/10.1002/jcc.25534","host_type":""},{"url":"https://api.library.uq.edu.au/view/UQ:263ec19/thumbnail_UQ263ec19_OA_t.jpg","host_type":""},{"url":"https://api.library.uq.edu.au/view/UQ:263ec19","host_type":""},{"url":"https://api.library.uq.edu.au/view/UQ:263ec19/UQ263ec19_OA.pdf","host_type":""},{"url":"https://espace.library.uq.edu.au/view/UQ:263ec19/thumbnail_UQ263ec19_OA_t.jpg","host_type":""},{"url":"https://espace.library.uq.edu.au/view/UQ:263ec19","host_type":""},{"url":"https://espace.library.uq.edu.au/view/UQ:263ec19/UQ263ec19_OA.pdf","host_type":""},{"url":"https://dx.doi.org/10.1002/jcc.25534","host_type":""},{"url":"https://doi.org/https://doi.org/10.1002/jcc.25534","host_type":""}],"fields_of_study":["Protein Structure and Dynamics","Computational Drug Discovery Methods","Enzyme Structure and Function","0301 basic medicine","0303 health sciences","03 medical and health sciences"],"mesh_terms":["Deep Learning","Proteins","Solvents","Protein Structure, Secondary"],"keywords":["Sequence (biology)","Computer science","Protein structure prediction","Sequence learning","Artificial intelligence","Chemistry","Computational biology","Protein structure","Biochemistry","Biology","Secondary structure prediction","Contact Prediction","Backbone Angles","Solvent Accessibility Prediction","Proteins","612","1600 Chemistry","Protein Structure, Secondary","Deep Learning","Physical chemistry","Theoretical and computational chemistry","Theoretical and computational chemistry not elsewhere classified","Solvents","Nanotechnology","2605 Computational Mathematics"],"sdg_mappings":[],"linked_datasets":[{"doi":"10.6084/m9.figshare.14355247.v1","title":"Additional file 1 of Prediction and analysis of multiple protein lysine modified sites based on conditional wasserstein generative adversarial networks","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.14355247","title":"Additional file 1 of Prediction and analysis of multiple protein lysine modified sites based on conditional wasserstein generative adversarial networks","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.14355250.v1","title":"Additional file 2 of Prediction and analysis of multiple protein lysine modified sites based on conditional wasserstein generative adversarial networks","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.14355250","title":"Additional file 2 of Prediction and analysis of multiple protein lysine modified sites based on conditional wasserstein generative adversarial networks","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.14355253.v1","title":"Additional file 3 of Prediction and analysis of multiple protein lysine modified sites based on conditional wasserstein generative adversarial networks","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.14355253","title":"Additional file 3 of Prediction and analysis of multiple protein lysine modified sites based on conditional wasserstein generative adversarial networks","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.14355256.v1","title":"Additional file 4 of Prediction and analysis of multiple protein lysine modified sites based on conditional wasserstein generative adversarial networks","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.14355256","title":"Additional file 4 of Prediction and analysis of multiple protein lysine modified sites based on conditional wasserstein generative adversarial networks","publisher":"figshare","resource_type":"JournalArticle"}],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-28T05:16:52.712259Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}