{"doi":"10.1101/791293","title":"Proteome-scale discovery of protein interactions with residue-level resolution using sequence coevolution","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>\n                  The majority of protein interactions in most organisms are unknown, and experimental methods for determining protein interactions can yield divergent results. Here we use an orthogonal, purely computational method based on sequence coevolution to discover protein interactions at large scale. In the model organism\n                  <jats:italic>Escherichia coli</jats:italic>\n                  , 53% of protein pairs in the proteome are eligible for our method given currently available sequenced genomes. When assaying the entire cell envelope proteome, which is understudied due to experimental challenges, we found 620 likely interactions and their predicted structures, increasing the space of known interactions by 529. Our results show that genomic sequencing data can be used to predict and resolve protein interactions to atomic resolution at large scale. Predictions and code are freely available at\n                  <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://marks.hms.harvard.edu/ecolicomplex\">https://marks.hms.harvard.edu/ecolicomplex</jats:ext-link>\n                  and\n                  <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://github.com/debbiemarkslab/EVcouplings\">https://github.com/debbiemarkslab/EVcouplings</jats:ext-link>\n                </jats:p>","journal":null,"year":null,"id":630462,"datarank":0.29188652235829704,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"self_citation_contribution":0.29188652235829704,"citation_network_contribution":0.0,"self_endowment_contribution":0.29188652235829704,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":6,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":629874,"name":"Hadeer Elhabashy","orcid":"0000-0002-4677-7064","position":1,"is_corresponding":false},{"id":261649,"name":"Kelly P. Brock","orcid":"0000-0002-5236-3773","position":2,"is_corresponding":false},{"id":629875,"name":"Rohan Maddamsetti","orcid":"0000-0003-3370-092X","position":3,"is_corresponding":false},{"id":105895,"name":"Oliver Kohlbacher","orcid":"0000-0003-1739-4598","position":4,"is_corresponding":false},{"id":56406,"name":"Debora S. Marks","orcid":"0000-0001-9388-2281","position":5,"is_corresponding":false},{"id":260809,"name":"Anna G. Green","orcid":"0000-0001-7548-3682","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Proteome-scale discovery of protein interactions with residue-level resolution using sequence coevolution","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>\n                  The majority of protein interactions in most organisms are unknown, and experimental methods for determining protein interactions can yield divergent results. Here we use an orthogonal, purely computational method based on sequence coevolution to discover protein interactions at large scale. In the model organism\n                  <jats:italic>Escherichia coli</jats:italic>\n                  , 53% of protein pairs in the proteome are eligible for our method given currently available sequenced genomes. When assaying the entire cell envelope proteome, which is understudied due to experimental challenges, we found 620 likely interactions and their predicted structures, increasing the space of known interactions by 529. Our results show that genomic sequencing data can be used to predict and resolve protein interactions to atomic resolution at large scale. Predictions and code are freely available at\n                  <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://marks.hms.harvard.edu/ecolicomplex\">https://marks.hms.harvard.edu/ecolicomplex</jats:ext-link>\n                  and\n                  <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://github.com/debbiemarkslab/EVcouplings\">https://github.com/debbiemarkslab/EVcouplings</jats:ext-link>\n                </jats:p>","is_dataset_classified":null,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19910364","pmcid":null,"openalex_id":"https://openalex.org/W2977939402","authors":[],"funders":[{"funder_name":"National Science Foundation","grant_id":"1144152","title":"Graduate Research Fellowship Program (GRFP)"},{"funder_name":"National Institutes of Health","grant_id":"1S10RR028832-01","title":"Orchestra: A high performance biomedical supercomputing collaborative"}],"total_grants":2,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2019,"count":1},{"year":2021,"count":5}],"oa_status":"green","license":"cc-by-nc-nd","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2019/10/22/791293.full.pdf","host_type":"repository"},{"url":"https://www.biorxiv.org/content/biorxiv/early/2019/10/22/791293.full.pdf","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.1101/791293","host_type":"publisher"},{"url":"https://doi.org/10.1101/791293","host_type":"repository"},{"url":"https://dx.doi.org/10.1101/791293","host_type":""},{"url":"http://dx.doi.org/10.1101/791293","host_type":""}],"fields_of_study":["Protein Structure and Dynamics","Genomics and Phylogenetic Studies","Microbial Metabolic Engineering and Bioproduction","0301 basic medicine","0303 health sciences","03 medical and health sciences"],"mesh_terms":[],"keywords":["Proteome","Computational biology","Coevolution","Genome","Biology","Protein sequencing","Sequence (biology)","Code (set theory)","Scale (ratio)","Genomics","Computer science","Genetics","Evolutionary biology","Peptide sequence","Gene","Physics"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-05T21:21:22.675973Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}