{"doi":"10.1101/2024.07.17.603980","title":"PINDER: The protein interaction dataset and evaluation resource","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>\n                  Protein-protein interactions (PPIs) are fundamental to understanding biological processes and play a key role in therapeutic advancements. As deep-learning docking methods for PPIs gain traction, benchmarking protocols and datasets tailored for effective training and evaluation of their generalization capabilities and performance across real-world scenarios become imperative. Aiming to overcome limitations of existing approaches, we introduce PINDER, a comprehensive annotated dataset that uses structural clustering to derive non-redundant interface-based data splits and includes\n                  <jats:italic>holo</jats:italic>\n                  (bound),\n                  <jats:italic>apo</jats:italic>\n                  (unbound), and computationally predicted structures. PINDER consists of 2,319,564 dimeric PPI systems (and up to 25 million augmented PPIs) and 1,955 high-quality test PPIs with interface data leakage removed. Additionally, PINDER provides a test subset with 180 dimers for comparison to AlphaFold-Multimer without any interface leakage with respect to its training set. Unsurprisingly, the PINDER benchmark reveals that the performance of existing docking models is highly overestimated when evaluated on leaky test sets. Most importantly, by retraining DiffDock-PP on PINDER interface-clustered splits, we show that interface cluster-based sampling of the training split, along with the diverse and less leaky validation split, leads to strong generalization improvements.\n                </jats:p>","journal":null,"year":null,"id":600180,"datarank":0.4887144807032224,"base_score":3.258096538021482,"endowment":3.258096538021482,"self_citation_contribution":0.4887144807032224,"citation_network_contribution":0.0,"self_endowment_contribution":0.4887144807032224,"citer_contribution":0.0,"corpus_percentile":60.8,"corpus_rank":5125,"citation_count":25,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":803839,"name":"Mehmet Akdel","orcid":"0000-0002-6092-3494","position":1,"is_corresponding":false},{"id":242903,"name":"Alexander Goncearenco","orcid":"0000-0002-9738-7146","position":2,"is_corresponding":false},{"id":1538489,"name":"Guoqing Zhou","orcid":"0000-0002-4000-8467","position":3,"is_corresponding":false},{"id":139197,"name":"Graham T. Holt","orcid":"0000-0003-3354-9748","position":4,"is_corresponding":false},{"id":1538490,"name":"David Baugher","orcid":null,"position":5,"is_corresponding":false},{"id":566538,"name":"Dejun Lin","orcid":null,"position":6,"is_corresponding":false},{"id":557690,"name":"Yusuf Adeshina","orcid":"0000-0001-8389-4203","position":7,"is_corresponding":false},{"id":1538491,"name":"Thomas Castiglione","orcid":null,"position":8,"is_corresponding":false},{"id":443656,"name":"Xiaoyun Wang","orcid":"0000-0001-7780-1488","position":9,"is_corresponding":false},{"id":1366505,"name":"Céline Marquet","orcid":"0000-0002-8691-5791","position":10,"is_corresponding":false},{"id":1538492,"name":"Matt McPartlon","orcid":null,"position":11,"is_corresponding":false},{"id":1538493,"name":"Tomas Geffner","orcid":null,"position":12,"is_corresponding":false},{"id":1538494,"name":"Emanuele Rossi","orcid":"0009-0008-5882-4519","position":13,"is_corresponding":false},{"id":1159770,"name":"Gabriele Corso","orcid":"0000-0002-1963-8755","position":14,"is_corresponding":false},{"id":841525,"name":"H. Stärk","orcid":"0000-0002-4463-326X","position":15,"is_corresponding":false},{"id":1538495,"name":"Zachary Carpenter","orcid":null,"position":16,"is_corresponding":false},{"id":1538496,"name":"Emine Kucukbenli","orcid":null,"position":17,"is_corresponding":false},{"id":1538497,"name":"Michael Bronstein","orcid":null,"position":18,"is_corresponding":false},{"id":1538498,"name":"Luca Naef","orcid":"0000-0001-5955-0347","position":19,"is_corresponding":false},{"id":1538488,"name":"Daniel Kovtun","orcid":"0009-0005-5616-9707","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"PINDER: The protein interaction dataset and evaluation resource","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>\n                  Protein-protein interactions (PPIs) are fundamental to understanding biological processes and play a key role in therapeutic advancements. As deep-learning docking methods for PPIs gain traction, benchmarking protocols and datasets tailored for effective training and evaluation of their generalization capabilities and performance across real-world scenarios become imperative. Aiming to overcome limitations of existing approaches, we introduce PINDER, a comprehensive annotated dataset that uses structural clustering to derive non-redundant interface-based data splits and includes\n                  <jats:italic>holo</jats:italic>\n                  (bound),\n                  <jats:italic>apo</jats:italic>\n                  (unbound), and computationally predicted structures. PINDER consists of 2,319,564 dimeric PPI systems (and up to 25 million augmented PPIs) and 1,955 high-quality test PPIs with interface data leakage removed. Additionally, PINDER provides a test subset with 180 dimers for comparison to AlphaFold-Multimer without any interface leakage with respect to its training set. Unsurprisingly, the PINDER benchmark reveals that the performance of existing docking models is highly overestimated when evaluated on leaky test sets. Most importantly, by retraining DiffDock-PP on PINDER interface-clustered splits, we show that interface cluster-based sampling of the training split, along with the diverse and less leaky validation split, leads to strong generalization improvements.\n                </jats:p>","is_dataset_classified":null,"base_score":0.0,"endowment":0.0,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"21097893","pmcid":null,"openalex_id":null,"authors":[],"funders":[],"total_grants":0,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[],"oa_status":"green","license":"cc-by","oa_locations":[{"url":"https://doi.org/10.1101/2024.07.17.603980","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.1101/2024.07.17.603980","host_type":"publisher"}],"fields_of_study":[],"mesh_terms":[],"keywords":[],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-29T12:24:54.129504Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}