{"doi":"10.5772/intechopen.1013230","title":"Perspectives Chapter: Data-Centric Strategies for Machine Learning-Driven Therapeutic Peptide Design – Challenges and Perspectives","abstract":"<jats:p>Machine learning is increasingly applied to the discovery and rational design of therapeutic peptides, offering powerful capabilities for functional prediction, pharmacological profiling, and de novo sequence generation. However, the reliability and generalizability of such approaches critically depend on the quality, completeness, and interoperability of the datasets on which they are built. This chapter provides a comprehensive overview of the computational strategies and best practices for constructing peptide datasets suitable for machine learning applications. We review widely used peptide repositories and examine strategies for data extraction, preprocessing, and standardization, with emphasis on mitigating annotation errors, redundancy, and incomplete metadata. Approaches for redundancy reduction––including sequence homology clustering, structure-based grouping, and embedding-driven similarity analysis––are compared in terms of their impact on dataset diversity and downstream performance. We also discuss challenges posed by scarce and imbalanced datasets, highlighting data-efficient strategies such as transfer learning, positive-unlabeled learning, and contrastive learning. The chapter concludes by outlining future directions, including the development of a peptide-specific markup language, standardized reporting formats, and widespread adoption of FAIR-compliant practices. By combining rigorous preprocessing with interoperable data standards, the field can enable more robust, interpretable, and transferable ML models, accelerating the discovery of safe and effective therapeutic peptides.</jats:p>","journal":"Data Quality Matters - Best Practices for Integrity and Assurance","year":2026,"id":639823,"datarank":0.10397207708399181,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"self_citation_contribution":0.10397207708399181,"citation_network_contribution":0.0,"self_endowment_contribution":0.10397207708399181,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":1,"citer_count":1,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1662652,"name":"Sebastián Escobedo","orcid":null,"position":1,"is_corresponding":false},{"id":1662653,"name":"Norma Murillo-Acevedo","orcid":null,"position":2,"is_corresponding":false},{"id":1662655,"name":"Nicole Soto-García","orcid":null,"position":3,"is_corresponding":false},{"id":1662656,"name":"Diego Fernández-Villegas","orcid":null,"position":4,"is_corresponding":false},{"id":784100,"name":"Diego Sandoval","orcid":null,"position":5,"is_corresponding":false},{"id":1662657,"name":"Anamaría Daza","orcid":null,"position":6,"is_corresponding":false},{"id":1662651,"name":"David Medina-Ortiz","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Perspectives Chapter: Data-Centric Strategies for Machine Learning-Driven Therapeutic Peptide Design – Challenges and Perspectives","abstract":"<jats:p>Machine learning is increasingly applied to the discovery and rational design of therapeutic peptides, offering powerful capabilities for functional prediction, pharmacological profiling, and de novo sequence generation. However, the reliability and generalizability of such approaches critically depend on the quality, completeness, and interoperability of the datasets on which they are built. This chapter provides a comprehensive overview of the computational strategies and best practices for constructing peptide datasets suitable for machine learning applications. We review widely used peptide repositories and examine strategies for data extraction, preprocessing, and standardization, with emphasis on mitigating annotation errors, redundancy, and incomplete metadata. Approaches for redundancy reduction––including sequence homology clustering, structure-based grouping, and embedding-driven similarity analysis––are compared in terms of their impact on dataset diversity and downstream performance. We also discuss challenges posed by scarce and imbalanced datasets, highlighting data-efficient strategies such as transfer learning, positive-unlabeled learning, and contrastive learning. The chapter concludes by outlining future directions, including the development of a peptide-specific markup language, standardized reporting formats, and widespread adoption of FAIR-compliant practices. By combining rigorous preprocessing with interoperable data standards, the field can enable more robust, interpretable, and transferable ML models, accelerating the discovery of safe and effective therapeutic peptides.</jats:p>","is_dataset_classified":null,"base_score":0.0,"endowment":0.0,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19767382","pmcid":null,"openalex_id":null,"authors":[],"funders":[],"total_grants":0,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[],"oa_status":"hybrid","license":"cc-by","oa_locations":[{"url":"https://www.intechopen.com/citation-pdf-url/1230184","host_type":"publisher"},{"url":"https://intech-files.s3.amazonaws.com/a04Tc00000ClgZiIAJ/a09Tc000002lwx3IAA/Final-Perspectives%20Chapter%20DataCentric%20Strategies%20for%20%20%282026-01-27%2009%3A43%3A17%29.pdf","host_type":"publisher"}],"fields_of_study":[],"mesh_terms":[],"keywords":[],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-07T02:38:50.008572Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}