{"doi":"10.3390/biom13030498","title":"On the Best Way to Cluster NCI-60 Molecules","abstract":"<jats:p>Machine learning-based models have been widely used in the early drug-design pipeline. To validate these models, cross-validation strategies have been employed, including those using clustering of molecules in terms of their chemical structures. However, the poor clustering of compounds will compromise such validation, especially on test molecules dissimilar to those in the training set. This study aims at finding the best way to cluster the molecules screened by the National Cancer Institute (NCI)-60 project by comparing hierarchical, Taylor–Butina, and uniform manifold approximation and projection (UMAP) clustering methods. The best-performing algorithm can then be used to generate clusters for model validation strategies. This study also aims at measuring the impact of removing outlier molecules prior to the clustering step. Clustering results are evaluated using three well-known clustering quality metrics. In addition, we compute an average similarity matrix to assess the quality of each cluster. The results show variation in clustering quality from method to method. The clusters obtained by the hierarchical and Taylor–Butina methods are more computationally expensive to use in cross-validation strategies, and both cluster the molecules poorly. In contrast, the UMAP method provides the best quality, and therefore we recommend it to analyze this highly valuable dataset.</jats:p>","journal":"Biomolecules","year":2023,"id":637032,"datarank":0.4887144807032224,"base_score":3.258096538021482,"endowment":3.258096538021482,"self_citation_contribution":0.4887144807032224,"citation_network_contribution":0.0,"self_endowment_contribution":0.4887144807032224,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":25,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":12262,"name":"Pedro J. Ballester","orcid":"0000-0002-4078-743X","position":1,"is_corresponding":false},{"id":1653793,"name":"Saiveth Hernández-Hernández","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"On the Best Way to Cluster NCI-60 Molecules","abstract":"<jats:p>Machine learning-based models have been widely used in the early drug-design pipeline. To validate these models, cross-validation strategies have been employed, including those using clustering of molecules in terms of their chemical structures. However, the poor clustering of compounds will compromise such validation, especially on test molecules dissimilar to those in the training set. This study aims at finding the best way to cluster the molecules screened by the National Cancer Institute (NCI)-60 project by comparing hierarchical, Taylor–Butina, and uniform manifold approximation and projection (UMAP) clustering methods. The best-performing algorithm can then be used to generate clusters for model validation strategies. This study also aims at measuring the impact of removing outlier molecules prior to the clustering step. Clustering results are evaluated using three well-known clustering quality metrics. In addition, we compute an average similarity matrix to assess the quality of each cluster. The results show variation in clustering quality from method to method. The clusters obtained by the hierarchical and Taylor–Butina methods are more computationally expensive to use in cross-validation strategies, and both cluster the molecules poorly. In contrast, the UMAP method provides the best quality, and therefore we recommend it to analyze this highly valuable dataset.</jats:p>","is_dataset_classified":null,"base_score":3.258096538021482,"endowment":3.258096538021482,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"36979433","pmcid":"PMC10046274","openalex_id":"https://openalex.org/W4323565291","authors":[],"funders":[{"funder_name":"National Council of Sciences and Technology of Mexico (CONACYT)","grant_id":"775584","title":null},{"funder_name":"Royal Society for a Royal Society Wolfson Fellowship","grant_id":"","title":null},{"funder_name":"Wolfson Foundation","grant_id":"","title":null}],"total_grants":3,"fwci":3.5581,"citation_percentile":0.94147069,"influential_citations":0,"citation_trend":[{"year":2023,"count":1},{"year":2024,"count":11},{"year":2025,"count":11},{"year":2026,"count":2}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://www.mdpi.com/2218-273X/13/3/498/pdf?version=1678764532","host_type":"journal"},{"url":"https://www.mdpi.com/2218-273X/13/3/498/pdf?version=1678764532","host_type":"publisher"},{"url":"https://www.mdpi.com/2218-273X/13/3/498/pdf","host_type":"publisher"},{"url":"https://doi.org/10.3390/biom13030498","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/36979433","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/10046274","host_type":"repository"},{"url":"https://doaj.org/article/7d187048d9894295a87124325bb3e03a","host_type":"repository"},{"url":"http://hdl.handle.net/10044/1/108527","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC10046274/pdf/biomolecules-13-00498.pdf","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC10046274","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC10046274?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Computational Drug Discovery Methods","Analytical Chemistry and Chromatography","Machine Learning in Materials Science"],"mesh_terms":["Machine Learning","Algorithms","United States","Drug Design","Cluster Analysis","National Cancer Institute (U.S.)"],"keywords":["Cluster analysis","Computer science","Data mining","Hierarchical clustering","Set (abstract data type)","Outlier","Cluster (spacecraft)","Correlation clustering","Single-linkage clustering","Artificial intelligence","Machine learning","Pattern recognition (psychology)","CURE data clustering algorithm","Clustering","Model validation","Small molecules","Nci-60 Panel"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-06T18:10:24.473752Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}