{"doi":"10.1101/2024.01.18.576317","title":"CHOIR improves significance-based detection of cell types and states from single-cell data","abstract":"<jats:title>Abstract</jats:title>\n                <jats:p>\n                  Clustering is a critical step in the analysis of single-cell data, as it enables the discovery and characterization of putative cell types and states. However, most popular clustering tools do not subject clustering results to statistical inference testing, leading to risks of overclustering or underclustering data and often resulting in ineffective identification of cell types with widely differing prevalence. To address these challenges, we present CHOIR (\n                  <jats:underline>c</jats:underline>\n                  lustering\n                  <jats:underline>h</jats:underline>\n                  ierarchy\n                  <jats:underline>o</jats:underline>\n                  ptimization by iterative random forests), which applies a framework of random forest classifiers and permutation tests across a hierarchical clustering tree to statistically determine which clusters represent distinct populations. We demonstrate the enhanced performance of CHOIR through extensive benchmarking against 14 existing clustering methods across 100 simulated and 4 real single-cell RNA-seq, ATAC-seq, spatial transcriptomic, and multi-omic datasets. CHOIR can be applied to any single-cell data type and provides a flexible, scalable, and robust solution to the important challenge of identifying biologically relevant cell groupings within heterogeneous single-cell data.\n                </jats:p>","journal":null,"year":null,"id":609283,"datarank":0.41588830833596724,"base_score":2.772588722239781,"endowment":2.772588722239781,"self_citation_contribution":0.41588830833596724,"citation_network_contribution":0.0,"self_endowment_contribution":0.41588830833596724,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":15,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":250334,"name":"Lennart Mucke","orcid":"0000-0001-6256-9559","position":1,"is_corresponding":false},{"id":15489,"name":"M. Ryan Corces","orcid":"0000-0001-7465-7652","position":2,"is_corresponding":false},{"id":570946,"name":"Cathrine Petersen","orcid":"0000-0002-5821-9828","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"CHOIR improves significance-based detection of cell types and states from single-cell data","abstract":"Clustering is a critical step in the analysis of single-cell data, as it enables the discovery and characterization of putative cell types and states. However, most popular clustering tools do not subject clustering results to statistical inference testing, leading to risks of overclustering or underclustering data and often resulting in ineffective identification of cell types with widely differing prevalence. To address these challenges, we present CHOIR (clustering hierarchy optimization by iterative random forests), which applies a framework of random forest classifiers and permutation tests across a hierarchical clustering tree to statistically determine which clusters represent distinct populations. We demonstrate the enhanced performance of CHOIR through extensive benchmarking against 14 existing clustering methods across 100 simulated and 4 real single-cell RNA-seq, ATAC-seq, spatial transcriptomic, and multi-omic datasets. CHOIR can be applied to any single-cell data type and provides a flexible, scalable, and robust solution to the important challenge of identifying biologically relevant cell groupings within heterogeneous single-cell data.","is_dataset_classified":null,"base_score":2.772588722239781,"endowment":2.772588722239781,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"38328105","pmcid":null,"openalex_id":"https://openalex.org/W4391146032","authors":[],"funders":[{"funder_name":"National Institutes of Health","grant_id":"3P01AG073082-03S1","title":"Decoding the Multifactorial Etiology of Neural Network Dysfunction in Alzheimer's Disease"},{"funder_name":"National Institutes of Health","grant_id":"1R21AG085079-01","title":"Transcriptomic and Proteomic Analysis of Tau-dependent E/I Imbalance"},{"funder_name":"National Institutes of Health","grant_id":"1U01AG072573-01","title":"Multi-omic functional assessment of novel AD variants using high-throughput and single-cell technologies"},{"funder_name":"National Institutes of Health","grant_id":"5UM1HG012076-04","title":"Single-cell Mapping Center for Human Regulatory Elements and Gene Activity"}],"total_grants":4,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2023,"count":1},{"year":2024,"count":5},{"year":2025,"count":4},{"year":2026,"count":5}],"oa_status":"green","license":"cc-by-nc-nd","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2024/01/23/2024.01.18.576317.full.pdf","host_type":"repository"},{"url":"https://www.biorxiv.org/content/biorxiv/early/2024/01/23/2024.01.18.576317.full.pdf","host_type":"repository"},{"url":"https://doi.org/10.1101/2024.01.18.576317","host_type":"repository"},{"url":"https://pubmed.ncbi.nlm.nih.gov/38328105","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/10849522","host_type":"repository"},{"url":"https://escholarship.org/uc/item/87n2w8tm","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC10849522/pdf/nihpp-2024.01.18.576317v2.pdf","host_type":"repository"},{"url":"https://doi.org/10.1038/s41588-025-02148-8","host_type":""},{"url":"https://pubmed.ncbi.nlm.nih.gov/40195561","host_type":""},{"url":"http://dx.doi.org/10.1101/2024.01.18.576317","host_type":""}],"fields_of_study":["Single-cell and spatial transcriptomics","Gene expression and cancer classification","Gene Regulatory Network Analysis","0206 medical engineering","02 engineering and technology"],"mesh_terms":[],"keywords":["Cluster analysis","Computer science","Hierarchical clustering","Data mining","Random forest","Benchmarking","Identification (biology)","Machine learning","Biology","Sequence Analysis, RNA","Gene Expression Profiling","Humans","Computational Biology","Single-Cell Analysis","Transcriptome","Article","Algorithms","Software"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Life in Land"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-31T05:31:55.951699Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}