{"doi":"10.1101/gr.261115.120","title":"Alignment of single-cell RNA-seq samples without overcorrection using kernel density matching","abstract":"Single-cell RNA sequencing (scRNA-seq) technology is poised to replace bulk cell RNA sequencing for many biological and medical applications as it allows users to measure gene expression levels in a cell type-specific manner. However, data produced by scRNA-seq often exhibit batch effects that can be specific to a cell type, to a sample, or to an experiment, which prevent integration or comparisons across multiple experiments. Here, we present Dmatch, a method that leverages an external expression atlas of human primary cells and kernel density matching to align multiple scRNA-seq experiments for downstream biological analysis. Dmatch facilitates alignment of scRNA-seq data sets with cell types that may overlap only partially and thus allows integration of multiple distinct scRNA-seq experiments to extract biological insights. In simulation, Dmatch compares favorably to other alignment methods, both in terms of reducing sample-specific clustering and in terms of avoiding overcorrection. When applied to scRNA-seq data collected from clinical samples in a healthy individual and five autoimmune disease patients, Dmatch enabled cell type-specific differential gene expression comparisons across biopsy sites and disease conditions and uncovered a shared population of pro-inflammatory monocytes across biopsy sites in RA patients. We further show that Dmatch increases the number of eQTLs mapped from population scRNA-seq data. Dmatch is fast, scalable, and improves the utility of scRNA-seq for several important applications. Dmatch is freely available online.","journal":"Genome Research","year":2021,"id":213679,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":16,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9493,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":566192,"name":"Qi Zhang","orcid":"0000-0001-7959-817X","position":1,"is_corresponding":false},{"id":295012,"name":"Zepeng Mu","orcid":"0000-0002-7717-3247","position":2,"is_corresponding":false},{"id":555649,"name":"Lili Wang","orcid":"0000-0001-6357-8051","position":3,"is_corresponding":false},{"id":566747,"name":"Zhaohui Zheng","orcid":null,"position":4,"is_corresponding":false},{"id":566748,"name":"Jinlin Miao","orcid":null,"position":5,"is_corresponding":false},{"id":566193,"name":"Ping Zhu","orcid":"0000-0002-4888-2685","position":6,"is_corresponding":false},{"id":295014,"name":"Yang Li","orcid":"0000-0002-0736-251X","position":7,"is_corresponding":false},{"id":256389,"name":"Mengjie Chen","orcid":"0000-0003-1579-087X","position":0,"is_corresponding":true}],"reference_count":38,"raw_metadata":null,"created_at":"2026-07-18T23:52:36.886828Z","pmid":"33741686","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}