{"doi":"10.1093/bioinformatics/btz292","title":"scMatch: a single-cell gene expression profile annotation tool using reference datasets","abstract":"<jats:title>Abstract</jats:title>\n               <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>Single-cell RNA sequencing (scRNA-seq) measures gene expression at the resolution of individual cells. Massively multiplexed single-cell profiling has enabled large-scale transcriptional analyses of thousands of cells in complex tissues. In most cases, the true identity of individual cells is unknown and needs to be inferred from the transcriptomic data. Existing methods typically cluster (group) cells based on similarities of their gene expression profiles and assign the same identity to all cells within each cluster using the averaged expression levels. However, scRNA-seq experiments typically produce low-coverage sequencing data for each cell, which hinders the clustering process.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>We introduce scMatch, which directly annotates single cells by identifying their closest match in large reference datasets. We used this strategy to annotate various single-cell datasets and evaluated the impacts of sequencing depth, similarity metric and reference datasets. We found that scMatch can rapidly and robustly annotate single cells with comparable accuracy to another recent cell annotation tool (SingleR), but that it is quicker and can handle larger reference datasets. We demonstrate how scMatch can handle large customized reference gene expression profiles that combine data from multiple sources, thus empowering researchers to identify cell populations in any complex tissue with the desired precision.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation</jats:title>\n                  <jats:p>scMatch (Python code) and the FANTOM5 reference dataset are freely available to the research community here https://github.com/forrest-lab/scMatch.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information</jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.</jats:p>\n               </jats:sec>","journal":"Bioinformatics","year":2019,"id":589420,"datarank":4.4792866241804985,"base_score":5.075173815233827,"endowment":5.075173815233827,"self_citation_contribution":0.7612760722850741,"citation_network_contribution":3.7180105518954245,"self_endowment_contribution":0.7612760722850741,"citer_contribution":3.7180105518954245,"corpus_percentile":null,"corpus_rank":null,"citation_count":159,"citer_count":137,"citers_with_citation_signal":104,"citers_with_endowment":104,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1508045,"name":"Elena Denisenko","orcid":null,"position":1,"is_corresponding":false},{"id":77580,"name":"Alistair R. R. Forrest","orcid":"0000-0003-4543-1675","position":2,"is_corresponding":false},{"id":1203810,"name":"Rui Hou","orcid":"0000-0001-6571-1514","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"scMatch: a single-cell gene expression profile annotation tool using reference datasets","abstract":"<jats:title>Abstract</jats:title>\n               <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>Single-cell RNA sequencing (scRNA-seq) measures gene expression at the resolution of individual cells. Massively multiplexed single-cell profiling has enabled large-scale transcriptional analyses of thousands of cells in complex tissues. In most cases, the true identity of individual cells is unknown and needs to be inferred from the transcriptomic data. Existing methods typically cluster (group) cells based on similarities of their gene expression profiles and assign the same identity to all cells within each cluster using the averaged expression levels. However, scRNA-seq experiments typically produce low-coverage sequencing data for each cell, which hinders the clustering process.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>We introduce scMatch, which directly annotates single cells by identifying their closest match in large reference datasets. We used this strategy to annotate various single-cell datasets and evaluated the impacts of sequencing depth, similarity metric and reference datasets. We found that scMatch can rapidly and robustly annotate single cells with comparable accuracy to another recent cell annotation tool (SingleR), but that it is quicker and can handle larger reference datasets. We demonstrate how scMatch can handle large customized reference gene expression profiles that combine data from multiple sources, thus empowering researchers to identify cell populations in any complex tissue with the desired precision.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation</jats:title>\n                  <jats:p>scMatch (Python code) and the FANTOM5 reference dataset are freely available to the research community here https://github.com/forrest-lab/scMatch.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information</jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.</jats:p>\n               </jats:sec>","is_dataset_classified":null,"base_score":5.075173815233827,"endowment":5.075173815233827,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"31028376","pmcid":"PMC6853649","openalex_id":"https://openalex.org/W2942380588","authors":[],"funders":[{"funder_name":"Australian National Health and Medical Research Council Fellowship","grant_id":"APP1154524","title":null},{"funder_name":"Cancer Council of Western Australia","grant_id":"","title":null},{"funder_name":"Cancer Research Trust ‘Enabling","grant_id":"","title":null},{"funder_name":"Australian Government Research Training Programme","grant_id":"","title":null},{"funder_name":"Cancer Research Trust","grant_id":"","title":null},{"funder_name":"Australian Government and the Government of Western Australia","grant_id":"","title":null}],"total_grants":6,"fwci":9.2528,"citation_percentile":0.98786834,"influential_citations":0,"citation_trend":[{"year":2019,"count":9},{"year":2020,"count":19},{"year":2021,"count":50},{"year":2022,"count":25},{"year":2023,"count":15},{"year":2024,"count":15},{"year":2025,"count":19},{"year":2026,"count":7}],"oa_status":"hybrid","license":"cc-by","oa_locations":[{"url":"https://academic.oup.com/bioinformatics/article-pdf/35/22/4688/30706721/btz292.pdf","host_type":"journal"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/35/22/4688/30706721/btz292.pdf","host_type":"publisher"},{"url":"http://academic.oup.com/bioinformatics/advance-article-pdf/doi/10.1093/bioinformatics/btz292/28665318/btz292.pdf","host_type":"publisher"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/35/22/4688/48978090/bioinformatics_35_22_4688.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/bioinformatics/btz292","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/31028376","host_type":"repository"},{"url":"https://admin.research-repository.uwa.edu.au/en/publications/8e87f9a0-2048-4a08-873e-07c7c4b33b0e","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/6853649","host_type":"repository"},{"url":"http://www.scopus.com/inward/record.url?scp=85073804186&partnerID=8YFLogxK","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC6853649","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC6853649?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Single-cell and spatial transcriptomics","Microfluidic and Bio-sensing Technologies","Cell Image Analysis Techniques","Algorithms","Gene Expression Profiling","Single-Cell Analysis","Software","Transcriptome"],"mesh_terms":["Algorithms","Software","Gene Expression Profiling","Single-Cell Analysis","Transcriptome"],"keywords":["Computer science","Annotation","Cluster analysis","Python (programming language)","Source code","Data mining","Computational biology","Gene expression profiling","RNA-Seq","Transcriptome","Gene","Gene expression","Biology","Artificial intelligence","Genetics"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[{"name":"geo"},{"name":"igsr"},{"name":"doi"}],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-23T18:13:11.021297Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}