{"doi":"10.1101/2022.01.18.476789","title":"SERM: a self-consistent deep learning solution for rapid and accurate gene expression recovery","abstract":"ABSTRACT Single cell RNA sequencing (scRNA-seq) is a promising technique to determine the states of individual cells and classify novel cell subtypes. Computationally, the processing of scRNA-seq data presents a daunting challenge because of the noisy nature and humongous size and dimensionality of the data. Compromised solution by omitting the genes with low expression is commonly taken in current scRNA-seq analysis, which leads to inaccurate gene counts. In this paper, we introduce a broadly applicable data-driven gene expression recovery framework, referred to as the self-consistent expression recovery machine (SERM), to impute the missing gene expression. Using deep learning, SERM first learns from a subset of the noisy gene expression data to estimate the underlying data distribution. SERM then recovers the overall gene expression data by imposing a self-consistency on the gene expression matrix, thus ensuring that the expression levels are similarly distributed in different parts of the matrix. We show that SERM significantly improves the accuracy of gene imputation with at least 100-fold increase in computational efficiency in comparison to the state-of-the-art techniques. Thus SERM promises to provide an urgently needed computational solution for rapid and accurate recovery of big genomic expression data. SERM is available as a web-based computational tool ( https://www.analyxus.com/compute/serm ) and its source codes can be found in https://github.com/xinglab-ai/self-consistent-expression-recovery-machine .","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2022,"id":305503,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9521,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2022-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":692481,"name":"Jen‐Yeu Wang","orcid":"0000-0002-2278-4470","position":1,"is_corresponding":false},{"id":946115,"name":"Hongyi Ren","orcid":"0000-0001-6999-0370","position":2,"is_corresponding":false},{"id":258393,"name":"Xiaomeng Li","orcid":"0000-0003-1105-8083","position":3,"is_corresponding":false},{"id":239067,"name":"Masoud Badiei Khuzani","orcid":null,"position":4,"is_corresponding":false},{"id":946116,"name":"Shengtian Sang","orcid":"0000-0003-0851-9520","position":5,"is_corresponding":false},{"id":258391,"name":"Lequan Yu","orcid":"0000-0002-9315-6527","position":6,"is_corresponding":false},{"id":285647,"name":"Liyue Shen","orcid":"0000-0001-5942-3196","position":7,"is_corresponding":false},{"id":301438,"name":"Wei Zhao","orcid":"0000-0002-6182-4746","position":8,"is_corresponding":false},{"id":258394,"name":"Lei Xing","orcid":"0000-0003-2536-5359","position":9,"is_corresponding":false},{"id":724558,"name":"Md Tauhidul Islam","orcid":"0000-0001-6259-632X","position":0,"is_corresponding":true}],"reference_count":42,"raw_metadata":null,"created_at":"2026-07-19T00:32:41.290883Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}