{"doi":"10.1101/2020.07.06.189639","title":"Compression of quantification uncertainty for scRNA-seq counts","abstract":"Abstract Motivation Quantification estimates of gene expression from single-cell RNA-seq (scRNA-seq) data have inherent uncertainty due to reads that map to multiple genes. Many existing scRNA-seq quantification pipelines ignore multi-mapping reads and therefore underestimate expected read counts for many genes. alevin accounts for multi-mapping reads and allows for the generation of “inferential replicates”, which reflect quantification uncertainty. Previous methods have shown improved performance when incorporating these replicates into statistical analyses, but storage and use of these replicates increases computation time and memory requirements. Results We demonstrate that storing only the mean and variance from a set of inferential replicates (“compression”) is sufficient to capture gene-level quantification uncertainty. Using these values, we generate “pseudo-inferential” replicates from a negative binomial distribution and propose a general procedure for incorporating these replicates into a proposed statistical testing framework. We show reduced false positives when applying this procedure to trajectory-based differential expression analyses. We additionally extend the Swish method to incorporate pseudo-inferential replicates and demonstrate improvements in computation time and memory consumption without any loss in performance. Lastly, we show that the removal of multi-mapping reads can result in significant underestimation of counts for functionally important genes in a real dataset. Availability and implementation makeInfReps and splitSwish are implemented in the development branch of the R/Bioconductor fishpond package available at http://bioconductor.org/packages/devel/bioc/html/fishpond.html . Sample code to calculate the uncertainty-aware p -values can be found on GitHub at https://github.com/skvanburen/scUncertaintyPaperCode . Contact michaelisaiahlove@gmail.com","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2020,"id":125117,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9473,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2020-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":497433,"name":"Hirak Sarkar","orcid":"0000-0003-3636-7384","position":1,"is_corresponding":false},{"id":52013,"name":"Avi Srivastava","orcid":"0000-0001-9798-2079","position":2,"is_corresponding":false},{"id":457307,"name":"Naim U. Rashid","orcid":"0000-0002-7274-8259","position":3,"is_corresponding":false},{"id":87821,"name":"Rob Patro","orcid":"0000-0001-8463-1675","position":4,"is_corresponding":false},{"id":29945,"name":"Michael I. Love","orcid":"0000-0001-8401-0545","position":5,"is_corresponding":false},{"id":570815,"name":"Scott Van Buren","orcid":"0000-0002-5722-4342","position":0,"is_corresponding":true}],"reference_count":56,"raw_metadata":null,"created_at":"2026-07-18T23:15:15.482227Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}