{"doi":"10.1101/2023.10.21.562237","title":"Genotype prediction of 336,463 samples from public expression data","abstract":"Tens of thousands of RNA-sequencing experiments comprising hundreds of thousands of individual samples have now been performed. These data represent a broad range of experimental conditions, sequencing technologies, and hypotheses under study. The Recount project has aggregated and uniformly processed hundreds of thousands of publicly available RNA-seq samples. Most of these samples only include RNA expression measurements; genotype data for these same samples would enable a wide range of analyses including variant prioritization, eQTL analysis, and studies of allele specific expression. Here, we developed a statistical model based on the existing reference and alternative read counts from the RNA-seq experiments available through Recount3 to predict genotypes at autosomal biallelic loci in coding regions. We demonstrate the accuracy of our model using large-scale studies that measured both gene expression and genotype genome-wide. We show that our predictive model is highly accurate with 99.5% overall accuracy, 99.6% major allele accuracy, and 90.4% minor allele accuracy. Our model is robust to tissue and study effects, provided the coverage is high enough. We applied this model to genotype all the samples in Recount 3 and provide the largest ready-to-use expression repository containing genotype information. We illustrate that the predicted genotype from RNA-seq data is sufficient to unravel the underlying population structure of samples in Recount3 using Principal Component Analysis.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2023,"id":392687,"datarank":0.31777762967852,"base_score":1.791759469228055,"endowment":1.791759469228055,"self_citation_contribution":0.26876392038420827,"citation_network_contribution":0.04901370929431174,"self_endowment_contribution":0.26876392038420827,"citer_contribution":0.04901370929431174,"corpus_percentile":46.723911193625746,"corpus_rank":6888,"citation_count":5,"citer_count":4,"citers_with_citation_signal":2,"citers_with_endowment":2,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.944,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1166257,"name":"Christopher Lo","orcid":"0000-0001-8400-0902","position":1,"is_corresponding":false},{"id":561417,"name":"Siruo Wang","orcid":null,"position":2,"is_corresponding":false},{"id":16225,"name":"Jeffrey T. Leek","orcid":"0000-0002-2873-2671","position":3,"is_corresponding":false},{"id":29385,"name":"Kasper D. Hansen","orcid":"0000-0003-0086-0687","position":4,"is_corresponding":false},{"id":557830,"name":"Afrooz Razi","orcid":"0000-0003-4556-9383","position":0,"is_corresponding":true}],"reference_count":33,"raw_metadata":null,"created_at":"2026-07-19T01:18:56.730121Z","pmid":"38559266","pmcid":"PMC10979922","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}