{"doi":"10.1093/nar/gkaa881","title":"A flexible, interpretable, and accurate approach for imputing the expression of unmeasured genes","abstract":"While there are >2 million publicly-available human microarray gene-expression profiles, these profiles were measured using a variety of platforms that each cover a pre-defined, limited set of genes. Therefore, key to reanalyzing and integrating this massive data collection are methods that can computationally reconstitute the complete transcriptome in partially-measured microarray samples by imputing the expression of unmeasured genes. Current state-of-the-art imputation methods are tailored to samples from a specific platform and rely on gene-gene relationships regardless of the biological context of the target sample. We show that sparse regression models that capture sample-sample relationships (termed SampleLASSO), built on-the-fly for each new target sample to be imputed, outperform models based on fixed gene relationships. Extensive evaluation involving three machine learning algorithms (LASSO, k-nearest-neighbors, and deep-neural-networks), two gene subsets (GPL96-570 and LINCS), and multiple imputation tasks (within and across microarray/RNA-seq datasets) establishes that SampleLASSO is the most accurate model. Additionally, we demonstrate the biological interpretability of this method by showing that, for imputing a target sample from a certain tissue, SampleLASSO automatically leverages training samples from the same tissue. Thus, SampleLASSO is a simple, yet powerful and flexible approach for harmonizing large-scale gene-expression data.","journal":"Nucleic Acids Research","year":2020,"id":87100,"datarank":0.7004765882792239,"base_score":2.5649493574615367,"endowment":2.5649493574615367,"self_citation_contribution":0.38474240361923057,"citation_network_contribution":0.31573418465999337,"self_endowment_contribution":0.38474240361923057,"citer_contribution":0.31573418465999337,"corpus_percentile":null,"corpus_rank":null,"citation_count":12,"citer_count":11,"citers_with_citation_signal":8,"citers_with_endowment":8,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.645,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2020-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":443449,"name":"Jacob L Canfield","orcid":"0000-0001-9018-0751","position":1,"is_corresponding":false},{"id":443450,"name":"Deepak Singla","orcid":"0000-0001-7699-7079","position":2,"is_corresponding":false},{"id":34964,"name":"Krishnan, Arjun","orcid":"0000-0002-7980-4110","position":3,"is_corresponding":false},{"id":443448,"name":"Christopher A Mancuso","orcid":"0000-0003-3081-2758","position":0,"is_corresponding":true}],"reference_count":50,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-18T21:59:31.247492Z","pmid":"33074331","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}