{"doi":"10.1101/2023.04.13.536784","title":"Assessing transcriptomic re-identification risks using discriminative sequence models","abstract":"Abstract Gene expression data provides molecular insights into the functional impact of genetic variation, for example through expression quantitative trait loci (eQTL). With an improving understanding of the association between genotypes and gene expression comes a greater concern that gene expression profiles could be matched to genotype profiles of the same individuals in another dataset, known as a linking attack. Prior works demonstrating such a risk could analyze only a fraction of eQTLs that are independent due to restrictive model assumptions, leaving the full extent of this risk incompletely understood. To address this challenge, we introduce the discriminative sequence model (DSM), a novel probabilistic framework for predicting a sequence of genotypes based on gene expression data. By modeling the joint distribution over all known eQTLs in a genomic region, DSM improves the power of linking attacks with necessary calibration for linkage disequilibrium and redundant predictive signals. We demonstrate greater linking accuracy of DSM compared to existing approaches across a range of attack scenarios and datasets including up to 22K individuals, suggesting that DSM helps uncover a substantial additional risk overlooked by previous studies. Our work provides a unified framework for assessing the privacy risks of sharing diverse omics datasets beyond transcriptomics.","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2023,"id":394441,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9401,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1169031,"name":"Daniel Fridman","orcid":"0000-0002-8669-9478","position":1,"is_corresponding":false},{"id":3585,"name":"Bonnie Berger","orcid":"0000-0002-2724-7228","position":2,"is_corresponding":false},{"id":230752,"name":"Hyunghoon Cho","orcid":"0000-0002-2713-0150","position":3,"is_corresponding":false},{"id":554095,"name":"Shuvom Sadhuka","orcid":"0009-0001-9225-8716","position":0,"is_corresponding":true}],"reference_count":50,"raw_metadata":null,"created_at":"2026-07-19T01:19:14.590635Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}