{"doi":"10.1109/access.2021.3093005","title":"Fast and Scalable Private Genotype Imputation Using Machine Learning and Partially Homomorphic Encryption","abstract":"The recent advances in genome sequencing technologies provide unprecedented opportunities to understand the relationship between human genetic variation and diseases. However, genotyping whole genomes from a large cohort of individuals is still cost prohibitive. Imputation methods to predict genotypes of missing genetic variants are widely used, especially for genome-wide association studies. Accurate genotype imputation requires complex statistical methods. Due to the data and computing-intensive nature of the problem, imputation is increasingly outsourced, raising serious privacy concerns. In this work, we investigate solutions for fast, scalable, and accurate privacy-preserving genotype imputation using Machine Learning (ML) and a standardized homomorphic encryption scheme, Paillier cryptosystem. ML-based privacy-preserving inference has been largely optimized for computation-heavy non-linear functions in a single-output multi-class classification setting. However, having a large number of multi-class outputs per genome per individual calls for further optimizations and/or approximations specific to this application. Here we explore the effectiveness of linear models for genotype imputation to convert them to privacy-preserving equivalents using standardized homomorphic encryption schemes. Our results show that performance of our privacy-preserving genotype imputation method is equivalent to the state-of-the-art plaintext solutions, achieving up to 99% micro area under curve score, even on real-world large-scale datasets up to 80,000 targets.","journal":"IEEE Access","year":2021,"id":162603,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":46,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9444,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":668906,"name":"Eduardo Chielle","orcid":"0000-0002-1938-912X","position":1,"is_corresponding":false},{"id":6259,"name":"Gamze Gürsoy","orcid":"0000-0002-1352-8686","position":2,"is_corresponding":false},{"id":680062,"name":"Oleg Mazonka","orcid":"0000-0001-5131-9044","position":3,"is_corresponding":false},{"id":108504,"name":"Mark Gerstein","orcid":"0000-0002-9746-3719","position":4,"is_corresponding":false},{"id":668907,"name":"Michail Maniatakos","orcid":"0000-0001-6899-0651","position":5,"is_corresponding":false},{"id":680061,"name":"Esha Sarkar","orcid":"0000-0002-4473-7368","position":0,"is_corresponding":true}],"reference_count":39,"raw_metadata":null,"created_at":"2026-07-18T23:45:05.641178Z","pmid":"34476144","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}