{"doi":"10.1093/hmg/ddaa192","title":"Implicit bias of encoded variables: frameworks for addressing structured bias in EHR–GWAS data","abstract":"<jats:title>Abstract</jats:title>\n               <jats:p>The ‘discovery’ stage of genome-wide association studies required amassing large, homogeneous cohorts. In order to attain clinically useful insights, we must now consider the presentation of disease within our clinics and, by extension, within our medical records. Large-scale use of electronic health record (EHR) data can help to understand phenotypes in a scalable manner, incorporating lifelong and whole-phenome context.</jats:p>\n               <jats:p>However, extending analyses to incorporate EHR and biobank-based analyses will require careful consideration of phenotype definition. Judgements and clinical decisions that occur ‘outside’ the system inevitably contain some degree of bias and become encoded in EHR data. Any algorithmic approach to phenotypic characterization that assumes non-biased variables will generate compounded biased conclusions. Here, we discuss and illustrate potential biases inherent within EHR analyses, how these may be compounded across time and suggest frameworks for large-scale phenotypic analysis to minimize and uncover encoded bias.</jats:p>","journal":"Human Molecular Genetics","year":2020,"id":619561,"datarank":0.519860385419959,"base_score":3.4657359027997265,"endowment":3.4657359027997265,"self_citation_contribution":0.519860385419959,"citation_network_contribution":0.0,"self_endowment_contribution":0.519860385419959,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":31,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":531739,"name":"Carina Seah","orcid":"0000-0003-1604-1838","position":1,"is_corresponding":false},{"id":1598924,"name":"Jessica S Johnson","orcid":null,"position":2,"is_corresponding":false},{"id":35431,"name":"Laura M. Huckins","orcid":"0000-0002-5369-6502","position":3,"is_corresponding":false},{"id":1598922,"name":"Hillary R Dueñas","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Implicit bias of encoded variables: frameworks for addressing structured bias in EHR–GWAS data","abstract":"<jats:title>Abstract</jats:title>\n               <jats:p>The ‘discovery’ stage of genome-wide association studies required amassing large, homogeneous cohorts. In order to attain clinically useful insights, we must now consider the presentation of disease within our clinics and, by extension, within our medical records. Large-scale use of electronic health record (EHR) data can help to understand phenotypes in a scalable manner, incorporating lifelong and whole-phenome context.</jats:p>\n               <jats:p>However, extending analyses to incorporate EHR and biobank-based analyses will require careful consideration of phenotype definition. Judgements and clinical decisions that occur ‘outside’ the system inevitably contain some degree of bias and become encoded in EHR data. Any algorithmic approach to phenotypic characterization that assumes non-biased variables will generate compounded biased conclusions. Here, we discuss and illustrate potential biases inherent within EHR analyses, how these may be compounded across time and suggest frameworks for large-scale phenotypic analysis to minimize and uncover encoded bias.</jats:p>","is_dataset_classified":null,"base_score":3.4657359027997265,"endowment":3.4657359027997265,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"32879975","pmcid":"PMC7530523","openalex_id":"https://openalex.org/W3083461656","authors":[],"funders":[{"funder_name":"National Institute of Mental Health","grant_id":"R01MH121923","title":null},{"funder_name":"National Institutes of Health","grant_id":"5R01MH121923-03","title":"3/4: Leveraging EHR-linked biobanks for deep phenotyping, polygenic risk score modeling, and outcomes analysis in psychiatric disorders"}],"total_grants":2,"fwci":2.4023,"citation_percentile":0.89177815,"influential_citations":0,"citation_trend":[{"year":2021,"count":3},{"year":2022,"count":8},{"year":2023,"count":4},{"year":2024,"count":6},{"year":2025,"count":8},{"year":2026,"count":2}],"oa_status":"hybrid","license":"cc-by-nc","oa_locations":[{"url":"https://academic.oup.com/hmg/article-pdf/29/R1/R33/33939096/ddaa192.pdf","host_type":"journal"},{"url":"https://academic.oup.com/hmg/article-pdf/29/R1/R33/33939096/ddaa192.pdf","host_type":"publisher"},{"url":"http://academic.oup.com/hmg/article-pdf/29/R1/R33/33939096/ddaa192.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/hmg/ddaa192","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/32879975","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/7530523","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC7530523","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC7530523?pdf=render","host_type":"Europe_PMC"},{"url":"http://dx.doi.org/10.1093/hmg/ddaa192","host_type":""},{"url":"https://dx.doi.org/10.1093/hmg/ddaa192","host_type":""}],"fields_of_study":["Genetic Associations and Epidemiology","Functional Brain Connectivity Studies","Neonatal and fetal brain pathology","0301 basic medicine","0303 health sciences","03 medical and health sciences","Computational Biology","Disease","Electronic Health Records","Genome-Wide Association Study","Humans","Phenotype","Polymorphism, Single Nucleotide","Prejudice"],"mesh_terms":["Disease","Humans","Phenotype","Prejudice","Computational Biology","Polymorphism, Single Nucleotide","Genome-Wide Association Study","Electronic Health Records"],"keywords":["Biology","Genome-wide association study","Computational biology","MEDLINE","Genetics","Single-nucleotide polymorphism","Gene","Genotype","Phenotype","Invited Review Article","Electronic Health Records","Humans","Disease","Polymorphism, Single Nucleotide","Prejudice"],"sdg_mappings":[{"sdg_number":3,"sdg_label":"3. Good health"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[{"name":"doi"}],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-03T07:36:43.118182Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}