{"doi":"10.1093/aje/kwaf043","title":"Pitfalls of imputing using incomplete auxiliary variables","abstract":"To the Editor: We enjoyed reading Madley-Dowd et al.’s1 excellent paper on the dangers of including incomplete auxiliary variables in imputation models. Their simulations explored the performance of imputation estimators of the association between a complete exposure variable, X, and an incomplete outcome, Y, when including an incomplete auxiliary variable, Z, in the imputation model. Surprisingly, when Y was independent of its missingness indicator given X and Z, then even if the missingness indicator for Z itself was independent from all other variables (measured and unmeasured), imputation estimates were substantially biased (Figure 1). It is striking that standard imputation algorithms fail for this graph, because it is possible—and indeed, straightforward—to design an unbiased imputation estimator for this graph.2,3,4,5 Let Z and Y denote the observed versions of the auxiliary variable and outcome, which contain missing values, and let |${Z}^{(1)}$| and |${Y}^{(1)}$| be their unobserved, underlying counterparts that do not have missing data. Their missingness indicators are |${M}_Z$| and |${M}_Y$|⁠, respectively. Thus, |$Z={Z}^{(1)}$| if |${M}_Z=1$| and Z = ? (missing) if |${M}_Z=0$|⁠, and likewise for Y. The conditional distribution of interest is identified by the g-formula:","journal":"American Journal of Epidemiology","year":2025,"id":564038,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9514,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":332299,"name":"Ilya Shpitser","orcid":"0000-0003-2571-7326","position":1,"is_corresponding":false},{"id":32607,"name":"Maya B. Mathur","orcid":"0000-0001-6698-2607","position":0,"is_corresponding":true}],"reference_count":16,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-19T02:56:17.117043Z","pmid":"40207659","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}