{"doi":"10.1145/3743093.3771065","title":"Position-Enhanced Gradient Attack (PEGA) on Medical Language Models","abstract":"Federated Learning (FL) enables collaborative training of language models on sensitive clinical notes without sharing the data. However, this paradigm is vulnerable to gradient inversion attacks that can reconstruct private data from shared gradients. We find that state-of-the-art attacks are less effective in the medical domain, failing to overcome the unique challenges posed by its specialized vocabulary and unstructured format. To address this, we introduce the Position-Enhanced Gradient Attack (PEGA), a novel attack that makes gradients position-aware by optimizing token and position embeddings simultaneously. PEGA employs two key innovations: a periodic sorting of positional embeddings to resolve token order ambiguity and a late-stage embedding replacement strategy to correct hard-to-recover critical tokens. To evaluate the leakage of sensitive data more directly, we also propose the Unified PHI-Recall (UPHI), a new metric measuring the recovery of Protected Health Information. Experiments on the MIMIC-III dataset show that PEGA significantly outperforms leading attacks like TAG and LAMP, particularly in its ability to reconstruct identifiable patient information, exposing a more severe and nuanced privacy risk in federated medical NLP.","journal":null,"year":2025,"id":584383,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.956,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":346057,"name":"Christopher B. Stanley","orcid":"0000-0002-4226-7710","position":1,"is_corresponding":false},{"id":378744,"name":"John Gounley","orcid":"0000-0001-8424-4982","position":2,"is_corresponding":false},{"id":329182,"name":"Heidi A. Hanson","orcid":"0000-0003-0056-196X","position":3,"is_corresponding":false},{"id":1497210,"name":"Chang Ge","orcid":"0000-0001-8788-4379","position":4,"is_corresponding":false},{"id":1497211,"name":"Caiwen Ding","orcid":"0000-0003-0891-1231","position":5,"is_corresponding":false},{"id":1497209,"name":"Nuo Xu","orcid":"0000-0001-6148-2830","position":0,"is_corresponding":true}],"reference_count":7,"raw_metadata":null,"created_at":"2026-07-19T02:59:11.978098Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}