{"doi":"10.64898/2026.05.28.26353952","title":"The Multimodal Anonymizer: a fully local multi-agent AI system for medical data deidentification","abstract":"<jats:title>Abstract</jats:title>\n                <jats:sec>\n                  <jats:title>Background</jats:title>\n                  <jats:p>Safe reuse of multimodal hospital data for AI development is limited by the absence of reliable, context-aware deidentification across multimodal data and longitudinal patient data. Existing approaches are largely modality-specific and can indiscriminately remove clinically important information.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Methods</jats:title>\n                  <jats:p>We developed the Multimodal Anonymizer, a modular, locally deployable multi-agent framework integrating multimodal large language models, task-specific neural networks and rule-based transformations. We evaluated 16 orchestrator model configurations on a benchmark built from publicly available data and hospital data from our institution. The benchmark dataset included data from different origins: 250 MIMIC-IV patients with synthetically injected personally identifiable information (PII) supplemented with head CT, face images, handwriting, audio, German clinical-text datasets and local data. Primary outcomes were deidentification sensitivity and preservation of clinically important content; secondary analyses examined model characteristics, reproducibility, and performance against leading market and open-source solutions.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>The best local configuration—the orchestrator being Qwen3-VL-235B-A22B-Thinking—achieved near-complete deidentification across all datasets, with per-patient sensitivity of 98.80% (95%-CI 97.20; 100), and per-PII sensitivity of 99.82% (95%-CI 99.76; 99.88). Critical clinical preservation was 99.60% (95%-CI 98.80; 100) per-patient, and clinical preservation was 99.61% (95%-CI 99.51; 99.71) per-file. All modalities achieved at least 98.30% sensitivity (lower bound 95%-CI). On our local data, the system achieved a deidentification sensitivity of 100% per-patient and per-PII; and a critical clinical preservation of 100% per-patient as well as a clinical preservation of 99.97% (95%-CI 99.91; 100) per-file. When comparing orchestrators, the leading local models were similar to proprietary models (GPT-5.2) in deidentification sensitivity while showing higher deidentification specificity. The Multimodal Anonymizer outperformed previous tools on most modalities.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Conclusion</jats:title>\n                  <jats:p>Near-complete, utility-preserving deidentification of multimodal clinical data is achievable with a unified, locally deployable multi-agent system, enabling safer large-scale reuse of hospital data for research and AI development.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Graphical Abstract</jats:title>\n                  <jats:fig id=\"ufig1\" position=\"float\" orientation=\"portrait\" fig-type=\"figure\">\n                    <jats:graphic xmlns:xlink=\"http://www.w3.org/1999/xlink\" xlink:href=\"26353952v1_ufig1\" position=\"float\" orientation=\"portrait\"/>\n                  </jats:fig>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Highlights</jats:title>\n                  <jats:list list-type=\"bullet\">\n                    <jats:list-item>\n                      <jats:p>Framework for deidentification of multimodal clinical data.</jats:p>\n                    </jats:list-item>\n                    <jats:list-item>\n                      <jats:p>Multimodal deidentification with preservation of clinically relevant content.</jats:p>\n                    </jats:list-item>\n                    <jats:list-item>\n                      <jats:p>On-premises plug-and-play deployment for local data processing.</jats:p>\n                    </jats:list-item>\n                    <jats:list-item>\n                      <jats:p>Evaluation of 16 model configurations and comparison with existing tools.</jats:p>\n                    </jats:list-item>\n                    <jats:list-item>\n                      <jats:p>Assessment on external, multilingual and site-specific datasets.</jats:p>\n                    </jats:list-item>\n                  </jats:list>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Short Description</jats:title>\n                  <jats:p>The Multimodal Anonymizer is a fully local, multi-agent system that prepares multimodal clinical records for privacy-preserving reuse by coordinating multimodal large language model reasoning, specialist neural networks, rule-based transformations, and iterative verification. Across benchmarks spanning text, tables, PDFs, imaging, metadata, filenames, audio, and handwriting, its best configuration using a local open-source multimodal large language model achieved 98.80% patient-level deidentification sensitivity and 99.60% preservation of clinically critical content, performing comparably to proprietary models and outperforming established deidentification tools across most modalities.</jats:p>\n                </jats:sec>","journal":null,"year":null,"id":654048,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1706829,"name":"Foo Wei Ten","orcid":null,"position":1,"is_corresponding":false},{"id":1706830,"name":"Kyle S. Krüger","orcid":null,"position":2,"is_corresponding":false},{"id":1706831,"name":"Robin C. Geyer","orcid":null,"position":3,"is_corresponding":false},{"id":1706832,"name":"Tobias Röschl","orcid":null,"position":4,"is_corresponding":false},{"id":1706833,"name":"Matthias I Gröschel","orcid":null,"position":5,"is_corresponding":false},{"id":473339,"name":"Paul Rostin","orcid":"0000-0002-1955-8333","position":6,"is_corresponding":false},{"id":13430,"name":"Roland Eils","orcid":"0000-0002-0034-4036","position":7,"is_corresponding":false},{"id":1706834,"name":"Martin Spott","orcid":null,"position":8,"is_corresponding":false},{"id":38092,"name":"Fabian Prasser","orcid":"0000-0003-3172-3095","position":9,"is_corresponding":false},{"id":1217304,"name":"Alexander Meyer","orcid":"0000-0002-6944-2478","position":10,"is_corresponding":false},{"id":1706835,"name":"Julian Madrid","orcid":"0000-0001-5135-6873","position":11,"is_corresponding":false},{"id":1706828,"name":"Anja Hirsch","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"The Multimodal Anonymizer: a fully local multi-agent AI system for medical data deidentification","abstract":"<jats:title>Abstract</jats:title>\n                <jats:sec>\n                  <jats:title>Background</jats:title>\n                  <jats:p>Safe reuse of multimodal hospital data for AI development is limited by the absence of reliable, context-aware deidentification across multimodal data and longitudinal patient data. Existing approaches are largely modality-specific and can indiscriminately remove clinically important information.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Methods</jats:title>\n                  <jats:p>We developed the Multimodal Anonymizer, a modular, locally deployable multi-agent framework integrating multimodal large language models, task-specific neural networks and rule-based transformations. We evaluated 16 orchestrator model configurations on a benchmark built from publicly available data and hospital data from our institution. The benchmark dataset included data from different origins: 250 MIMIC-IV patients with synthetically injected personally identifiable information (PII) supplemented with head CT, face images, handwriting, audio, German clinical-text datasets and local data. Primary outcomes were deidentification sensitivity and preservation of clinically important content; secondary analyses examined model characteristics, reproducibility, and performance against leading market and open-source solutions.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>The best local configuration—the orchestrator being Qwen3-VL-235B-A22B-Thinking—achieved near-complete deidentification across all datasets, with per-patient sensitivity of 98.80% (95%-CI 97.20; 100), and per-PII sensitivity of 99.82% (95%-CI 99.76; 99.88). Critical clinical preservation was 99.60% (95%-CI 98.80; 100) per-patient, and clinical preservation was 99.61% (95%-CI 99.51; 99.71) per-file. All modalities achieved at least 98.30% sensitivity (lower bound 95%-CI). On our local data, the system achieved a deidentification sensitivity of 100% per-patient and per-PII; and a critical clinical preservation of 100% per-patient as well as a clinical preservation of 99.97% (95%-CI 99.91; 100) per-file. When comparing orchestrators, the leading local models were similar to proprietary models (GPT-5.2) in deidentification sensitivity while showing higher deidentification specificity. The Multimodal Anonymizer outperformed previous tools on most modalities.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Conclusion</jats:title>\n                  <jats:p>Near-complete, utility-preserving deidentification of multimodal clinical data is achievable with a unified, locally deployable multi-agent system, enabling safer large-scale reuse of hospital data for research and AI development.</jats:p>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Graphical Abstract</jats:title>\n                  <jats:fig id=\"ufig1\" position=\"float\" orientation=\"portrait\" fig-type=\"figure\">\n                    <jats:graphic xmlns:xlink=\"http://www.w3.org/1999/xlink\" xlink:href=\"26353952v1_ufig1\" position=\"float\" orientation=\"portrait\"/>\n                  </jats:fig>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Highlights</jats:title>\n                  <jats:list list-type=\"bullet\">\n                    <jats:list-item>\n                      <jats:p>Framework for deidentification of multimodal clinical data.</jats:p>\n                    </jats:list-item>\n                    <jats:list-item>\n                      <jats:p>Multimodal deidentification with preservation of clinically relevant content.</jats:p>\n                    </jats:list-item>\n                    <jats:list-item>\n                      <jats:p>On-premises plug-and-play deployment for local data processing.</jats:p>\n                    </jats:list-item>\n                    <jats:list-item>\n                      <jats:p>Evaluation of 16 model configurations and comparison with existing tools.</jats:p>\n                    </jats:list-item>\n                    <jats:list-item>\n                      <jats:p>Assessment on external, multilingual and site-specific datasets.</jats:p>\n                    </jats:list-item>\n                  </jats:list>\n                </jats:sec>\n                <jats:sec>\n                  <jats:title>Short Description</jats:title>\n                  <jats:p>The Multimodal Anonymizer is a fully local, multi-agent system that prepares multimodal clinical records for privacy-preserving reuse by coordinating multimodal large language model reasoning, specialist neural networks, rule-based transformations, and iterative verification. Across benchmarks spanning text, tables, PDFs, imaging, metadata, filenames, audio, and handwriting, its best configuration using a local open-source multimodal large language model achieved 98.80% patient-level deidentification sensitivity and 99.60% preservation of clinically critical content, performing comparably to proprietary models and outperforming established deidentification tools across most modalities.</jats:p>\n                </jats:sec>","is_dataset_classified":null,"base_score":0.0,"endowment":0.0,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":null,"pmcid":null,"openalex_id":null,"authors":[],"funders":[],"total_grants":0,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[],"oa_status":"green","license":"cc-by-nc-nd","oa_locations":[{"url":"https://doi.org/10.64898/2026.05.28.26353952","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.64898/2026.05.28.26353952","host_type":"publisher"}],"fields_of_study":[],"mesh_terms":[],"keywords":[],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-11T03:02:59.211730Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}