{"doi":"10.1093/jamia/ocag067","title":"Disparate language and model effects on AI-based translation and recognition of genetic conditions","abstract":"<jats:title>Abstract</jats:title>\n                  <jats:sec>\n                    <jats:title>Introduction</jats:title>\n                    <jats:p>Artificial intelligence (AI) is increasingly prevalent. Patients and clinicians may use AI-based tools in many different languages.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Objective</jats:title>\n                    <jats:p>To investigate AI translation tools for descriptions of genetic conditions and how AI identification of genetic conditions is affected by translations.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Materials and Methods</jats:title>\n                    <jats:p>We used Neural machine translation (NMT) and large language-model (LLM) translation to translate descriptions of 40 genetic conditions into 191 and 93 languages, respectively. Excluding translations retaining English medical terms verbatim, we respectively focused on 139 and 70 languages. After assessing translations, we assessed the ability of 3 proprietary and 3 open-weight general LLMs to identify conditions in the translations. We analyzed how accuracy was affected by the conditions’ prevalence in the literature, and attributes of the languages (the script, language family, and prevalence of the language in training sources). We also investigated adaptive translation for select languages.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Results</jats:title>\n                    <jats:p>We found significant differences in condition identification based on the translation method, condition, language, and prediction model. The accuracy of some models was more affected than others by factors like the conditions’ literature prevalence, language script, family, and language prevalence. Adaptive translation for select languages did not improve translations or diagnostic accuracy with the 3 tested LLMs. However, further analysis with 1 language showed that this approach was more effective with smaller LLMs.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions</jats:title>\n                    <jats:p>AI-based translation has variable performance, which can affect the ability of AI models to recognize genetic conditions. These findings should inform safe medical AI use to support consistent performance in different languages.</jats:p>\n                  </jats:sec>","journal":"Journal of the American Medical Informatics Association","year":2026,"id":609448,"datarank":0.16479184330021646,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"self_citation_contribution":0.16479184330021646,"citation_network_contribution":0.0,"self_endowment_contribution":0.16479184330021646,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":422653,"name":"Irini Manoli","orcid":"0000-0003-1543-2941","position":1,"is_corresponding":false},{"id":1566448,"name":"Shubha R Phadke","orcid":null,"position":2,"is_corresponding":false},{"id":445802,"name":"Chanika Phornphutkul","orcid":"0000-0002-0130-9101","position":3,"is_corresponding":false},{"id":1566449,"name":"Jonathan D Raymond","orcid":null,"position":4,"is_corresponding":false},{"id":1566450,"name":"Benjamin D Solomon","orcid":null,"position":5,"is_corresponding":false},{"id":719367,"name":"Dat Duong","orcid":"0000-0003-3022-9378","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Disparate language and model effects on AI-based translation and recognition of genetic conditions","abstract":"<jats:title>Abstract</jats:title>\n                  <jats:sec>\n                    <jats:title>Introduction</jats:title>\n                    <jats:p>Artificial intelligence (AI) is increasingly prevalent. Patients and clinicians may use AI-based tools in many different languages.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Objective</jats:title>\n                    <jats:p>To investigate AI translation tools for descriptions of genetic conditions and how AI identification of genetic conditions is affected by translations.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Materials and Methods</jats:title>\n                    <jats:p>We used Neural machine translation (NMT) and large language-model (LLM) translation to translate descriptions of 40 genetic conditions into 191 and 93 languages, respectively. Excluding translations retaining English medical terms verbatim, we respectively focused on 139 and 70 languages. After assessing translations, we assessed the ability of 3 proprietary and 3 open-weight general LLMs to identify conditions in the translations. We analyzed how accuracy was affected by the conditions’ prevalence in the literature, and attributes of the languages (the script, language family, and prevalence of the language in training sources). We also investigated adaptive translation for select languages.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Results</jats:title>\n                    <jats:p>We found significant differences in condition identification based on the translation method, condition, language, and prediction model. The accuracy of some models was more affected than others by factors like the conditions’ literature prevalence, language script, family, and language prevalence. Adaptive translation for select languages did not improve translations or diagnostic accuracy with the 3 tested LLMs. However, further analysis with 1 language showed that this approach was more effective with smaller LLMs.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions</jats:title>\n                    <jats:p>AI-based translation has variable performance, which can affect the ability of AI models to recognize genetic conditions. These findings should inform safe medical AI use to support consistent performance in different languages.</jats:p>\n                  </jats:sec>","is_dataset_classified":null,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"42090313","pmcid":null,"openalex_id":"https://openalex.org/W7160430808","authors":[],"funders":[{"funder_name":"National Human Genome Research Institute of the National Institutes of Health","grant_id":"","title":null},{"funder_name":"Intramural Research Program of the National Human Genome Research Institute of the National Institutes of Health","grant_id":"","title":null}],"total_grants":2,"fwci":14.2794,"citation_percentile":0.99026764,"influential_citations":0,"citation_trend":[{"year":2026,"count":2}],"oa_status":"green","license":"public-domain","oa_locations":[{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC13317944/","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC13317944/","host_type":"repository"},{"url":"https://academic.oup.com/jamia/advance-article-pdf/doi/10.1093/jamia/ocag067/68231624/ocag067.pdf","host_type":"publisher"},{"url":"https://academic.oup.com/jamia/article-pdf/33/7/1284/68231624/ocag067.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/jamia/ocag067","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/42090313","host_type":"repository"}],"fields_of_study":["Genomics and Rare Diseases","Artificial Intelligence in Healthcare and Education","Genetic Associations and Epidemiology"],"mesh_terms":["Large Language Models","Intelligent Systems","Artificial Intelligence","Humans","Language","Natural Language Processing","Translating","Neural Networks, Computer"],"keywords":["Translation (biology)","Identification (biology)","Machine translation","Affect (linguistics)","Genetic algorithm","Translation","Artificial intelligence","Medical genetics","Medical Genomics","Large Language Model"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Quality Education"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-31T07:58:05.769050Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}