{"doi":"10.1093/jamia/ocaa180","title":"Representation of EHR data for predictive modeling: a comparison between UMLS and other terminologies","abstract":"OBJECTIVE: Predictive disease modeling using electronic health record data is a growing field. Although clinical data in their raw form can be used directly for predictive modeling, it is a common practice to map data to standard terminologies to facilitate data aggregation and reuse. There is, however, a lack of systematic investigation of how different representations could affect the performance of predictive models, especially in the context of machine learning and deep learning. MATERIALS AND METHODS: We projected the input diagnoses data in the Cerner HealthFacts database to Unified Medical Language System (UMLS) and 5 other terminologies, including CCS, CCSR, ICD-9, ICD-10, and PheWAS, and evaluated the prediction performances of these terminologies on 2 different tasks: the risk prediction of heart failure in diabetes patients and the risk prediction of pancreatic cancer. Two popular models were evaluated: logistic regression and a recurrent neural network. RESULTS: For logistic regression, using UMLS delivered the optimal area under the receiver operating characteristics (AUROC) results in both dengue hemorrhagic fever (81.15%) and pancreatic cancer (80.53%) tasks. For recurrent neural network, UMLS worked best for pancreatic cancer prediction (AUROC 82.24%), second only (AUROC 85.55%) to PheWAS (AUROC 85.87%) for dengue hemorrhagic fever prediction. DISCUSSION/CONCLUSION: In our experiments, terminologies with larger vocabularies and finer-grained representations were associated with better prediction performances. In particular, UMLS is consistently 1 of the best-performing ones. We believe that our work may help to inform better designs of predictive models, although further investigation is warranted.","journal":"Journal of the American Medical Informatics Association","year":2020,"id":97334,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":33,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.951,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2020-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":49267,"name":"Firat Tiryaki","orcid":null,"position":1,"is_corresponding":false},{"id":347167,"name":"Yujia Zhou","orcid":"0000-0003-0889-2261","position":2,"is_corresponding":false},{"id":478778,"name":"Yang Xiang","orcid":"0000-0001-5252-0831","position":3,"is_corresponding":false},{"id":23317,"name":"Cui Tao","orcid":"0000-0002-4267-1924","position":4,"is_corresponding":false},{"id":12534,"name":"Hua Xu","orcid":"0000-0002-5274-4672","position":5,"is_corresponding":false},{"id":24932,"name":"Degui Zhi","orcid":"0000-0001-7754-1890","position":6,"is_corresponding":false},{"id":347169,"name":"Laila Rasmy","orcid":"0000-0002-2644-4908","position":0,"is_corresponding":true}],"reference_count":27,"raw_metadata":null,"created_at":"2026-07-18T22:35:34.520494Z","pmid":"32930711","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}