{"doi":"10.1002/hsr2.1438","title":"Machine learning prediction of mild cognitive impairment and its progression to Alzheimer's disease","abstract":"It is estimated that the number of people with dementia will reach 78 million by 2030 and 139 million by 2050, costing over 2.8 trillion dollars worldwide.1 Effective screening for mild cognitive impairment (MCI) as a risk factor for developing Alzheimer's disease (AD) is a crucial step in helping aging population with their needs.2 Early detection and automated screening for MCI and dementia could offer opportunities for deliberate study and recruitment into trials for developing other potentially useful therapeutics or interventions.3-5 Here, we systematically compare multiple automated machine learning (ML) models in predicting MCI and its progression to AD using real-world structured and unstructured electronic health records (EHRs) data. Our objective is to comprehensively evaluate the predictive accuracy, measured by the area under the curve (AUC) of the receiver operating characteristic (ROC), for future MCI and progression to AD based on routine EHR data, among a diverse population of primary care patients aged 65 years or older. This is a retrospective cohort study using Stanford Healthcare data from 1999 to 2022. The use of this data for this study was approved by Stanford's Institutional Review Board. Our data are formatted in the Observational Medical Outcomes Partnership (OMOP) model.6 The cohort consists of 157,804 (MCI and non-MCI) patients, who had at least one primary care visit after reaching the age of 65; with an average age of 73 and 57.7% were females. 15.1% of patients were Asian, 6.4% were Black, 0.2% were American Indian, 0.9% were Native Hawaiian, 64.3% were White, and 13.1% had other/unknown races or declined to state their race. Our study includes two main components: (a) MCI prediction and (b) MCI to AD progression prediction. We extracted 531,387 primary care visits (for all 157,804 patients in our cohort; each patient has multiple visits) where the patients were at least 65 years old at the time of their appointment. All historical EHR records, including diagnoses, prescriptions, procedures, and clinical notes before the primary care visits, were extracted. Note clinical note features are pre-processed and extracted in the form of standardized SNOMED structure concepts from patients' notes as part of OMOP data model.7 The OMOP Common Data Model standardizes healthcare data for research. By standardizing the representation of patient information and healthcare data elements, OMOP enables researchers to produce reliable evidence, conduct large-scale and multisite studies, and develop predictive models using data from multiple institutions, enhancing our understanding of health outcomes and treatment effectiveness. MCI prediction component was created using supervised ML models including logistic regression,8 random forest,9 and xgboost10 to predict MCI diagnosis within 1 year of primary care visit and using 480 predictors extracted from structured and unstructured EHR data. Models were trained using data in or before 2019 and tested using data in 2020 and after. The second component, MCI to AD progression prediction model, was trained using 7425 MCI patients' data and 373 predictors extracted from structured and unstructured EHR data before MCI onset. Further, we analyzed and presented possible risk factors for progression from MCI to AD in our data. Table 1 shows the MCI and MCI to AD progression prediction results. Random forest was the best-performing model in predicting MCI onset as well as predicting its progression to AD. Additionally, we utilized age-stratified test data to evaluate the performance of our models. We divided our test data sets into distinct age groups (65–74, 75–85, and 85+ years old), and tested our models separately on each age group. For MCI prediction, the random forest model outperformed the other models in the age groups of 65–74 (ROC-AUC = 64.3 ± $\\,\\pm \\,$ 1.2), 75–84 (ROC-AUC = 60.6 ± $\\,\\pm \\,$ 1.4), and 85 years and older (ROC-AUC = 60.8 ± $\\,\\pm \\,$ 2.2). Similarl","journal":"Health Science Reports","year":2023,"id":343451,"datarank":0.6798541599588015,"base_score":2.70805020110221,"endowment":2.70805020110221,"self_citation_contribution":0.40620753016533157,"citation_network_contribution":0.2736466297934699,"self_endowment_contribution":0.40620753016533157,"citer_contribution":0.2736466297934699,"corpus_percentile":null,"corpus_rank":null,"citation_count":14,"citer_count":13,"citers_with_citation_signal":7,"citers_with_endowment":7,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9586,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":717334,"name":"Morteza Noshad","orcid":null,"position":1,"is_corresponding":false},{"id":1081455,"name":"VJ Periyakoil","orcid":null,"position":2,"is_corresponding":false},{"id":227523,"name":"Jonathan H. Chen","orcid":"0000-0002-4387-8740","position":3,"is_corresponding":false},{"id":513828,"name":"Sajjad Fouladvand","orcid":"0000-0002-9869-1836","position":0,"is_corresponding":true}],"reference_count":14,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-19T01:11:21.758449Z","pmid":"37867782","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}