{"doi":"10.1002/dad2.70171","title":"A vision transformer approach for fully automated and scalable dementia screening using clock drawing test images","abstract":"<jats:title>Abstract</jats:title><jats:sec><jats:title>INTRODUCTION</jats:title><jats:p>The clock drawing test (CDT) screens for dementia but requires trained scorers and lacks standardized criteria. Thus, we developed an automated vision transformer (ViT)‐based diagnostic system with convolutional neural network preprocessing for analyzing hand‐drawn CDT images.</jats:p></jats:sec><jats:sec><jats:title>METHODS</jats:title><jats:p>The architecture implements fine‐tuned ViT feature extraction with linear classification for dementia prediction. Training used the National Health and Aging Trends Study (NHATS) dataset (<jats:italic>n</jats:italic> = 54,027), with testing on an independent clinical cohort from the Toronto Dementia Research Alliance (TDRA; <jats:italic>n</jats:italic> = 862; 522 dementia, 340 normal cognition).</jats:p></jats:sec><jats:sec><jats:title>RESULTS</jats:title><jats:p>The ViT approach predicted dementia with 76.5% balanced accuracy, outperforming human‐scored features (74.3%) and existing deep learning models (MiniVGG = 73.3%, MobileNetV2 = 72.3%, relevance factor variational autoencoder = 69.1%) on the TDRA dataset.</jats:p></jats:sec><jats:sec><jats:title>DISCUSSION</jats:title><jats:p>This pen‐and‐paper compatible diagnostic system enables scalable remote cognitive screening through automated CDT image analysis that is competitive with human‐scored features, potentially increasing diagnostic accessibility for diverse populations across varied socioeconomic contexts.</jats:p></jats:sec><jats:sec><jats:title>HIGHLIGHTS</jats:title><jats:p><jats:list list-type=\"bullet\">\n<jats:list-item><jats:p>The vision transformer model achieves 76.5% accuracy in dementia detection from clock drawing tests, outperforming human scoring and existing deep learning methods.</jats:p></jats:list-item>\n<jats:list-item><jats:p>Novel convolutional neural network–based preprocessing automatically handles challenging image quality issues like shadows, irrelevant markings, and improper cropping.</jats:p></jats:list-item>\n<jats:list-item><jats:p>The system requires only a photo of a hand‐drawn clock test, enabling scalable remote screening accessible across socioeconomic contexts.</jats:p></jats:list-item>\n<jats:list-item><jats:p>A feature‐extraction model trained on 54,027 samples demonstrates robust generalization to an independent clinical dataset of 862 patients.</jats:p></jats:list-item>\n<jats:list-item><jats:p>This fully automated approach eliminates the need for trained scorers while maintaining diagnostic accuracy above manual methods.</jats:p></jats:list-item>\n</jats:list></jats:p></jats:sec>","journal":"Alzheimer's &amp; Dementia: Diagnosis, Assessment &amp; Disease Monitoring","year":2025,"id":612105,"datarank":0.24141568686511508,"base_score":1.6094379124341003,"endowment":1.6094379124341003,"self_citation_contribution":0.24141568686511508,"citation_network_contribution":0.0,"self_endowment_contribution":0.24141568686511508,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":4,"citer_count":3,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":652167,"name":"Morris Freedman","orcid":"0000-0003-4003-6158","position":1,"is_corresponding":false},{"id":431865,"name":"Sandra E. Black","orcid":"0000-0001-7093-8289","position":2,"is_corresponding":false},{"id":355527,"name":"Daniel Felsky","orcid":"0000-0003-1831-9848","position":3,"is_corresponding":false},{"id":112785,"name":"Sanjeev Kumar","orcid":null,"position":4,"is_corresponding":false},{"id":1151745,"name":"Bradley Pugh","orcid":null,"position":5,"is_corresponding":false},{"id":402085,"name":"Stephen C. Strother","orcid":"0000-0002-3198-217X","position":6,"is_corresponding":false},{"id":344569,"name":"David F. Tang‐Wai","orcid":"0000-0001-9747-9967","position":7,"is_corresponding":false},{"id":253935,"name":"Maria Carmela Tartaglia","orcid":"0000-0002-5944-8497","position":8,"is_corresponding":false},{"id":494024,"name":"Bradley R. Buchsbaum","orcid":"0000-0002-1108-4866","position":9,"is_corresponding":false},{"id":1575858,"name":"Michael B. Bone","orcid":"0000-0002-4644-6230","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"A vision transformer approach for fully automated and scalable dementia screening using clock drawing test images","abstract":"<jats:title>Abstract</jats:title><jats:sec><jats:title>INTRODUCTION</jats:title><jats:p>The clock drawing test (CDT) screens for dementia but requires trained scorers and lacks standardized criteria. Thus, we developed an automated vision transformer (ViT)‐based diagnostic system with convolutional neural network preprocessing for analyzing hand‐drawn CDT images.</jats:p></jats:sec><jats:sec><jats:title>METHODS</jats:title><jats:p>The architecture implements fine‐tuned ViT feature extraction with linear classification for dementia prediction. Training used the National Health and Aging Trends Study (NHATS) dataset (<jats:italic>n</jats:italic> = 54,027), with testing on an independent clinical cohort from the Toronto Dementia Research Alliance (TDRA; <jats:italic>n</jats:italic> = 862; 522 dementia, 340 normal cognition).</jats:p></jats:sec><jats:sec><jats:title>RESULTS</jats:title><jats:p>The ViT approach predicted dementia with 76.5% balanced accuracy, outperforming human‐scored features (74.3%) and existing deep learning models (MiniVGG = 73.3%, MobileNetV2 = 72.3%, relevance factor variational autoencoder = 69.1%) on the TDRA dataset.</jats:p></jats:sec><jats:sec><jats:title>DISCUSSION</jats:title><jats:p>This pen‐and‐paper compatible diagnostic system enables scalable remote cognitive screening through automated CDT image analysis that is competitive with human‐scored features, potentially increasing diagnostic accessibility for diverse populations across varied socioeconomic contexts.</jats:p></jats:sec><jats:sec><jats:title>HIGHLIGHTS</jats:title><jats:p><jats:list list-type=\"bullet\">\n<jats:list-item><jats:p>The vision transformer model achieves 76.5% accuracy in dementia detection from clock drawing tests, outperforming human scoring and existing deep learning methods.</jats:p></jats:list-item>\n<jats:list-item><jats:p>Novel convolutional neural network–based preprocessing automatically handles challenging image quality issues like shadows, irrelevant markings, and improper cropping.</jats:p></jats:list-item>\n<jats:list-item><jats:p>The system requires only a photo of a hand‐drawn clock test, enabling scalable remote screening accessible across socioeconomic contexts.</jats:p></jats:list-item>\n<jats:list-item><jats:p>A feature‐extraction model trained on 54,027 samples demonstrates robust generalization to an independent clinical dataset of 862 patients.</jats:p></jats:list-item>\n<jats:list-item><jats:p>This fully automated approach eliminates the need for trained scorers while maintaining diagnostic accuracy above manual methods.</jats:p></jats:list-item>\n</jats:list></jats:p></jats:sec>","is_dataset_classified":null,"base_score":1.3862943611198906,"endowment":1.3862943611198906,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"41001422","pmcid":"PMC12457074","openalex_id":"https://openalex.org/W4414564277","authors":[],"funders":[{"funder_name":"National Institutes of Health","grant_id":"5U01AG032947-05","title":"National Study of Disability Trends and Dynamics"}],"total_grants":1,"fwci":2.4373,"citation_percentile":0.89895947,"influential_citations":0,"citation_trend":[{"year":2026,"count":3}],"oa_status":"gold","license":"cc-by-nc","oa_locations":[{"url":"https://onlinelibrary.wiley.com/doi/pdfdirect/10.1002/dad2.70171","host_type":"journal"},{"url":"https://onlinelibrary.wiley.com/doi/pdfdirect/10.1002/dad2.70171","host_type":"publisher"},{"url":"https://alz-journals.onlinelibrary.wiley.com/doi/pdf/10.1002/dad2.70171","host_type":"publisher"},{"url":"https://doi.org/10.1002/dad2.70171","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/41001422","host_type":"repository"},{"url":"https://doaj.org/article/200e34195d834d818ddc324fb276fc84","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/12457074","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12457074/","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC12457074","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC12457074?pdf=render","host_type":"Europe_PMC"},{"url":"https://doi.org/10.1002/alz70858_097670","host_type":""},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12739198/","host_type":""},{"url":"https://pubmed.ncbi.nlm.nih.gov/41001422/","host_type":""}],"fields_of_study":["Dementia and Cognitive Impairment Research","Pressure Ulcer Prevention and Management","EEG and Brain-Computer Interfaces","0202 electrical engineering, electronic engineering, information engineering","02 engineering and technology"],"mesh_terms":[],"keywords":["Preprocessor","Convolutional neural network","Scalability","Transformer","Deep learning","Artificial neural network","Data pre-processing","Cognitive Assessment","Clock Drawing Test","Dementia Screening","Vision Transformer","Dementia Care Research and Psychosocial Factors","Research Article"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-02T01:33:51.473901Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}