{"doi":"10.17615/mw2x-r483","title":"Evaluating Picture Description Speech for Dementia Detection using Image-text Alignment","abstract":"Using picture description speech for dementia detection has been extensively studied. However, past research has concentrated on distinguishing speech patterns between healthy individuals and dementia patients without directly utilizing the picture content. This paper introduces the first dementia detection models using images and description texts, leveraging knowledge from pre-trained image-text alignment models. It explores the distinction between dementia and healthy samples based on the text’s relevance to the image and its focused area, suggesting this approach could improve dementia detection accuracy. Specifically, we use the text’s relevance to the picture to rank and filter the sentences of the samples. We also identified focused areas of the picture as topics and categorized the sentences according to the focused areas. We propose three advanced models that pre-processed the samples based on their relevance to the picture, sub-image, and focused areas. The evaluation results show that our advanced models, with knowledge of the picture and large image-text alignment models, achieve state-of-the-art performance with the best detection accuracy at 83.44%, which is higher than the text-only baseline model at 79.91%. Lastly, we visualize the sample and picture results to explain the advantages of our models.","journal":"UNC Libraries","year":2025,"id":583948,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9589,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":658315,"name":"Youxiang Zhu","orcid":"0009-0000-7294-1596","position":1,"is_corresponding":false},{"id":272507,"name":"John A. Batsis","orcid":"0000-0002-0845-4416","position":2,"is_corresponding":false},{"id":634806,"name":"Brian MacWhinney","orcid":"0000-0002-4988-1342","position":3,"is_corresponding":false},{"id":658316,"name":"Xiaohui Liang","orcid":"0000-0003-4064-2393","position":4,"is_corresponding":false},{"id":315571,"name":"Robert M. Roth","orcid":"0000-0003-4374-6569","position":5,"is_corresponding":false},{"id":1074491,"name":"Nana Lin","orcid":"0009-0009-1714-6405","position":0,"is_corresponding":true}],"reference_count":0,"raw_metadata":null,"created_at":"2026-07-19T02:59:07.970345Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}