{"doi":"10.1093/geroni/igaf120","title":"Advanced topic modeling with large language models: analyzing social media content from dementia caregivers","abstract":"Background and Objectives: While traditional topic modeling methods have been applied to analyze social media content from dementia caregivers, they often struggle with semantic understanding and coherent topic generation. This study explores the direct application of large language models (LLMs) for topic modeling of caregiver tweets, aiming to leverage their advanced semantic comprehension capabilities. Research Design and Methods: We analyzed 231 870 tweets from dementia caregivers after preprocessing using ChatGPT as the primary topic modeling tool. To address context length limitations, we developed a 2-stage approach: first splitting the dataset into 226 batches of 1000 tweets each for initial topic extraction, then combining these results through a second-stage prompt for final topic synthesis. We compared our approach against 11 baseline methods, including Latent Dirichlet Allocation (LDA), Gibbs Sampling Dirichlet Multinomial Mixture Model (GSDMM), their term-weighted variants, and state-of-the-art BERTopic models. Topic quality was evaluated using Sentence-BERT-based coherence scores, and topic comprehensiveness was assessed through both ChatGPT and human expert evaluation. Results: Our LLM-based approach achieved a coherence score of 0.358, significantly outperforming all baseline methods. Traditional approaches like GSDMM (0.317) and LDA (0.320), their term-weighted variants (ranging from 0.264 to 0.302), and BERTopic variants (approximately 0.30) showed lower coherence scores. The 2-stage batching strategy effectively handled the large dataset while maintaining topic quality and representativeness. Expert evaluation confirmed the topics' relevance to caregiver experiences and their comprehensive coverage of key themes. Discussion and Implications: This study introduces a novel methodology for applying LLMs to large-scale topic modeling tasks, demonstrating superior performance over traditional and state-of-the-art approaches. The significant improvement in coherence scores suggests that LLMs can better capture the semantic relationships within topics. Our approach addresses key challenges in context length limitations and prompt engineering, while providing more coherent and interpretable insights into caregiver experiences that can inform targeted support strategies.","journal":"Innovation in Aging","year":2025,"id":527916,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9556,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":869622,"name":"Bojian Hou","orcid":"0000-0002-3894-4547","position":1,"is_corresponding":false},{"id":226456,"name":"Amy Zheng","orcid":"0009-0004-4051-2796","position":2,"is_corresponding":false},{"id":573148,"name":"Yanbo Feng","orcid":null,"position":3,"is_corresponding":false},{"id":481140,"name":"Ari Z Klein","orcid":"0000-0002-8281-3464","position":4,"is_corresponding":false},{"id":96346,"name":"Karen O’Connor","orcid":"0000-0001-7709-3813","position":5,"is_corresponding":false},{"id":1013077,"name":"Shu Yang","orcid":"0000-0002-8507-7191","position":6,"is_corresponding":false},{"id":1358078,"name":"Tianqi Shang","orcid":null,"position":7,"is_corresponding":false},{"id":412309,"name":"George Demiris","orcid":"0000-0002-6318-5829","position":8,"is_corresponding":false},{"id":76217,"name":"Graciela Gonzalez-Hernandez","orcid":null,"position":9,"is_corresponding":false},{"id":49224,"name":"Li Shen","orcid":"0000-0002-5443-0503","position":10,"is_corresponding":false},{"id":1357741,"name":"Weiqing He","orcid":"0009-0006-5219-6289","position":0,"is_corresponding":true}],"reference_count":23,"raw_metadata":null,"created_at":"2026-07-19T02:50:44.062153Z","pmid":"41458886","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}