{"doi":"10.1093/jamiaopen/ooaf102","title":"Automating inductive thematic analyses of health content using large language models: a proof-of-concept study using social media data","abstract":"Abstract Objectives Large language models (LLMs) face challenges in inductive thematic analysis, a task requiring deep interpretive, domain-specific expertise. We evaluated the feasibility of using LLMs to replicate expert-driven thematic analysis of social media data. Materials and Methods Using 2 temporally nonintersecting Reddit datasets on xylazine (n = 286 and 686, for model optimization and validation, respectively) with 12 expert-derived themes, we evaluated 5 LLMs against expert coding. We modeled the task as a series of binary classifications, rather than a single, multilabel classification, employing zero-, single-, and few-shot prompting strategies and measuring performance via accuracy, precision, recall, and F1 score. Results On the validation set, GPT-4o with 2-shot prompting performed best (accuracy: 90.9%; F1 score: 0.71). For high-prevalence themes, model-derived thematic distributions closely mirrored expert classifications (eg, xylazine: 13.6% vs 17.8%; medications for opioid use disorders: 16.5% vs 17.8%). Conclusion Our findings suggest that few-shot LLM-based approaches can automate thematic analyses, offering a scalable supplement for qualitative research.","journal":"JAMIA Open","year":2025,"id":542345,"datarank":0.16479184330021646,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"self_citation_contribution":0.16479184330021646,"citation_network_contribution":0.0,"self_endowment_contribution":0.16479184330021646,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.6045,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1431926,"name":"Ritvik Ranjan","orcid":null,"position":1,"is_corresponding":false},{"id":1030502,"name":"Sahithi Lakamana","orcid":"0000-0003-1304-7484","position":2,"is_corresponding":false},{"id":808229,"name":"Anthony Spadaro","orcid":"0000-0002-0941-4651","position":3,"is_corresponding":false},{"id":73053,"name":"Selen Bozkurt","orcid":"0000-0003-1234-2158","position":4,"is_corresponding":false},{"id":570754,"name":"Jeanmarie Perrone","orcid":"0000-0002-3396-6333","position":5,"is_corresponding":false},{"id":96345,"name":"Abeed Sarker","orcid":"0000-0001-7358-544X","position":6,"is_corresponding":false},{"id":1223660,"name":"JaMor Hairston","orcid":"0000-0001-6069-5869","position":0,"is_corresponding":true}],"reference_count":20,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-19T02:52:55.809501Z","pmid":"40985037","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}