{"doi":"10.1093/jamia/ocae010","title":"Overview of the 8th Social Media Mining for Health Applications (#SMM4H) shared tasks at the AMIA 2023 Annual Symposium","abstract":"OBJECTIVE: The aim of the Social Media Mining for Health Applications (#SMM4H) shared tasks is to take a community-driven approach to address the natural language processing and machine learning challenges inherent to utilizing social media data for health informatics. In this paper, we present the annotated corpora, a technical summary of participants' systems, and the performance results. METHODS: The eighth iteration of the #SMM4H shared tasks was hosted at the AMIA 2023 Annual Symposium and consisted of 5 tasks that represented various social media platforms (Twitter and Reddit), languages (English and Spanish), methods (binary classification, multi-class classification, extraction, and normalization), and topics (COVID-19, therapies, social anxiety disorder, and adverse drug events). RESULTS: In total, 29 teams registered, representing 17 countries. In general, the top-performing systems used deep neural network architectures based on pre-trained transformer models. In particular, the top-performing systems for the classification tasks were based on single models that were pre-trained on social media corpora. CONCLUSION: To facilitate future work, the datasets-a total of 61 353 posts-will remain available by request, and the CodaLab sites will remain active for a post-evaluation phase.","journal":"Journal of the American Medical Informatics Association","year":2024,"id":452505,"datarank":0.9433871245783999,"base_score":2.833213344056216,"endowment":2.833213344056216,"self_citation_contribution":0.42498200160843247,"citation_network_contribution":0.5184051229699674,"self_endowment_contribution":0.42498200160843247,"citer_contribution":0.5184051229699674,"corpus_percentile":78.58745261855032,"corpus_rank":2769,"citation_count":16,"citer_count":12,"citers_with_citation_signal":11,"citers_with_endowment":11,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.7821,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2024-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":4963,"name":"Juan M. Banda","orcid":"0000-0001-8499-824X","position":1,"is_corresponding":false},{"id":273824,"name":"Yuting Guo","orcid":"0000-0002-8919-0888","position":2,"is_corresponding":false},{"id":1163426,"name":"Ana Lucía Schmidt","orcid":"0000-0002-3352-5677","position":3,"is_corresponding":false},{"id":488870,"name":"Dongfang Xu","orcid":"0000-0003-0828-1102","position":4,"is_corresponding":false},{"id":1274724,"name":"Ivan Amaro","orcid":"0000-0002-8813-5510","position":5,"is_corresponding":false},{"id":1163427,"name":"Raul Rodriguez‐Esteban","orcid":"0000-0002-9494-9609","position":6,"is_corresponding":false},{"id":96345,"name":"Abeed Sarker","orcid":"0000-0001-7358-544X","position":7,"is_corresponding":false},{"id":21659,"name":"Graciela Gonzalez‐Hernandez","orcid":"0000-0002-6416-9556","position":8,"is_corresponding":false},{"id":481140,"name":"Ari Z Klein","orcid":"0000-0002-8281-3464","position":0,"is_corresponding":true}],"reference_count":13,"raw_metadata":null,"created_at":"2026-07-19T02:02:54.415477Z","pmid":"38218723","pmcid":"PMC10990511","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}