{"doi":"10.2196/preprints.84791","title":"Cross-Cultural Differences in Public Discourse on COVID-19 Vaccination in the United States and South Korea: Cross-Sectional Analysis Using Natural Language Processing (Preprint)","abstract":"<sec>\n                  <title>BACKGROUND</title>\n                  <p>The COVID-19 vaccine was introduced as a crucial tool to combat the pandemic. However, concerns about its effectiveness, side effects, and misinformation spread remain. Prior research largely relied on survey-based approaches with limited populations. To address these limitations, social media offers a broader, more naturalistic lens into public discourse on COVID-19 vaccination. Accordingly, our study leverages social media data to identify factors shaping vaccine-related information needs, perceptions, and communication dynamics.</p>\n                </sec>\n                <sec>\n                  <title>OBJECTIVE</title>\n                  <p>This study investigated public discourse about COVID-19 vaccines on community-driven question-and-answer sites in the United States (Quora; Quora, Inc) and South Korea (Naver Knowledge-iN; Naver Corp) to identify cross-national similarities and differences in vaccine-related information needs, sentiment patterns, and public perceptions over time.</p>\n                </sec>\n                <sec>\n                  <title>METHODS</title>\n                  <p>We analyzed publicly available COVID-19 vaccine–related questions and answers posted between June 27, 2020, and June 27, 2021, on 2 community-driven question-and-answer platforms: Quora (United States) and Naver Knowledge-iN (South Korea). After preprocessing and sample-size matching, the dataset included 3952 question-answer pairs per platform, with one community-selected (most upvoted) answer analyzed per question. Natural language processing (NLP) techniques were applied for topic classification and sentiment analysis. Questions were categorized using a hybrid topic modeling approach combining Latent Dirichlet Allocation (LDA) and Top2Vec, identifying 5 topics on Quora and 7 topics on Naver Knowledge-iN. Answer sentiments were classified using an ensemble of Bidirectional Encoder Representations from Transformers (BERT; Google LLC)– and Efficiently Learning an Encoder that Classifies Token Replacements Accurately (ELECTRA; Google LLC)–based transformer models, and temporal sentiment trends were examined using monthly aggregation.</p>\n                </sec>\n                <sec>\n                  <title>RESULTS</title>\n                  <p>Five shared information needs emerged, including effects of vaccines, variants, government policy, visiting overseas, and different vaccines, while South Korea uniquely exhibited vaccination appointments (711/3952, 18%) and school and education (513/3592, 13%). Negative sentiment predominated in US (Quora) answers across 4 of 5 topics, whereas positive sentiment exceeded 50% (498/790, 337/474, 367/592, 218/316, 348/553, 562/711, and 364/513) across all 7 topics on Naver Knowledge-iN. Temporally, US sentiment exhibited multiple positive-negative crossovers, whereas Korean sentiment stabilized toward positivity after February 2021, coinciding with the national vaccine rollout. Question-answer sentiment pairs showed contrasting interaction patterns, including negative-negative pairs dominated in the United States (eg, 504/978, 51.5% for different vaccines), while in South Korea, positive-positive and negative-positive pairs accounted for more than 63% (498/790, 337/474, 367/592, 218/316, 348/553, 562/711, and 364/513) of interactions in 7 topics, with positive-positive pairs most prevalent in 6 of 7 topics, except for variants.</p>\n                </sec>\n                <sec>\n                  <title>CONCLUSIONS</title>\n                  <p>Public perceptions of COVID-19 vaccines and related information needs differ between the 2 countries, shaped by cultural context, trust in government, and information-seeking environments. Analysis of social question and answer data from the 2 countries reveals shared information needs but divergent sentiment patterns. These findings highlight the value of social media data for public health research and the need for culturally and platform-specific communication strategies.</p>\n                </sec>","journal":null,"year":null,"id":625080,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":342335,"name":"Sou Hyun Jang","orcid":"0000-0003-2265-4753","position":1,"is_corresponding":false},{"id":68523,"name":"Haewoon Kwak","orcid":"0000-0003-1418-0834","position":2,"is_corresponding":false},{"id":1395006,"name":"Jaeyoung Choi","orcid":"0000-0001-8264-1482","position":3,"is_corresponding":false},{"id":1616190,"name":"Yong Jeong Yi","orcid":"0000-0003-1736-2107","position":4,"is_corresponding":false},{"id":1616186,"name":"Sangpil Youm","orcid":"0000-0001-7234-0395","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Cross-Cultural Differences in Public Discourse on COVID-19 Vaccination in the United States and South Korea: Cross-Sectional Analysis Using Natural Language Processing (Preprint)","abstract":"<sec>\n                  <title>BACKGROUND</title>\n                  <p>The COVID-19 vaccine was introduced as a crucial tool to combat the pandemic. However, concerns about its effectiveness, side effects, and misinformation spread remain. Prior research largely relied on survey-based approaches with limited populations. To address these limitations, social media offers a broader, more naturalistic lens into public discourse on COVID-19 vaccination. Accordingly, our study leverages social media data to identify factors shaping vaccine-related information needs, perceptions, and communication dynamics.</p>\n                </sec>\n                <sec>\n                  <title>OBJECTIVE</title>\n                  <p>This study investigated public discourse about COVID-19 vaccines on community-driven question-and-answer sites in the United States (Quora; Quora, Inc) and South Korea (Naver Knowledge-iN; Naver Corp) to identify cross-national similarities and differences in vaccine-related information needs, sentiment patterns, and public perceptions over time.</p>\n                </sec>\n                <sec>\n                  <title>METHODS</title>\n                  <p>We analyzed publicly available COVID-19 vaccine–related questions and answers posted between June 27, 2020, and June 27, 2021, on 2 community-driven question-and-answer platforms: Quora (United States) and Naver Knowledge-iN (South Korea). After preprocessing and sample-size matching, the dataset included 3952 question-answer pairs per platform, with one community-selected (most upvoted) answer analyzed per question. Natural language processing (NLP) techniques were applied for topic classification and sentiment analysis. Questions were categorized using a hybrid topic modeling approach combining Latent Dirichlet Allocation (LDA) and Top2Vec, identifying 5 topics on Quora and 7 topics on Naver Knowledge-iN. Answer sentiments were classified using an ensemble of Bidirectional Encoder Representations from Transformers (BERT; Google LLC)– and Efficiently Learning an Encoder that Classifies Token Replacements Accurately (ELECTRA; Google LLC)–based transformer models, and temporal sentiment trends were examined using monthly aggregation.</p>\n                </sec>\n                <sec>\n                  <title>RESULTS</title>\n                  <p>Five shared information needs emerged, including effects of vaccines, variants, government policy, visiting overseas, and different vaccines, while South Korea uniquely exhibited vaccination appointments (711/3952, 18%) and school and education (513/3592, 13%). Negative sentiment predominated in US (Quora) answers across 4 of 5 topics, whereas positive sentiment exceeded 50% (498/790, 337/474, 367/592, 218/316, 348/553, 562/711, and 364/513) across all 7 topics on Naver Knowledge-iN. Temporally, US sentiment exhibited multiple positive-negative crossovers, whereas Korean sentiment stabilized toward positivity after February 2021, coinciding with the national vaccine rollout. Question-answer sentiment pairs showed contrasting interaction patterns, including negative-negative pairs dominated in the United States (eg, 504/978, 51.5% for different vaccines), while in South Korea, positive-positive and negative-positive pairs accounted for more than 63% (498/790, 337/474, 367/592, 218/316, 348/553, 562/711, and 364/513) of interactions in 7 topics, with positive-positive pairs most prevalent in 6 of 7 topics, except for variants.</p>\n                </sec>\n                <sec>\n                  <title>CONCLUSIONS</title>\n                  <p>Public perceptions of COVID-19 vaccines and related information needs differ between the 2 countries, shaped by cultural context, trust in government, and information-seeking environments. Analysis of social question and answer data from the 2 countries reveals shared information needs but divergent sentiment patterns. These findings highlight the value of social media data for public health research and the need for culturally and platform-specific communication strategies.</p>\n                </sec>","is_dataset_classified":null,"base_score":0.0,"endowment":0.0,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19767382","pmcid":null,"openalex_id":"https://openalex.org/W7133697518","authors":[],"funders":[],"total_grants":0,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://doi.org/10.2196/preprints.84791","host_type":""},{"url":"https://doi.org/10.2196/preprints.84791","host_type":""}],"fields_of_study":["Computational and Text Analysis Methods","Misinformation and Its Impacts","Vaccine Coverage and Hesitancy"],"mesh_terms":[],"keywords":["Social media","Misinformation","Topic model","Latent Dirichlet allocation","Public opinion","Encoder","Sentiment analysis","Perception"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Good health and well-being"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-04T05:44:37.902161Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}