{"doi":"10.1111/cts.13299","title":"Feeling anxious yet? Interpreting findings on drug safety from large healthcare databases","abstract":"In this issue of Clinical and Translational Science, Yun-Han Wang and colleagues report on a potential link between treatment with proton pump inhibitors (PPIs) and incident depression and anxiety disorders in children.1 The proposed mechanism between PPI use and anxiety and depression is related to the “microbiota-gut-brain axis.” The microbiota-gut-brain axis and associated health problems have been attributed to multiple inter-related pathways, including microbiota-derived production of neurotransmitters (e.g., serotonin); communications among afferent, efferent, autonomic, and limbic nerves; immune-mediated pathways (e.g., cytokines); and endocrine pathways (e.g., hypothalamic–pituitary–adrenal axis).2 PPIs are known to dysregulate the microbiome, and microbiome dysregulation has been linked to mental health disorders in animal models and in adult populations.2, 3 To investigate this pediatric drug safety question, Wang and colleagues utilized nationwide databases that cover medication dispensing and hospital and emergency department encounters in Sweden. The authors observed an increase in depression and anxiety in children who initiated PPI treatment, a finding consistent across multiple secondary and sensitivity analyses. Pharmacoepidemiological studies, such as the one by Wang et al.4 can provide valuable insights on pediatric drug safety, particularly when randomized controlled trials (RCTs) are unavailable, underpowered, or infeasible to address outcomes of interest. Understanding the strengths and weaknesses of real-world data from large healthcare databases for pharmacoepidemiological research and associated threats to validity are key to interpreting findings. Pharmacoepidemiological studies investigating drug safety often utilize large healthcare databases, similar to the data utilized by Wang et al. These types of data include national administrative data (sometimes termed “registries,” including the Swedish registry data used by Wang et al.), electronic health record (EHR) data, and insurance claims. Although aspects of each data source vary, many share critical features that permit pharmacoepidemiologic research. A key component to these data sources is the availability of detailed information on prescribed or pharmacy-dispensed medications and dates of prescribing or dispensing, with additional details such as days’ supply and dosage. Details on medications in the context of longitudinal data allow assessment of treatment patterns and identification of persons initiating treatment. The implementation of the “new user design” is a key feature of pharmacoepidemiologic studies that minimizes bias by restricting the study population to individuals who, analogous to RCTs, newly initiate the treatments of interest.5 Further, these data often capture diagnoses and procedures from healthcare encounters in inpatient, emergency, or outpatient settings with associated dates. This patient-level clinical information allows implementation of inclusion and exclusion criteria, characterization of the study population, adjustment for confounders, and identification of outcomes. Additionally, the large size of these datasets allows for assessment of rare outcomes, rare therapeutics, and restricted or vulnerable study populations lacking robust data from RCTs (e.g., children and pregnant women). Variations and nuances across these databases are important to consider when selecting a data source for a particular research question and when interpretating results.6 For example, certain databases may be restricted to primary care encounters, may capture only written prescriptions or only dispensed medications, and may lack data on medications received in inpatient settings. Although large healthcare databases have key strengths for pediatric pharmacoepidemiological safety studies, threats to validity remain. Pharmacoepidemiologists attempt to address these concerns through study design and analysis. Three key topics in interpreting","journal":"Clinical and Translational Science","year":2022,"id":294502,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.95,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2022-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":247077,"name":"Daniel B. Horton","orcid":"0000-0002-1831-1339","position":1,"is_corresponding":false},{"id":504255,"name":"Tobias Gerhard","orcid":"0000-0002-8598-5771","position":2,"is_corresponding":false},{"id":322703,"name":"Greta Bushnell","orcid":"0000-0002-4056-8405","position":0,"is_corresponding":true}],"reference_count":11,"raw_metadata":null,"created_at":"2026-07-19T00:31:01.450041Z","pmid":"35578775","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}