{"doi":"10.1016/j.jad.2025.04.014","title":"Estimating depression severity in narrative clinical notes using large language models","abstract":"Background Depression treatment guidelines emphasize measurement-based care using patient-reported outcome measures, yet their impact on narrative documentation quality remains underexplored. Methods We sampled 15,000 narrative clinical outpatient notes from the electronic health record of a large academic medical center, reflecting visits between January 2, 2019 and January 30, 2024, for which a 9-item Patient Health Questionnaire (PHQ-9) was completed at the same time. After censoring PHQ-9 scores from notes, we estimated severity of depressive symptoms with a foundational large language model (gpt4o-08-06) in a HIPAA-compliant enclave. We estimated correlation between true PHQ-9 and model-estimated score and examined the predictive performance of the model for moderate or greater depressive symptoms. Results Mean age was 46.3 years (SD 14.9); 9083 (60.6 %) identified as female. 925 (6.2 %) identified as Asian, 638 (4.3 %) as Black, 853 (5.7 %) as another race, and 12,187 (81.2 %) as White. A total of 1044 (7.0 %) identified as Hispanic ethnicity, while 12,699 (84.7 %) were non-Hispanic. Mean measured PHQ-9 score was 1.23 (SD 3.45); 721 (4.8 %) met criteria for moderate or greater depressive symptoms. LLM-predicted PHQ-9 scores were modestly correlated with actual scores (r 2 = 0.264 (95 % CI 0.252–0.276)); PPV for moderate or greater depression was 0.309 (95 % CI 0.302–0.317). Performance was consistent across demographic subgroups, with modest differences identified by race, ethnicity, and sex. Conclusion A foundational LLM performed poorly but consistently across subgroups in imputing PHQ-9 scores from notes when actual PHQ-9 reporting was ablated. This result suggests the extent to which inclusion of PROMs may impoverish documentation of psychiatric symptoms.","journal":"Journal of Affective Disorders","year":2025,"id":516572,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":9,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.7311,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":274491,"name":"Víctor M. Castro","orcid":"0000-0001-7390-6354","position":1,"is_corresponding":false},{"id":1382345,"name":"Roy H Perlis","orcid":null,"position":2,"is_corresponding":false},{"id":808901,"name":"Thomas H. McCoy","orcid":"0000-0002-5624-0439","position":0,"is_corresponding":true}],"reference_count":29,"raw_metadata":null,"created_at":"2026-07-19T02:48:49.486328Z","pmid":"40187432","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}