{"doi":"10.31234/osf.io/uvamz","title":"Key language markers of depression on social media depend on race","abstract":"Depression has robust natural language correlates, and can increasingly be measured in language using predictive models. However, despite evidence that language use varies as a function of individual demographic features (e.g., age, gender), previous work has not systematically examined whether and how depression’s association with language varies by race. Here, we examined how race moderates the relationship between language features (i.e., first-person pronouns and negative emotions) and self-reported depression, in a matched sample of Black and White participants. Analyses revealed moderating effects of race: while depression severity predicts I-usage in White individuals, it does not in Black individuals. Negative emotion language use varies by race; as depression severity increased, White individuals used more belongingness and self-deprecation language, while Black individuals used more loneliness and worry language. Machine learning models trained to predict depression severity performed poorly when tested on Black individuals, even when they were trained exclusively using the language of Black individuals. In contrast, analogous models tested on White individuals performed relatively well. We ensured an equal number of Black and White individuals with similar age and gender distribution in our sample, thereby confirming that the performance disparity is not due to sampling bias. Our study reveals surprising race-based differences in the expression of psychological traits, like depression, in natural language, and highlights the need to better understand these effects, especially before language-based models for detecting psychological phenomena are integrated into clinical practice.","journal":"PsyArXiv (OSF Preprints)","year":2023,"id":397655,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9596,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1141996,"name":"Elizabeth Cameron Stade","orcid":"0000-0001-6409-848X","position":1,"is_corresponding":false},{"id":509750,"name":"Salvatore Giorgi","orcid":"0000-0001-7381-6295","position":2,"is_corresponding":false},{"id":1173484,"name":"Ashley Francisco","orcid":"0000-0002-4912-3787","position":3,"is_corresponding":false},{"id":15862,"name":"Lyle Ungar","orcid":"0000-0003-2047-1443","position":4,"is_corresponding":false},{"id":311274,"name":"Brenda Curtis","orcid":"0000-0002-2511-3322","position":5,"is_corresponding":false},{"id":710201,"name":"Sharath Chandra Guntuku","orcid":"0000-0002-2929-0035","position":6,"is_corresponding":false},{"id":1159947,"name":"Sunny Rai","orcid":"0000-0002-0677-3747","position":0,"is_corresponding":true}],"reference_count":29,"raw_metadata":null,"created_at":"2026-07-19T01:19:39.497368Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}