{"doi":"10.2196/25457","title":"Classification of the Disposition of Patients Hospitalized with COVID-19: Reading Discharge Summaries Using Natural Language Processing","abstract":"BACKGROUND: Medical notes are a rich source of patient data; however, the nature of unstructured text has largely precluded the use of these data for large retrospective analyses. Transforming clinical text into structured data can enable large-scale research studies with electronic health records (EHR) data. Natural language processing (NLP) can be used for text information retrieval, reducing the need for labor-intensive chart review. Here we present an application of NLP to large-scale analysis of medical records at 2 large hospitals for patients hospitalized with COVID-19. OBJECTIVE: Our study goal was to develop an NLP pipeline to classify the discharge disposition (home, inpatient rehabilitation, skilled nursing inpatient facility [SNIF], and death) of patients hospitalized with COVID-19 based on hospital discharge summary notes. METHODS: Text mining and feature engineering were applied to unstructured text from hospital discharge summaries. The study included patients with COVID-19 discharged from 2 hospitals in the Boston, Massachusetts area (Massachusetts General Hospital and Brigham and Women's Hospital) between March 10, 2020, and June 30, 2020. The data were divided into a training set (70%) and hold-out test set (30%). Discharge summaries were represented as bags-of-words consisting of single words (unigrams), bigrams, and trigrams. The number of features was reduced during training by excluding n-grams that occurred in fewer than 10% of discharge summaries, and further reduced using least absolute shrinkage and selection operator (LASSO) regularization while training a multiclass logistic regression model. Model performance was evaluated using the hold-out test set. RESULTS: The study cohort included 1737 adult patients (median age 61 [SD 18] years; 55% men; 45% White and 16% Black; 14% nonsurvivors and 61% discharged home). The model selected 179 from a vocabulary of 1056 engineered features, consisting of combinations of unigrams, bigrams, and trigrams. The top features contributing most to the classification by the model (for each outcome) were the following: \"appointments specialty,\" \"home health,\" and \"home care\" (home); \"intubate\" and \"ARDS\" (inpatient rehabilitation); \"service\" (SNIF); \"brief assessment\" and \"covid\" (death). The model achieved a micro-average area under the receiver operating characteristic curve value of 0.98 (95% CI 0.97-0.98) and average precision of 0.81 (95% CI 0.75-0.84) in the testing set for prediction of discharge disposition. CONCLUSIONS: A supervised learning-based NLP approach is able to classify the discharge disposition of patients hospitalized with COVID-19. This approach has the potential to accelerate and increase the scale of research on patients' discharge disposition that is possible with EHR data.","journal":"JMIR Medical Informatics","year":2020,"id":87889,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":11,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9603,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2020-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":305327,"name":"Haoqi Sun","orcid":"0000-0002-5041-8312","position":1,"is_corresponding":false},{"id":336247,"name":"Aayushee Jain","orcid":"0000-0002-5018-3234","position":2,"is_corresponding":false},{"id":336249,"name":"Haitham Alabsi","orcid":"0000-0001-6354-4679","position":3,"is_corresponding":false},{"id":249909,"name":"Laura Brenner","orcid":"0000-0001-6752-0471","position":4,"is_corresponding":false},{"id":336250,"name":"Elissa Ye","orcid":"0000-0003-4851-6543","position":5,"is_corresponding":false},{"id":336251,"name":"Wendong Ge","orcid":"0000-0003-1557-5336","position":6,"is_corresponding":false},{"id":336256,"name":"Sarah Isabel Collens","orcid":"0000-0001-7010-7266","position":7,"is_corresponding":false},{"id":336248,"name":"Michael J. Leone","orcid":"0000-0002-0218-8612","position":8,"is_corresponding":false},{"id":45250,"name":"Sudeshna Das","orcid":"0000-0002-9486-6811","position":9,"is_corresponding":false},{"id":311026,"name":"Gregory K. Robbins","orcid":"0000-0001-7545-5817","position":10,"is_corresponding":false},{"id":311025,"name":"Shibani S. Mukerji","orcid":"0000-0002-5677-6954","position":11,"is_corresponding":false},{"id":280809,"name":"M. Brandon Westover","orcid":"0000-0003-4803-312X","position":12,"is_corresponding":false},{"id":446095,"name":"Marta Fernandes","orcid":"0000-0002-7203-2832","position":0,"is_corresponding":true}],"reference_count":28,"raw_metadata":null,"created_at":"2026-07-18T22:00:29.831859Z","pmid":"33449908","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}