{"doi":"10.3390/covid5120205","title":"Developing a Long COVID Case Definition: Using Machine Learning to Distinguish Long COVID Based on Symptom Presentation","abstract":"Efforts have been made to develop a case definition for Long COVID, with results differing on whether the case definition should be specific and exclusive, or broad and easily generalizable. Each of these methods has been subject to limitations. As most efforts have focused on symptoms, inclusion criteria have often relied on the binary occurrence of a symptom. The current study uses a more detailed measure that considers the frequency and severity of symptoms in a sample of individuals with Long COVID and matched controls who recovered from acute SARS-CoV-2 infection. Patients were diagnosed with Long COVID in a systematic process involving their completion of quantitative questionnaires, qualitative interviews, a physical examination, and general laboratory testing to rule out other diagnoses. Since samples were comparatively small given the number of symptoms investigated, Leave One Out Cross-Validation (LOOCV) was used to develop LASSO regression models to determine which symptoms best distinguished Long COVID from recovered controls. An ideal threshold for classifying Long COVID based on symptomatology was developed using a receiver operator characteristics (ROC) curve. The model presented in this article identified Long COVID with high accuracy. The importance of smell/taste was lessened in the current study, and gastrointestinal symptoms took on greater prominence in our study. It is possible to achieve high accuracy in differentiating those with Long COVID from those who have recovered. It is important to specify criteria of Long COVID and to measure symptoms comprehensively to identify those with Long COVID. Reliably identifying those who have developed Long COVID will help in the formulation of treatment strategies.","journal":"COVID","year":2025,"id":537540,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":1,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9544,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":902835,"name":"Jacob Furst","orcid":"0009-0006-9019-2182","position":1,"is_corresponding":false},{"id":1423778,"name":"Lauren Ruesink","orcid":null,"position":2,"is_corresponding":false},{"id":432163,"name":"Ben Z. Katz","orcid":"0000-0003-0149-2388","position":3,"is_corresponding":false},{"id":530664,"name":"Leonard A. Jason","orcid":"0000-0002-9972-4425","position":0,"is_corresponding":true}],"reference_count":0,"raw_metadata":null,"created_at":"2026-07-19T02:52:12.997494Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}