{"doi":"10.1093/jamiaopen/ooaf089","title":"A natural language processing pipeline for identifying pediatric long COVID symptoms and functional impacts in freeform clinical notes: a RECOVER study","abstract":"Objective: To develop a natural language processing (NLP) pipeline for unstructured electronic health record (EHR) data to identify symptoms and functional impacts associated with Long COVID in children. Materials and Methods: We analyzed 48 287 outpatient progress notes from 10 618 pediatric patients from 12 institutions. We evaluated notes obtained 28 to 179 days after a COVID-19 diagnosis or positive test. Two samples were examined: patients with evidence of Long COVID and patients with acute COVID but no evidence of Long COVID based on diagnostic codes. The pipeline identified clinical concepts associated with 21 symptoms and 4 functional impact categories. Subject matter experts (SMEs) screened a sample of 4586 terms from the NLP output to assess pipeline accuracy. Prevalence and concordance of each of the 25 concepts was compared between the 2 patient samples. Results: A binary assertion measure comparing SME and NLP assertions showed moderate accuracy (N = 4133; F1 = .80) and improved substantially when only high-confidence SME assertions were considered (N = 2043; F1 = .90). Overall, the 25 Long COVID concept categories were markedly more prevalent in the presumptive Long COVID cohort, and differences were noted between concepts identified in notes versus structured data. Discussion: This preliminary analysis illustrates the additional insight into a syndrome such as Long COVID gained from incorporating notes data, characterizing symptoms and functional impacts. Conclusion: These data support the importance of incorporating NLP methodology when possible into designing computable phenotypes and to accurately characterize patients with Long COVID.","journal":"JAMIA Open","year":2025,"id":534545,"datarank":0.16479184330021646,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"self_citation_contribution":0.16479184330021646,"citation_network_contribution":0.0,"self_endowment_contribution":0.16479184330021646,"citer_contribution":0.0,"corpus_percentile":29.844511487584125,"corpus_rank":8690,"citation_count":2,"citer_count":1,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.8013,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":12.5,"fair_percentile":28.767960868236013,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1353795,"name":"Cara Reedy","orcid":null,"position":1,"is_corresponding":false},{"id":984528,"name":"Vitaly Lorman","orcid":"0000-0003-0561-6587","position":2,"is_corresponding":false},{"id":367009,"name":"Ravi Jhaveri","orcid":"0000-0001-9921-0419","position":3,"is_corresponding":false},{"id":785879,"name":"Andrea Rivera-Sepúlveda","orcid":"0000-0002-8423-3981","position":4,"is_corresponding":false},{"id":1404038,"name":"Katherine S. Salamon","orcid":"0000-0003-0372-7670","position":5,"is_corresponding":false},{"id":958382,"name":"Payal Patel","orcid":"0000-0002-2984-9484","position":6,"is_corresponding":false},{"id":991667,"name":"Keith Morse","orcid":"0000-0002-8307-1642","position":7,"is_corresponding":false},{"id":1417133,"name":"Mattina Davenport","orcid":"0000-0002-5098-105X","position":8,"is_corresponding":false},{"id":555625,"name":"Lindsay G. Cowell","orcid":"0000-0003-1617-8244","position":9,"is_corresponding":false},{"id":991668,"name":"Levon Utidjian","orcid":"0000-0002-2189-108X","position":10,"is_corresponding":false},{"id":675064,"name":"Dimitri Christakis","orcid":"0000-0003-0726-7253","position":11,"is_corresponding":false},{"id":516763,"name":"Suchitra Rao","orcid":"0000-0002-0334-6301","position":12,"is_corresponding":false},{"id":516767,"name":"Marion R. Sills","orcid":"0000-0003-1322-0822","position":13,"is_corresponding":false},{"id":989174,"name":"Abigail Case","orcid":"0000-0002-5521-3076","position":14,"is_corresponding":false},{"id":111245,"name":"Eneida A. Mendonca","orcid":"0000-0003-4297-9221","position":15,"is_corresponding":false},{"id":271552,"name":"Bradley Taylor","orcid":"0000-0002-6414-4172","position":16,"is_corresponding":false},{"id":1417687,"name":"Jacqueline Rutter","orcid":null,"position":17,"is_corresponding":false},{"id":1417688,"name":"Aaron Thomas Martinez","orcid":null,"position":18,"is_corresponding":false},{"id":1201374,"name":"Rebecca Letts","orcid":"0009-0009-4747-7918","position":19,"is_corresponding":false},{"id":860340,"name":"L. Charles Bailey","orcid":"0000-0002-8967-0662","position":20,"is_corresponding":false},{"id":302148,"name":"Christopher B. Forrest","orcid":"0000-0003-1252-068X","position":21,"is_corresponding":false},{"id":305709,"name":"Iván Díaz","orcid":"0000-0001-9056-2047","position":22,"is_corresponding":false},{"id":1413011,"name":"Rachel Kenny","orcid":null,"position":23,"is_corresponding":false},{"id":342446,"name":"Jasmin Divers","orcid":"0000-0003-0120-7620","position":24,"is_corresponding":false},{"id":389936,"name":"Lorna E. Thorpe","orcid":"0000-0002-5535-2674","position":25,"is_corresponding":false},{"id":1038107,"name":"Hannah Mandel","orcid":"0000-0002-5685-9602","position":26,"is_corresponding":false},{"id":1289051,"name":"Jennifer Truong","orcid":"0000-0003-0255-0287","position":27,"is_corresponding":false},{"id":1381929,"name":"Shannon W. Wuller","orcid":"0000-0002-5763-3672","position":28,"is_corresponding":false},{"id":657634,"name":"Marc B. Rosenman","orcid":null,"position":30,"is_corresponding":false},{"id":451076,"name":"Sara J. Deakyne Davies","orcid":"0000-0002-8602-1624","position":32,"is_corresponding":false},{"id":892037,"name":"Christopher R. Forrest","orcid":"0000-0002-2934-9690","position":34,"is_corresponding":false},{"id":675067,"name":"Nathan M. Pajor","orcid":"0000-0002-1637-4942","position":35,"is_corresponding":false},{"id":1382386,"name":"Jyothi Priya Alekapatti Nandagopal","orcid":null,"position":36,"is_corresponding":false},{"id":809827,"name":"Alexander J. Stoddard","orcid":null,"position":38,"is_corresponding":false},{"id":460598,"name":"Kelly J. Kelleher","orcid":"0000-0002-4420-8044","position":39,"is_corresponding":false},{"id":536232,"name":"Yungui Huang","orcid":"0000-0003-3265-9902","position":40,"is_corresponding":false},{"id":1417689,"name":"Marion R Sills","orcid":null,"position":42,"is_corresponding":false},{"id":3872,"name":"Thuy Le","orcid":"0000-0002-3393-6580","position":43,"is_corresponding":false},{"id":992066,"name":"Daksha Ranade","orcid":null,"position":45,"is_corresponding":false},{"id":383225,"name":"Alan R. Schroeder","orcid":"0000-0002-9725-3279","position":46,"is_corresponding":false},{"id":437720,"name":"Mollie Cummins","orcid":"0000-0001-7078-8479","position":48,"is_corresponding":false},{"id":361482,"name":"Ramkiran Gouripeddi","orcid":"0000-0002-4345-9669","position":49,"is_corresponding":false},{"id":814080,"name":"Phillip Reeder","orcid":"0000-0002-8633-7872","position":51,"is_corresponding":false},{"id":805140,"name":"H. Timothy Bunnell","orcid":"0000-0002-3099-7572","position":0,"is_corresponding":true}],"reference_count":17,"raw_metadata":null,"created_at":"2026-07-19T02:51:47.434742Z","pmid":"40918941","pmcid":"PMC12409404","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":11.1111,"fair_a":6.25,"fair_i":60.0,"fair_r":41.6667,"fair_zscore":-0.8688,"fair_rationale":{"fair_score":12.5,"has_llm":true,"taxonomy_version":"fair_taxonomy_v5","dimensions":{"F":{"name":"Findable","score":11.11,"criteria":[{"key":"f_dataset_pid","label":"Persistent identifier for the data","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The paper does not provide any persistent identifier (DOI, Handle, ARK, or repository accession) for its own dataset.","anchors":["RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit","RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier'","FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'"],"scored":true,"signal":null},{"key":"f_repository_named","label":"Named repository","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The paper does not name a repository as the holder of the data; the data are held by the RECOVER program coordinating center with no repository name. [majority verdict 'no' (4/5 passes agreed)]","anchors":["RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed (","NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived","NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten"],"scored":true,"signal":null},{"key":"f_data_availability_statement","label":"Data-availability statement","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":"The results reported here are based on detailed individual-level patient data compiled as part of the RECOVER Program. Due to the high risk of reidentification based on the number of unique patterns in the data, particularly physician notes data, patient privacy regulations prohibit us from releasing the data publicly. The data are maintained in a secure enclave, with access managed by the program coordinating center to remain compliant with regulatory and program requirements. Please direct requests to access the data, either for reproduction of the work reported here or for other purposes, to recover@chop.edu.","grounded":false,"rationale":"The data availability statement points to a person (recover@chop.edu) for access, which is Colavizza category 1 (available on request). [downgraded to 'no' — no verifiable quote from the paper]","anchors":["Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li","Springer Nature research data policy — Data Availability Statements: standard statement templat","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes"],"scored":false,"signal":null},{"key":"f_discovery_metadata","label":"Description of the dataset as an object","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"We analyzed 48 287 outpatient progress notes from 10 618 pediatric patients from 12 institutions.","grounded":true,"rationale":"The dataset is described in running prose in the abstract, but there is no itemised inventory such as a section, table, or list of files.","anchors":["RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential)","FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability'","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'"],"scored":false,"signal":null},{"key":"f_dataset_cited","label":"Dataset formally cited","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No identifier for the dataset appears anywhere in the paper, either in the reference list or in the body text.","anchors":["FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first-","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes","FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'"],"scored":true,"signal":null}]},"A":{"name":"Accessible","score":6.25,"criteria":[{"key":"a_data_openly_accessible","label":"Access route free of preconditions","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":"Please direct requests to access the data, either for reproduction of the work reported here or for other purposes, to recover@chop.edu.","grounded":false,"rationale":"The only access route offered is a discretionary request to an email address, which is not a followable process and is classified as 'no'.","anchors":["RDA-A1.1-01D — 'Data is accessible through a free access protocol'","FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'"],"scored":true,"signal":null},{"key":"a_access_conditions_stated","label":"Access level labelled","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":"Please direct requests to access the data, either for reproduction of the work reported here or for other purposes, to recover@chop.edu.","grounded":false,"rationale":"The paper does not use an explicit access-level label from the standard vocabulary, but it describes the action of requesting access, which allows inference of the level. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (4/5 passes agreed)]","anchors":["FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","RDA-A1-01M — metadata contains information to enable the user to get access to the data","COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl"],"scored":false,"signal":null},{"key":"a_controlled_access_for_sensitive","label":"Gatekeeper for sensitive data","kind":"llm","weight":0.5,"fraction":0.5,"verdict":"partial","evidence":"The data are maintained in a secure enclave, with access managed by the program coordinating center to remain compliant with regulatory and program requirements. Please direct requests to access the data, either for reproduction of the work reported here or for other purposes, to recover@chop.edu.","grounded":false,"rationale":"The gatekeeper is an institutional entity (the program coordinating center) with a defined request route, not a natural person. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (2/5 passes agreed)]","anchors":["NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee","RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and ","NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse"],"scored":false,"signal":null},{"key":"a_timeline_retention","label":"Availability timing & retention","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The paper neither states a persistence commitment nor specifies when the data become available; it only mentions that they are maintained in a secure enclave with access managed.","anchors":["NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines","NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy'","RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'"],"scored":false,"signal":null}]},"I":{"name":"Interoperable","score":60.0,"criteria":[{"key":"i_open_nonproprietary_format","label":"Open file format","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No file format is named for the released data; the data are clinical notes but their format is not specified.","anchors":["FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co","RDA-R1.3-02D — data is expressed in a machine-understandable community standard","RDA-I1-01D — data uses a knowledge representation expressed in a standardised format"],"scored":true,"signal":null},{"key":"i_community_standard_vocabulary","label":"Community standard / vocabulary","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"To facilitate comparison of note-based clinical concepts with structured data, we assembled ICD-10-CM diagnostic code sets for each of the 25 clinical concepts","grounded":true,"rationale":"The paper names ICD-10-CM, which is a community-standard vocabulary registered in FAIRsharing, as applied to the data. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential)","RDA-R1.3-01D — 'Data complies with a community standard'","RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'"],"scored":false,"signal":null},{"key":"i_qualified_references","label":"Identifiers for the resources the data depend on","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"code for both the RE and word2vec training and application is available at https://github.com/RECOVER-Coordinating-Center/pediatric_nlp_manuscript_1","grounded":true,"rationale":"The paper gives a URL for the code, which is a locator for a resource other than the dataset itself.","anchors":["RDA-I3-01M — '(meta)data include references to other (meta)data'","RDA-I3-03M — 'metadata includes qualified references to other metadata'","FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'"],"scored":false,"signal":null}]},"R":{"name":"Reusable","score":41.67,"criteria":[{"key":"r_reuse_license","label":"Reuse licence","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No license or reuse terms are stated for the data; the CC-BY license applies only to the article, not the data.","anchors":["RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu","RDA-R1.1-02M — 'Metadata refers to a standard reuse licence'","RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'"],"scored":true,"signal":null},{"key":"r_provenance_methods","label":"Provenance of the data","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"The pipeline ( Figure 1 ) used openly available as well as commercially licensed Spark-NLP models from John Snow Labs (JSL).","grounded":true,"rationale":"The paper names specific software tools (Spark-NLP, JSL models) used to produce the data, satisfying the class-1 requirement. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa","FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati","W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance"],"scored":false,"signal":null},{"key":"r_documentation_codebook","label":"Documentation / codebook","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":"We identified 21 symptom concepts... An additional 4 clinical concepts addressed impairments in physical, cognitive, sleep, and school functioning. We then created a set of expressions—a pediatric Long COVID lexicon—to represent each concept.","grounded":false,"rationale":"Variable definitions are provided within the article text and tables, but no documentation object (e.g., README, codebook) is said to accompany the data. [downgraded to 'no' — no verifiable quote from the paper]","anchors":["RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'","NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t"],"scored":false,"signal":null},{"key":"r_versioning","label":"Snapshot identified","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No version token or date identifies a snapshot of the data; the time span (2020-2022) is a range, not a specific release. [majority verdict 'no' (3/5 passes agreed)]","anchors":["DataCite Metadata Schema 4.6 — the 'Version' property","RDA-R1.2-01M — provenance information (which version was used is provenance)","NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'"],"scored":true,"signal":null},{"key":"x_code_availability","label":"Analysis code available","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"code for both the RE and word2vec training and application is available at https://github.com/RECOVER-Coordinating-Center/pediatric_nlp_manuscript_1","grounded":true,"rationale":"The paper gives a GitHub URL for the code, which is a machine-resolvable locator in a version-controlled repository.","anchors":["NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code'","FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear","FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)"],"scored":true,"signal":null},{"key":"x_funding_attribution","label":"Funder and award number","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"This work was supported by the National Institutes of Health (NIH) Agreement OTA OT2HL161847-01 as part of the Researching COVID to Enhance Recovery (RECOVER) research Initiative.","grounded":true,"rationale":"The paper includes an award number (OT2HL161847-01) attached to a named funder (NIH).","anchors":["DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award","Crossref Funder Registry — canonical funder identifiers for funding metadata","RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco"],"scored":true,"signal":null}]}},"actions":[{"key":"f_dataset_pid","dimension":"F","label":"Persistent identifier for the data","action":"Mint or cite a persistent identifier for the dataset — a repository DOI or an accession from a registered repository — and print it in the paper. A bare URL is not persistent: it is the single most common cause of a dead data link five years after publication.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The paper does not provide any persistent identifier (DOI, Handle, ARK, or repository accession) for its own dataset.","gain":16.67,"priority":"essential","scored":true},{"key":"f_repository_named","dimension":"F","label":"Named repository","action":"Deposit the data in a repository registered in re3data/FAIRsharing (a domain repository such as GEO, SRA, dbGaP, PRIDE, or a generalist such as Zenodo, Dryad, Dataverse) and name it explicitly in the paper. A lab website is not an archive: it has no retention commitment and no accession.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The paper does not name a repository as the holder of the data; the data are held by the RECOVER program coordinating center with no repository name. [majority verdict 'no' (4/5 passes agreed)]","gain":16.67,"priority":"essential","scored":true},{"key":"a_data_openly_accessible","dimension":"A","label":"Access route free of preconditions","action":"Remove the precondition or justify it. Release the data at publication with no embargo, no registration wall, and no approval step — NIH's zero-embargo public- access rule (NOT-OD-25-101) has already made 'available at publication' the federal baseline for the article; the data should not lag behind it.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"Please direct requests to access the data, either for reproduction of the work reported here or for other purposes, to recover@chop.edu.","why":"The only access route offered is a discretionary request to an email address, which is not a followable process and is classified as 'no'.","gain":16.67,"priority":"essential","scored":true},{"key":"r_reuse_license","dimension":"R","label":"Reuse licence","action":"Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No license or reuse terms are stated for the data; the CC-BY license applies only to the article, not the data.","gain":16.67,"priority":"essential","scored":true},{"key":"f_dataset_cited","dimension":"F","label":"Dataset formally cited","action":"Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No identifier for the dataset appears anywhere in the paper, either in the reference list or in the body text.","gain":8.33,"priority":"important","scored":true},{"key":"i_open_nonproprietary_format","dimension":"I","label":"Open file format","action":"Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No file format is named for the released data; the data are clinical notes but their format is not specified.","gain":8.33,"priority":"important","scored":true},{"key":"r_versioning","dimension":"R","label":"Snapshot identified","action":"Version the deposit and cite the exact version analysed (a version-specific DOI, or an accession with its version suffix). A reader reproducing your work against 'the current release' is reproducing it against a different dataset.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No version token or date identifies a snapshot of the data; the time span (2020-2022) is a range, not a specific release. [majority verdict 'no' (3/5 passes agreed)]","gain":4.17,"priority":"useful","scored":true},{"key":"f_data_availability_statement","dimension":"F","label":"Data-availability statement","action":"Replace the statement with the repository template: name the repository and give the accession or DOI (Colavizza category 3). This is the only DAS class associated with a measured citation advantage; 'available on reasonable request' and 'within the article' are not.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"The results reported here are based on detailed individual-level patient data compiled as part of the RECOVER Program. Due to the high risk of reidentification based on the number of unique patterns in the data, particularly physician notes data, patient privacy regulations prohibit us from releasing the data publicly. The data are maintained in a secure enclave, with access managed by the program coordinating center to remain compliant with regulatory and program requirements. Please direct requests to access the data, either for reproduction of the work reported here or for other purposes, to recover@chop.edu.","why":"The data availability statement points to a person (recover@chop.edu) for access, which is Colavizza category 1 (available on request). [downgraded to 'no' — no verifiable quote from the paper]","gain":0.0,"priority":"essential","scored":false},{"key":"f_discovery_metadata","dimension":"F","label":"Description of the dataset as an object","action":"Add a 'Data Records' section: itemise every file in the deposit and every variable or sample it holds, with counts and units. Describe the dataset as an object in its own right, not as a by-product of the findings — this is what makes it discoverable to someone who is not looking for your paper.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"We analyzed 48 287 outpatient progress notes from 10 618 pediatric patients from 12 institutions.","why":"The dataset is described in running prose in the abstract, but there is no itemised inventory such as a section, table, or list of files.","gain":0.0,"priority":"essential","scored":false},{"key":"a_access_conditions_stated","dimension":"A","label":"Access level labelled","action":"State the access level in words, using the standard vocabulary: 'These data are open access' / 'These data are controlled access'. A reader — and a harvester — should not have to infer the access level from the presence of a download link.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"Please direct requests to access the data, either for reproduction of the work reported here or for other purposes, to recover@chop.edu.","why":"The paper does not use an explicit access-level label from the standard vocabulary, but it describes the action of requesting access, which allows inference of the level. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (4/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"r_documentation_codebook","dimension":"R","label":"Documentation / codebook","action":"Ship a README and a data dictionary IN the deposit — every file, every variable, its units, its allowed values, its missing-value codes. It is the cheapest single thing that makes a dataset usable by someone who was not in the lab, and a table buried in the article does not travel with the data.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"We identified 21 symptom concepts... An additional 4 clinical concepts addressed impairments in physical, cognitive, sleep, and school functioning. We then created a set of expressions—a pediatric Long COVID lexicon—to represent each concept.","why":"Variable definitions are provided within the article text and tables, but no documentation object (e.g., README, codebook) is said to accompany the data. [downgraded to 'no' — no verifiable quote from the paper]","gain":0.0,"priority":"important","scored":false},{"key":"a_controlled_access_for_sensitive","dimension":"A","label":"Gatekeeper for sensitive data","action":"Route sensitive data through an institutional gatekeeper — deposit in a controlled- access repository (dbGaP, EGA) with a Data Access Committee and a published DUA — rather than through the corresponding author's inbox. An author-gated dataset dies with the author's email address, and 'on reasonable request' has been shown repeatedly not to yield data.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"The data are maintained in a secure enclave, with access managed by the program coordinating center to remain compliant with regulatory and program requirements. Please direct requests to access the data, either for reproduction of the work reported here or for other purposes, to recover@chop.edu.","why":"The gatekeeper is an institutional entity (the program coordinating center) with a defined request route, not a natural person. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (2/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false},{"key":"a_timeline_retention","dimension":"A","label":"Availability timing & retention","action":"State when the data become available AND how long they will be retained — cite the repository's preservation policy. NIH DMS Element 4 asks for both; most papers give neither.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The paper neither states a persistence commitment nor specifies when the data become available; it only mentions that they are maintained in a secure enclave with access managed.","gain":0.0,"priority":"useful","scored":false}],"suggestions":["Mint or cite a persistent identifier for the dataset — a repository DOI or an accession from a registered repository — and print it in the paper. A bare URL is not persistent: it is the single most common cause of a dead data link five years after publication.","Deposit the data in a repository registered in re3data/FAIRsharing (a domain repository such as GEO, SRA, dbGaP, PRIDE, or a generalist such as Zenodo, Dryad, Dataverse) and name it explicitly in the paper. A lab website is not an archive: it has no retention commitment and no accession.","Remove the precondition or justify it. Release the data at publication with no embargo, no registration wall, and no approval step — NIH's zero-embargo public- access rule (NOT-OD-25-101) has already made 'available at publication' the federal baseline for the article; the data should not lag behind it.","Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v8","fulltext_source":"epmc_xml"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v8","fair_fulltext_source":"epmc_xml","fair_has_llm":true,"fair_computed_at":"2026-07-20T13:30:34.811570Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}