{"doi":"10.1038/s41597-024-04183-2","title":"Contextualized race and ethnicity annotations for clinical text from MIMIC-III","abstract":"Observational health research often relies on accurate and complete race and ethnicity (RE) patient information, such as characterizing cohorts, assessing quality/performance metrics of hospitals and health systems, and identifying health disparities. While the electronic health record contains structured data such as accessible patient-level RE data, it is often missing, inaccurate, or lacking granular details. Natural language processing models can be trained to identify RE in clinical text which can supplement missing RE data in clinical data repositories. Here we describe the Contextualized Race and Ethnicity Annotations for Clinical Text (C-REACT) Dataset, which comprises 12,000 patients and 17,281 sentences from their clinical notes in the MIMIC-III dataset. Using these sentences, two sets of reference standard annotations for RE data are made available with annotation guidelines. The first set of annotations comprise highly granular information related to RE, such as preferred language and country of origin, while the second set contains RE labels annotated by physicians. This dataset can support health systems' ability to use RE data to serve health equity goals.","journal":"Scientific Data","year":2024,"id":485795,"datarank":0.3693227727154428,"base_score":1.791759469228055,"endowment":1.791759469228055,"self_citation_contribution":0.26876392038420827,"citation_network_contribution":0.10055885233123449,"self_endowment_contribution":0.26876392038420827,"citer_contribution":0.10055885233123449,"corpus_percentile":51.465924034965575,"corpus_rank":6275,"citation_count":5,"citer_count":5,"citers_with_citation_signal":3,"citers_with_endowment":3,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.7788,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2024-01-01","fair_score":58.3333,"fair_percentile":72.8829104249465,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":777031,"name":"Adrienne Pichon","orcid":"0000-0002-9393-3410","position":1,"is_corresponding":false},{"id":1104352,"name":"Harry Reyes Nieva","orcid":"0000-0001-7774-2561","position":2,"is_corresponding":false},{"id":634015,"name":"Tony Sun","orcid":"0000-0002-6482-1188","position":3,"is_corresponding":false},{"id":1328857,"name":"Jun Li","orcid":"0000-0002-3637-1642","position":4,"is_corresponding":false},{"id":437556,"name":"Joshua W. Joseph","orcid":"0000-0001-9704-6635","position":5,"is_corresponding":false},{"id":1328858,"name":"Sivan Kinberg","orcid":"0000-0003-4407-6720","position":6,"is_corresponding":false},{"id":979080,"name":"Lauren R. Richter","orcid":"0000-0001-7319-0480","position":7,"is_corresponding":false},{"id":973571,"name":"Salvatore Crusco","orcid":null,"position":8,"is_corresponding":false},{"id":1329336,"name":"Kyle Kulas","orcid":null,"position":9,"is_corresponding":false},{"id":1329337,"name":"Shaan Ahmed","orcid":null,"position":10,"is_corresponding":false},{"id":1244200,"name":"Daniel J. Snyder","orcid":"0000-0002-6018-1443","position":11,"is_corresponding":false},{"id":1329338,"name":"Ashkon Rahbari","orcid":null,"position":12,"is_corresponding":false},{"id":244466,"name":"Benjamin L. Ranard","orcid":"0000-0002-9565-6939","position":13,"is_corresponding":false},{"id":1328859,"name":"Pallavi Juneja","orcid":"0000-0002-1995-9823","position":14,"is_corresponding":false},{"id":483824,"name":"Dina Demner‐Fushman","orcid":null,"position":15,"is_corresponding":false},{"id":675186,"name":"Noémie Elhadad","orcid":"0000-0001-9721-5240","position":16,"is_corresponding":false},{"id":307673,"name":"Oliver J. Bear Don’t Walk","orcid":"0000-0003-4310-9279","position":0,"is_corresponding":true}],"reference_count":51,"raw_metadata":null,"created_at":"2026-07-19T02:07:52.536246Z","pmid":"39638783","pmcid":"PMC11621419","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":72.2222,"fair_a":50.0,"fair_i":60.0,"fair_r":58.3333,"fair_zscore":0.9454,"fair_rationale":{"fair_score":58.33,"has_llm":true,"taxonomy_version":"fair_taxonomy_v5","dimensions":{"F":{"name":"Findable","score":72.22,"criteria":[{"key":"f_dataset_pid","label":"Persistent identifier for the data","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"The C-REACT Dataset project page on PhysioNet provides further details on the dataset and the access application process (https://doi.org/10.13026/t9ka-6k29)","grounded":false,"rationale":"The paper provides a DOI (10.13026/t9ka-6k29) for the dataset. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]","anchors":["RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit","RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier'","FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'"],"scored":true,"signal":null},{"key":"f_repository_named","label":"Named repository","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Both sets of reference standard annotations are available on PhysioNet 50.","grounded":true,"rationale":"PhysioNet is a named data repository. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed (","NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived","NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten"],"scored":true,"signal":null},{"key":"f_data_availability_statement","label":"Data-availability statement","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"Both sets of reference standard annotations are available on PhysioNet.","grounded":false,"rationale":"The statement points to a repository record (PhysioNet) with a DOI, corresponding to Colavizza category 3. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]","anchors":["Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li","Springer Nature research data policy — Data Availability Statements: standard statement templat","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes"],"scored":false,"signal":null},{"key":"f_discovery_metadata","label":"Description of the dataset as an object","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"There is one main folder containing two jsonl files with span-level RE indicators and sentence-level RE assignments as well as a subdirectory with raw RE assignment files.","grounded":true,"rationale":"The 'Data Records' section provides an itemised inventory of the dataset files.","anchors":["RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential)","FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability'","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'"],"scored":false,"signal":null},{"key":"f_dataset_cited","label":"Dataset formally cited","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"Bear Don’t Walk, O. J. et al. C-REACT: Contextualized Race and Ethnicity Annotations for Clinical Text. physionet.org https://doi.org/10.13026/C2JT1Q (2023).","grounded":false,"rationale":"The dataset appears as a reference-list entry (reference 50), cited in the text. [downgraded to 'partial' — no verifiable quote from the paper]","anchors":["FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first-","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes","FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'"],"scored":true,"signal":null}]},"A":{"name":"Accessible","score":50.0,"criteria":[{"key":"a_data_openly_accessible","label":"Access route free of preconditions","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"users are required to register, complete a credentialing process, and sign a data use agreement.","grounded":true,"rationale":"The access route carries a stated precondition (registration, credentialing, DUA), making it a followable process but not unconditional. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-A1.1-01D — 'Data is accessible through a free access protocol'","FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'"],"scored":true,"signal":null},{"key":"a_access_conditions_stated","label":"Access level labelled","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"While PhysioNet 51 data are freely accessible, users are required to register, complete a credentialing process, and sign a data use agreement.","grounded":true,"rationale":"The paper describes the access process but does not label the access level with a standard term like 'open access' or 'restricted access'. [majority verdict 'partial' (2/5 passes agreed)]","anchors":["FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","RDA-A1-01M — metadata contains information to enable the user to get access to the data","COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl"],"scored":false,"signal":null},{"key":"a_controlled_access_for_sensitive","label":"Gatekeeper for sensitive data","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"users are required to register, complete a credentialing process, and sign a data use agreement.","grounded":true,"rationale":"The paper names PhysioNet, an institutional repository, as the gatekeeper with a defined credentialing process and data use agreement. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee","RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and ","NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse"],"scored":false,"signal":null},{"key":"a_timeline_retention","label":"Availability timing & retention","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The paper does not state any retention period or permanent archival claim for the data. [majority verdict 'no' (4/5 passes agreed)]","anchors":["NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines","NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy'","RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'"],"scored":false,"signal":null}]},"I":{"name":"Interoperable","score":60.0,"criteria":[{"key":"i_open_nonproprietary_format","label":"Open file format","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"jsonl files","grounded":false,"rationale":"The dataset is provided in JSONL format, which is an open, community-standard format. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]","anchors":["FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co","RDA-R1.3-02D — data is expressed in a machine-understandable community standard","RDA-I1-01D — data uses a knowledge representation expressed in a standardised format"],"scored":true,"signal":null},{"key":"i_community_standard_vocabulary","label":"Community standard / vocabulary","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"U.S. census categories","grounded":true,"rationale":"The paper uses the U.S. census categories (a federally recognized community standard) for race and ethnicity labeling. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential)","RDA-R1.3-01D — 'Data complies with a community standard'","RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'"],"scored":false,"signal":null},{"key":"i_qualified_references","label":"Identifiers for the resources the data depend on","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":"All code is made publicly available through GitHub (https://github.com/elhadadlab/MIMIC_race_ethnicity_dataset).","grounded":false,"rationale":"The paper provides a URL for the code repository, which is an identifier for a resource other than the own dataset. [downgraded to 'no' — no verifiable quote from the paper]","anchors":["RDA-I3-01M — '(meta)data include references to other (meta)data'","RDA-I3-03M — 'metadata includes qualified references to other metadata'","FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'"],"scored":false,"signal":null}]},"R":{"name":"Reusable","score":58.33,"criteria":[{"key":"r_reuse_license","label":"Reuse licence","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"users are required to register, complete a credentialing process, and sign a data use agreement.","grounded":true,"rationale":"A data use agreement is named, which is a terms artefact, but it is not an open standard licence. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu","RDA-R1.1-02M — 'Metadata refers to a standard reuse licence'","RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'"],"scored":true,"signal":null},{"key":"r_provenance_methods","label":"Provenance of the data","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"We used NLTK40 to extract sentences and heuristics to handle clinical lists such as medication and condition lists.","grounded":true,"rationale":"The paper names the specific tool (NLTK) used to produce the data.","anchors":["RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa","FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati","W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance"],"scored":false,"signal":null},{"key":"r_documentation_codebook","label":"Documentation / codebook","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"The annotation guidelines for indicators are available in the PhysioNet dataset under the file name “File_1_Annotation_Guidelines_for_Race_and_Ethnicity_Indicators.docx”.","grounded":false,"rationale":"A documentation object (annotation guidelines) is named as accompanying the data in the repository. [downgraded to 'partial' — no verifiable quote from the paper]","anchors":["RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'","NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t"],"scored":false,"signal":null},{"key":"r_versioning","label":"Snapshot identified","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No version token or date is given for the C-REACT dataset itself.","anchors":["DataCite Metadata Schema 4.6 — the 'Version' property","RDA-R1.2-01M — provenance information (which version was used is provenance)","NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'"],"scored":true,"signal":null},{"key":"x_code_availability","label":"Analysis code available","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"All code is made publicly available through GitHub (https://github.com/elhadadlab/MIMIC_race_ethnicity_dataset).","grounded":false,"rationale":"A machine-resolvable code repository URL is provided. [downgraded to 'partial' — no verifiable quote from the paper]","anchors":["NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code'","FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear","FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)"],"scored":true,"signal":null},{"key":"x_funding_attribution","label":"Funder and award number","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"T15LM007079","grounded":true,"rationale":"The paper includes specific grant numbers (e.g., T15LM007079) from named funders. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award","Crossref Funder Registry — canonical funder identifiers for funding metadata","RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco"],"scored":true,"signal":null}]}},"actions":[{"key":"f_dataset_pid","dimension":"F","label":"Persistent identifier for the data","action":"Mint or cite a persistent identifier for the dataset — a repository DOI or an accession from a registered repository — and print it in the paper. A bare URL is not persistent: it is the single most common cause of a dead data link five years after publication. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"The C-REACT Dataset project page on PhysioNet provides further details on the dataset and the access application process (https://doi.org/10.13026/t9ka-6k29)","why":"The paper provides a DOI (10.13026/t9ka-6k29) for the dataset. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]","gain":8.33,"priority":"essential","scored":true},{"key":"a_data_openly_accessible","dimension":"A","label":"Access route free of preconditions","action":"Remove the precondition or justify it. Release the data at publication with no embargo, no registration wall, and no approval step — NIH's zero-embargo public- access rule (NOT-OD-25-101) has already made 'available at publication' the federal baseline for the article; the data should not lag behind it. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"users are required to register, complete a credentialing process, and sign a data use agreement.","why":"The access route carries a stated precondition (registration, credentialing, DUA), making it a followable process but not unconditional. [majority verdict 'partial' (4/5 passes agreed)]","gain":8.33,"priority":"essential","scored":true},{"key":"r_reuse_license","dimension":"R","label":"Reuse licence","action":"Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"users are required to register, complete a credentialing process, and sign a data use agreement.","why":"A data use agreement is named, which is a terms artefact, but it is not an open standard licence. [majority verdict 'partial' (4/5 passes agreed)]","gain":8.33,"priority":"essential","scored":true},{"key":"f_dataset_cited","dimension":"F","label":"Dataset formally cited","action":"Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the clinical / human-subjects repository accession (e.g. from dbGaP or the European Genome-phenome Archive (EGA)) in the reference list.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Bear Don’t Walk, O. J. et al. C-REACT: Contextualized Race and Ethnicity Annotations for Clinical Text. physionet.org https://doi.org/10.13026/C2JT1Q (2023).","why":"The dataset appears as a reference-list entry (reference 50), cited in the text. [downgraded to 'partial' — no verifiable quote from the paper]","gain":4.17,"priority":"important","scored":true},{"key":"i_open_nonproprietary_format","dimension":"I","label":"Open file format","action":"Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"jsonl files","why":"The dataset is provided in JSONL format, which is an open, community-standard format. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]","gain":4.17,"priority":"important","scored":true},{"key":"x_code_availability","dimension":"R","label":"Analysis code available","action":"Publish the analysis code in a public forge, archive a tagged release with a DOI (Zenodo/Software Heritage), and cite that DOI in the paper. NIH DMS Element 2 asks for the tools and code, not only the data — and 'available on request' is not a locator. Archive the analysis code in a versioned repository (GitHub + a Zenodo release DOI).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"All code is made publicly available through GitHub (https://github.com/elhadadlab/MIMIC_race_ethnicity_dataset).","why":"A machine-resolvable code repository URL is provided. [downgraded to 'partial' — no verifiable quote from the paper]","gain":4.17,"priority":"important","scored":true},{"key":"r_versioning","dimension":"R","label":"Snapshot identified","action":"Version the deposit and cite the exact version analysed (a version-specific DOI, or an accession with its version suffix). A reader reproducing your work against 'the current release' is reproducing it against a different dataset.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No version token or date is given for the C-REACT dataset itself.","gain":4.17,"priority":"useful","scored":true},{"key":"f_data_availability_statement","dimension":"F","label":"Data-availability statement","action":"Replace the statement with the repository template: name the repository and give the accession or DOI (Colavizza category 3). This is the only DAS class associated with a measured citation advantage; 'available on reasonable request' and 'within the article' are not.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Both sets of reference standard annotations are available on PhysioNet.","why":"The statement points to a repository record (PhysioNet) with a DOI, corresponding to Colavizza category 3. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]","gain":0.0,"priority":"essential","scored":false},{"key":"a_access_conditions_stated","dimension":"A","label":"Access level labelled","action":"State the access level in words, using the standard vocabulary: 'These data are open access' / 'These data are controlled access'. A reader — and a harvester — should not have to infer the access level from the presence of a download link.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"While PhysioNet 51 data are freely accessible, users are required to register, complete a credentialing process, and sign a data use agreement.","why":"The paper describes the access process but does not label the access level with a standard term like 'open access' or 'restricted access'. [majority verdict 'partial' (2/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"r_documentation_codebook","dimension":"R","label":"Documentation / codebook","action":"Ship a README and a data dictionary IN the deposit — every file, every variable, its units, its allowed values, its missing-value codes. It is the cheapest single thing that makes a dataset usable by someone who was not in the lab, and a table buried in the article does not travel with the data.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"The annotation guidelines for indicators are available in the PhysioNet dataset under the file name “File_1_Annotation_Guidelines_for_Race_and_Ethnicity_Indicators.docx”.","why":"A documentation object (annotation guidelines) is named as accompanying the data in the repository. [downgraded to 'partial' — no verifiable quote from the paper]","gain":0.0,"priority":"important","scored":false},{"key":"i_qualified_references","dimension":"I","label":"Identifiers for the resources the data depend on","action":"Cite by identifier every resource the data depend on — the source datasets' accessions, the reference build (GRCh38 / GCA_000001405.28), the cohort application number, the code DOI — and register those relations on the dataset record (IsDerivedFrom, IsSupplementTo). A name is not a link: it cannot be resolved, versioned, or followed by a machine.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"All code is made publicly available through GitHub (https://github.com/elhadadlab/MIMIC_race_ethnicity_dataset).","why":"The paper provides a URL for the code repository, which is an identifier for a resource other than the own dataset. [downgraded to 'no' — no verifiable quote from the paper]","gain":0.0,"priority":"useful","scored":false},{"key":"a_timeline_retention","dimension":"A","label":"Availability timing & retention","action":"State when the data become available AND how long they will be retained — cite the repository's preservation policy. NIH DMS Element 4 asks for both; most papers give neither.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The paper does not state any retention period or permanent archival claim for the data. [majority verdict 'no' (4/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false}],"suggestions":["Mint or cite a persistent identifier for the dataset — a repository DOI or an accession from a registered repository — and print it in the paper. A bare URL is not persistent: it is the single most common cause of a dead data link five years after publication. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","Remove the precondition or justify it. Release the data at publication with no embargo, no registration wall, and no approval step — NIH's zero-embargo public- access rule (NOT-OD-25-101) has already made 'available at publication' the federal baseline for the article; the data should not lag behind it. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the clinical / human-subjects repository accession (e.g. from dbGaP or the European Genome-phenome Archive (EGA)) in the reference list.","Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v8","fulltext_source":"unpaywall_pdf"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v8","fair_fulltext_source":"unpaywall_pdf","fair_has_llm":true,"fair_computed_at":"2026-07-20T12:56:12.073608Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}