{"doi":"10.1093/jncics/pkaf027","title":"Building research infrastructure to advance precision medicine in colorectal cancer","abstract":"BACKGROUND: Addressing critical gaps in precision medicine initiatives in colorectal cancer (CRC) requires building larger collaborative studies. METHODS: The Latino Colorectal Cancer Consortium (LC3) is a resource that harmonizes data collected in observational studies with data from individuals who identify as Hispanic/Latino with a diagnosis of primary colorectal adenocarcinoma. Data collected includes demographics, medical history, family history, and lifestyle risk factors from patient-completed surveys. Vital status, cause of death, treatment, and clinicopathological characteristics were obtained through medical chart abstraction, pathology reports, and/or linkage to state cancer registries. Blood, saliva, or normal colonic tissues were used to extract and genotype germline DNA. Tumor tissue (snap frozen or formalin-fixed paraffin-embedded) was evaluated by pathologists for diagnosis, tissue content, tumor cellularity, necrosis, immune infiltration, and additional histopathological characteristics. A centralized database with a virtual tumor repository was created to facilitate collaborative research. RESULTS: As of April 2024, LC3 assembled data from 2210 patients (diagnosed 1994 to 2023). The mean age at diagnosis was 57 (range: 19-93) years; 54.3% of participants were male, and 62.0% had been diagnosed with colon cancer. Surveys were completed by 1722 (77.8%) participants. Ongoing multi-omics profiling on up to 600 patients include: genome-wide germline genotyping, paired tumor/normal whole exome sequencing, bulk RNA-seq, T cell receptor immunosequencing, and multiplex immunofluorescence. CONCLUSIONS: This consortium fills an important gap in research infrastructure in CRC as well as improving precision medicine initiatives for all individuals.","journal":"JNCI Cancer Spectrum","year":2025,"id":563240,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":0.0,"corpus_rank":10062,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9231,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":52.0833,"fair_percentile":67.4411494955671,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1087524,"name":"Nicole C. Loroña","orcid":"0000-0002-8302-3157","position":1,"is_corresponding":false},{"id":1466351,"name":"Daniel Sobieski","orcid":null,"position":2,"is_corresponding":false},{"id":1466055,"name":"Marco Matejcic","orcid":"0009-0004-9468-0664","position":3,"is_corresponding":false},{"id":305218,"name":"Nathalie T. Nguyen","orcid":"0009-0001-4282-103X","position":4,"is_corresponding":false},{"id":1332613,"name":"Hannah J. Hoehn","orcid":"0000-0001-7056-9671","position":5,"is_corresponding":false},{"id":1240768,"name":"Diana B. Díaz","orcid":null,"position":6,"is_corresponding":false},{"id":1332614,"name":"Kritika Shankar","orcid":"0009-0007-4120-2900","position":7,"is_corresponding":false},{"id":287319,"name":"Eric Cockman","orcid":null,"position":8,"is_corresponding":false},{"id":1466056,"name":"Esther Jean‐Baptiste","orcid":"0000-0002-5677-3104","position":9,"is_corresponding":false},{"id":1466057,"name":"Ya-Yu Tsai","orcid":"0000-0002-5085-6400","position":10,"is_corresponding":false},{"id":920912,"name":"R. Blake Buchalter","orcid":"0000-0003-3605-904X","position":11,"is_corresponding":false},{"id":1466352,"name":"Karina Brito","orcid":null,"position":12,"is_corresponding":false},{"id":1466353,"name":"Rusche Wilson","orcid":null,"position":13,"is_corresponding":false},{"id":1466058,"name":"Domenico Coppola","orcid":"0009-0008-3540-2014","position":14,"is_corresponding":false},{"id":572174,"name":"Clifton G. Fulmer","orcid":"0000-0003-1453-9421","position":15,"is_corresponding":false},{"id":977315,"name":"Ozlen Saglam","orcid":"0000-0003-3216-239X","position":16,"is_corresponding":false},{"id":1143481,"name":"Alexandra F. Tassielli","orcid":"0009-0009-1445-8227","position":17,"is_corresponding":false},{"id":688340,"name":"Francisca Beato","orcid":"0009-0003-0147-4013","position":18,"is_corresponding":false},{"id":1143916,"name":"Ruifan Dai","orcid":null,"position":19,"is_corresponding":false},{"id":32930,"name":"Jennifer A. Freedman","orcid":"0000-0001-6771-4161","position":20,"is_corresponding":false},{"id":264005,"name":"Kristen S. Purrington","orcid":"0000-0002-5710-1692","position":21,"is_corresponding":false},{"id":455368,"name":"Bo Hu","orcid":"0000-0003-0606-7639","position":22,"is_corresponding":false},{"id":268272,"name":"Daniel J. McGrail","orcid":"0000-0002-6669-6069","position":23,"is_corresponding":false},{"id":299803,"name":"Heather M. Gibson","orcid":"0000-0003-2652-8252","position":24,"is_corresponding":false},{"id":317583,"name":"Kun Jiang","orcid":"0000-0001-8531-7813","position":25,"is_corresponding":false},{"id":1332922,"name":"Teresita Muñoz-Antonia","orcid":null,"position":26,"is_corresponding":false},{"id":426573,"name":"Idhaliz Flores","orcid":"0000-0003-2130-4508","position":27,"is_corresponding":false},{"id":1332923,"name":"Edna R. Gordián","orcid":null,"position":28,"is_corresponding":false},{"id":951453,"name":"José A. Oliveras Torres","orcid":null,"position":29,"is_corresponding":false},{"id":252204,"name":"Iona Cheng","orcid":"0000-0003-4132-2893","position":30,"is_corresponding":false},{"id":332144,"name":"Erin L. Van Blarigan","orcid":"0000-0001-5079-3385","position":31,"is_corresponding":false},{"id":897190,"name":"Seth Felder","orcid":"0000-0001-6975-8379","position":32,"is_corresponding":false},{"id":443628,"name":"Julian Sanchez","orcid":"0000-0002-7198-839X","position":33,"is_corresponding":false},{"id":288150,"name":"Jason B. Fleming","orcid":"0000-0002-8796-5627","position":34,"is_corresponding":false},{"id":695513,"name":"Erin M. Siegel","orcid":"0000-0003-1779-2510","position":35,"is_corresponding":false},{"id":1332921,"name":"Douglas W. Cress","orcid":null,"position":36,"is_corresponding":false},{"id":553577,"name":"Patricia Thompson","orcid":"0009-0007-3799-0334","position":37,"is_corresponding":false},{"id":107681,"name":"Mariana C. Stern","orcid":"0000-0001-9777-4683","position":38,"is_corresponding":false},{"id":318017,"name":"Jamie K. Teer","orcid":"0000-0003-4513-0282","position":39,"is_corresponding":false},{"id":286040,"name":"Jane C. Figueiredo","orcid":"0000-0001-8040-3341","position":40,"is_corresponding":false},{"id":107679,"name":"Stephanie L. Schmit","orcid":"0000-0001-5931-1194","position":0,"is_corresponding":true}],"reference_count":59,"raw_metadata":null,"created_at":"2026-07-19T02:56:13.374766Z","pmid":"40111849","pmcid":"PMC12140009","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":94.4444,"fair_a":56.25,"fair_i":0.0,"fair_r":37.5,"fair_zscore":0.698,"fair_rationale":{"fair_score":52.08,"has_llm":true,"taxonomy_version":"fair_taxonomy_v5","dimensions":{"F":{"name":"Findable","score":94.44,"criteria":[{"key":"f_dataset_pid","label":"Persistent identifier for the data","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Whole exome sequencing, genetic ancestry proportions, and core analysis variables will be available through dbGaP (phs003464).","grounded":true,"rationale":"The paper provides a dbGaP accession (phs003464), which is a persistent identifier scheme recognized by the FAIR rubric.","anchors":["RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit","RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier'","FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'"],"scored":true,"signal":null},{"key":"f_repository_named","label":"Named repository","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"dbGaP","grounded":true,"rationale":"The paper names dbGaP as the repository for the whole exome sequencing and core analysis variables, and dbGaP is a known data repository listed in re3data/FAIRsharing.","anchors":["RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed (","NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived","NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten"],"scored":true,"signal":null},{"key":"f_data_availability_statement","label":"Data-availability statement","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Whole exome sequencing, genetic ancestry proportions, and core analysis variables will be available through dbGaP (phs003464)","grounded":true,"rationale":"The statement points to a repository record (dbGaP with accession), which is Colavizza category 3. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li","Springer Nature research data policy — Data Availability Statements: standard statement templat","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes"],"scored":false,"signal":null},{"key":"f_discovery_metadata","label":"Description of the dataset as an object","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Table 2. Biospecimen and data resources obtained or to be generated (TBG) in the LC3.","grounded":true,"rationale":"The paper includes Table 2, an itemized inventory of the dataset's files, samples, and variables, which is a structured description of the dataset.","anchors":["RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential)","FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability'","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'"],"scored":false,"signal":null},{"key":"f_dataset_cited","label":"Dataset formally cited","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"Whole exome sequencing, genetic ancestry proportions, and core analysis variables will be available through dbGaP (phs003464).","grounded":true,"rationale":"The dataset identifier (phs003464) appears only in the body text (Data Availability Statement) and not in the reference list.","anchors":["FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first-","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes","FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'"],"scored":true,"signal":null}]},"A":{"name":"Accessible","score":56.25,"criteria":[{"key":"a_data_openly_accessible","label":"Access route free of preconditions","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"Applications will be reviewed based on scientific merit and consortium priorities for usage of nonrenewable biospecimen resources.","grounded":true,"rationale":"The text states a defined, followable precondition (review by the Steering Committee and DUA) for accessing the data, so the access route carries a defined precondition. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-A1.1-01D — 'Data is accessible through a free access protocol'","FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'"],"scored":true,"signal":null},{"key":"a_access_conditions_stated","label":"Access level labelled","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"Outside of public repositories, data and biospecimens are available on a collaborative basis through a standardized proposal system through the LC3 Steering Committee, with the proposal template available upon request from the corresponding authors.","grounded":true,"rationale":"The paper describes an access action (proposal and review) but does not apply an explicit access-level label such as 'open access' or 'restricted access'. [majority verdict 'partial' (3/5 passes agreed)]","anchors":["FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","RDA-A1-01M — metadata contains information to enable the user to get access to the data","COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl"],"scored":false,"signal":null},{"key":"a_controlled_access_for_sensitive","label":"Gatekeeper for sensitive data","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"Outside of public repositories, data and biospecimens are available on a collaborative basis through a standardized proposal system through the LC3 Steering Committee","grounded":true,"rationale":"The LC3 Steering Committee is an institutional gatekeeper that reviews applications for data access, which qualifies as a named institutional gatekeeper.","anchors":["NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee","RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and ","NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse"],"scored":false,"signal":null},{"key":"a_timeline_retention","label":"Availability timing & retention","kind":"llm","weight":0.5,"fraction":0.5,"verdict":"partial","evidence":"Whole exome sequencing, genetic ancestry proportions, and core analysis variables will be available through dbGaP (phs003464).","grounded":true,"rationale":"The text states when the data will become available (future) but does not specify how long they will persist, so it is an availability-timing statement only. [majority verdict 'partial' (3/5 passes agreed)]","anchors":["NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines","NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy'","RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'"],"scored":false,"signal":null}]},"I":{"name":"Interoperable","score":0.0,"criteria":[{"key":"i_open_nonproprietary_format","label":"Open file format","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No file format token (e.g., FASTQ, BAM, VCF) is mentioned for the released data; only assay types are described.","anchors":["FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co","RDA-R1.3-02D — data is expressed in a machine-understandable community standard","RDA-I1-01D — data uses a knowledge representation expressed in a standardised format"],"scored":true,"signal":null},{"key":"i_community_standard_vocabulary","label":"Community standard / vocabulary","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No data or metadata community standard (e.g., MIAME, BIDS, an ontology) is named in the text for the dataset.","anchors":["RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential)","RDA-R1.3-01D — 'Data complies with a community standard'","RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'"],"scored":false,"signal":null},{"key":"i_qualified_references","label":"Identifiers for the resources the data depend on","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No identifier for an external resource (other than the study's own dataset) is provided; the references are to the study's own dbGaP accessions or to papers without data identifiers.","anchors":["RDA-I3-01M — '(meta)data include references to other (meta)data'","RDA-I3-03M — 'metadata includes qualified references to other metadata'","FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'"],"scored":false,"signal":null}]},"R":{"name":"Reusable","score":37.5,"criteria":[{"key":"r_reuse_license","label":"Reuse licence","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No reuse licence is named for the data; the CC-BY-NC-ND licence applies to the article, not the data.","anchors":["RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu","RDA-R1.1-02M — 'Metadata refers to a standard reuse licence'","RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'"],"scored":true,"signal":null},{"key":"r_provenance_methods","label":"Provenance of the data","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"germline DNA was extracted and genotyped using the Illumina Infinium OncoArray-500K BeadChip and/or HumanOmni2.5Exome-8 BeadChip array for a subset of HCCS and TCC participants.","grounded":true,"rationale":"The paper names specific instruments and platforms (Illumina BeadChip arrays) used to produce the data. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa","FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati","W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance"],"scored":false,"signal":null},{"key":"r_documentation_codebook","label":"Documentation / codebook","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"Table 1. Distribution of demographic and clinical characteristics in the LC3 ( N = 2210).","grounded":true,"rationale":"Variable definitions are provided inside the article (Table 1), but no separate documentation object (e.g., README, codebook) is named as travelling with the data. [majority verdict 'partial' (3/5 passes agreed)]","anchors":["RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'","NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t"],"scored":false,"signal":null},{"key":"r_versioning","label":"Snapshot identified","kind":"llm","weight":0.5,"fraction":0.5,"verdict":"partial","evidence":"As of April 2024, LC3 assembled data from 2210 patients","grounded":true,"rationale":"The paper gives a date (April 2024) as the data cut-off point, but no version token is provided. [majority verdict 'partial' (3/5 passes agreed)]","anchors":["DataCite Metadata Schema 4.6 — the 'Version' property","RDA-R1.2-01M — provenance information (which version was used is provenance)","NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'"],"scored":true,"signal":null},{"key":"x_code_availability","label":"Analysis code available","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The paper does not provide a locator for the study's own code; it mentions third-party packages (Rmonize) but not the consortium's analysis code.","anchors":["NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code'","FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear","FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)"],"scored":true,"signal":null},{"key":"x_funding_attribution","label":"Funder and award number","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"R01CA155101","grounded":true,"rationale":"The paper lists multiple NIH grant numbers, including R01CA155101, which is an award/grant identifier.","anchors":["DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award","Crossref Funder Registry — canonical funder identifiers for funding metadata","RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco"],"scored":true,"signal":null}]}},"actions":[{"key":"r_reuse_license","dimension":"R","label":"Reuse licence","action":"Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No reuse licence is named for the data; the CC-BY-NC-ND licence applies to the article, not the data.","gain":16.67,"priority":"essential","scored":true},{"key":"a_data_openly_accessible","dimension":"A","label":"Access route free of preconditions","action":"Remove the precondition or justify it. Release the data at publication with no embargo, no registration wall, and no approval step — NIH's zero-embargo public- access rule (NOT-OD-25-101) has already made 'available at publication' the federal baseline for the article; the data should not lag behind it. For genomics / sequencing data, deposit in GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Applications will be reviewed based on scientific merit and consortium priorities for usage of nonrenewable biospecimen resources.","why":"The text states a defined, followable precondition (review by the Steering Committee and DUA) for accessing the data, so the access route carries a defined precondition. [majority verdict 'partial' (4/5 passes agreed)]","gain":8.33,"priority":"essential","scored":true},{"key":"i_open_nonproprietary_format","dimension":"I","label":"Open file format","action":"Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable. Prefer open genomics / sequencing formats such as FASTQ, BAM or VCF.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No file format token (e.g., FASTQ, BAM, VCF) is mentioned for the released data; only assay types are described.","gain":8.33,"priority":"important","scored":true},{"key":"x_code_availability","dimension":"R","label":"Analysis code available","action":"Publish the analysis code in a public forge, archive a tagged release with a DOI (Zenodo/Software Heritage), and cite that DOI in the paper. NIH DMS Element 2 asks for the tools and code, not only the data — and 'available on request' is not a locator. Archive the analysis code in a versioned repository (GitHub + a Zenodo release DOI).","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The paper does not provide a locator for the study's own code; it mentions third-party packages (Rmonize) but not the consortium's analysis code.","gain":8.33,"priority":"important","scored":true},{"key":"f_dataset_cited","dimension":"F","label":"Dataset formally cited","action":"Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the genomics / sequencing repository accession (e.g. from GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA)) in the reference list.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Whole exome sequencing, genetic ancestry proportions, and core analysis variables will be available through dbGaP (phs003464).","why":"The dataset identifier (phs003464) appears only in the body text (Data Availability Statement) and not in the reference list.","gain":4.17,"priority":"important","scored":true},{"key":"r_versioning","dimension":"R","label":"Snapshot identified","action":"Version the deposit and cite the exact version analysed (a version-specific DOI, or an accession with its version suffix). A reader reproducing your work against 'the current release' is reproducing it against a different dataset.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"As of April 2024, LC3 assembled data from 2210 patients","why":"The paper gives a date (April 2024) as the data cut-off point, but no version token is provided. [majority verdict 'partial' (3/5 passes agreed)]","gain":2.08,"priority":"useful","scored":true},{"key":"a_access_conditions_stated","dimension":"A","label":"Access level labelled","action":"State the access level in words, using the standard vocabulary: 'These data are open access' / 'These data are controlled access'. A reader — and a harvester — should not have to infer the access level from the presence of a download link.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Outside of public repositories, data and biospecimens are available on a collaborative basis through a standardized proposal system through the LC3 Steering Committee, with the proposal template available upon request from the corresponding authors.","why":"The paper describes an access action (proposal and review) but does not apply an explicit access-level label such as 'open access' or 'restricted access'. [majority verdict 'partial' (3/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"i_community_standard_vocabulary","dimension":"I","label":"Community standard / vocabulary","action":"Adopt and NAME your domain's data standard — the minimum-information checklist, metadata schema, or ontology your community uses (MIAME/MINSEQE, ISA-Tab, BIDS, an OBO ontology, HL7 FHIR/OMOP) — and say which one you followed. A reporting checklist standardises your paper; it does nothing for your data. In genomics / sequencing, describe the data with MIAME, MINSEQE or MIxS.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No data or metadata community standard (e.g., MIAME, BIDS, an ontology) is named in the text for the dataset.","gain":0.0,"priority":"important","scored":false},{"key":"r_documentation_codebook","dimension":"R","label":"Documentation / codebook","action":"Ship a README and a data dictionary IN the deposit — every file, every variable, its units, its allowed values, its missing-value codes. It is the cheapest single thing that makes a dataset usable by someone who was not in the lab, and a table buried in the article does not travel with the data.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Table 1. Distribution of demographic and clinical characteristics in the LC3 ( N = 2210).","why":"Variable definitions are provided inside the article (Table 1), but no separate documentation object (e.g., README, codebook) is named as travelling with the data. [majority verdict 'partial' (3/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"i_qualified_references","dimension":"I","label":"Identifiers for the resources the data depend on","action":"Cite by identifier every resource the data depend on — the source datasets' accessions, the reference build (GRCh38 / GCA_000001405.28), the cohort application number, the code DOI — and register those relations on the dataset record (IsDerivedFrom, IsSupplementTo). A name is not a link: it cannot be resolved, versioned, or followed by a machine.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No identifier for an external resource (other than the study's own dataset) is provided; the references are to the study's own dbGaP accessions or to papers without data identifiers.","gain":0.0,"priority":"useful","scored":false},{"key":"a_timeline_retention","dimension":"A","label":"Availability timing & retention","action":"State when the data become available AND how long they will be retained — cite the repository's preservation policy. NIH DMS Element 4 asks for both; most papers give neither.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Whole exome sequencing, genetic ancestry proportions, and core analysis variables will be available through dbGaP (phs003464).","why":"The text states when the data will become available (future) but does not specify how long they will persist, so it is an availability-timing statement only. [majority verdict 'partial' (3/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false}],"suggestions":["Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","Remove the precondition or justify it. Release the data at publication with no embargo, no registration wall, and no approval step — NIH's zero-embargo public- access rule (NOT-OD-25-101) has already made 'available at publication' the federal baseline for the article; the data should not lag behind it. For genomics / sequencing data, deposit in GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA).","Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable. Prefer open genomics / sequencing formats such as FASTQ, BAM or VCF.","Publish the analysis code in a public forge, archive a tagged release with a DOI (Zenodo/Software Heritage), and cite that DOI in the paper. NIH DMS Element 2 asks for the tools and code, not only the data — and 'available on request' is not a locator. Archive the analysis code in a versioned repository (GitHub + a Zenodo release DOI).","Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the genomics / sequencing repository accession (e.g. from GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA)) in the reference list."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v8","fulltext_source":"epmc_xml"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v8","fair_fulltext_source":"epmc_xml","fair_has_llm":true,"fair_computed_at":"2026-07-20T13:52:10.255233Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}