{"doi":"10.1093/jhered/esaf023","title":"A high-quality genome assembly for a desert-adapted rodent, Merriam’s kangaroo rat (<i>Dipodomys merriami</i>)","abstract":"Merriam's kangaroo rat (Dipodomys merriami) is a member of a unique family of primarily desert-adapted North American rodents (Heteromyidae). Of the 20 species in the genus, D. merriami is one of the most wide-ranging and ecologically flexible, inhabiting desert scrub, grassland, sagebrush steppe, and juniper-piñon woodland in the southwestern deserts of the United States and Mexico. We present a de novo reference genome for D. merriami generated from PacBio HiFi long-read and Omni-C chromatin proximity sequencing as a part of the California Conservation Genomics Project. The primary pseudo-haplotype assembly comprises 3,110 scaffolds, with a contig N50 of 8.6 Mb, scaffold N50 of 49.1 Mb, and a total length of 3.57 Gb. Further, a BUSCO completeness score of 97.8% suggests that the assembly is highly complete. This reference genome will serve as a resource for future studies of Dipodomys conservation genomics, desert adaptation, and phylogeography.","journal":"Journal of Heredity","year":2025,"id":564360,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":0.0,"corpus_rank":10062,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.7679,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":66.6667,"fair_percentile":86.48731274839498,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":92291,"name":"Merly Escalona","orcid":"0000-0003-0213-4777","position":1,"is_corresponding":false},{"id":730421,"name":"Oanh Nguyen","orcid":"0000-0003-1492-2084","position":2,"is_corresponding":false},{"id":730420,"name":"Mohan P A Marimuthu","orcid":"0000-0001-6121-3286","position":3,"is_corresponding":false},{"id":911241,"name":"Noravit Chumchim","orcid":"0000-0003-1139-9427","position":4,"is_corresponding":false},{"id":908281,"name":"Colin W Fairbairn","orcid":"0009-0007-8129-0095","position":5,"is_corresponding":false},{"id":24610,"name":"William Seligmann","orcid":"0000-0002-5762-3095","position":6,"is_corresponding":false},{"id":911242,"name":"Eric Beraut","orcid":"0000-0002-4443-6282","position":7,"is_corresponding":false},{"id":1467739,"name":"Christopher J Conroy","orcid":null,"position":8,"is_corresponding":false},{"id":993313,"name":"James L. Patton","orcid":"0000-0002-8709-7196","position":9,"is_corresponding":false},{"id":1108317,"name":"Rauri C. K. Bowie","orcid":"0000-0001-8328-6021","position":10,"is_corresponding":false},{"id":350939,"name":"Michael W. Nachman","orcid":"0000-0003-4321-5135","position":11,"is_corresponding":false},{"id":820818,"name":"Erin Voss","orcid":"0000-0003-2662-1190","position":0,"is_corresponding":true}],"reference_count":61,"raw_metadata":null,"created_at":"2026-07-19T02:56:20.933088Z","pmid":"40243028","pmcid":"PMC12277566","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":83.3333,"fair_a":68.75,"fair_i":20.0,"fair_r":41.6667,"fair_zscore":1.2752,"fair_rationale":{"fair_score":66.67,"has_llm":true,"taxonomy_version":"fair_taxonomy_v5","dimensions":{"F":{"name":"Findable","score":83.33,"criteria":[{"key":"f_dataset_pid","label":"Persistent identifier for the data","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).","grounded":true,"rationale":"The paper provides BioProject accessions (PRJNA851460, PRJNA851459) which are persistent identifiers registered in identifiers.org. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit","RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier'","FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'"],"scored":true,"signal":null},{"key":"f_repository_named","label":"Named repository","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Raw sequencing data for sample MVZ:Mamm:240054 (NCBI BioSample SAMN29046532) are deposited in the NCBI Short Read Archive (SRA) under accessions SRX17304138 - SRX17304140.","grounded":true,"rationale":"The paper names NCBI Short Read Archive and Dryad Data repository as holders of the data. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed (","NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived","NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten"],"scored":true,"signal":null},{"key":"f_data_availability_statement","label":"Data-availability statement","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate). Raw sequencing data for sample MVZ:Mamm:240054 (NCBI BioSample SAMN29046532) are deposited in the NCBI Short Read Archive (SRA) under accessions SRX17304138 - SRX17304140. Assembly scripts and other data for the analyses presented can be found at the following GitHub repository: www.github.com/ccgproject/ccgp_assembly . Preliminary annotation and mitochondrial genome sequence are available on the Dryad Data repository at https://doi.org/10.5061/dryad.x0k6djhtc .","grounded":true,"rationale":"The data-availability statement points to repositories (NCBI, Dryad) with accessions and DOIs, meeting Colavizza category 3. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li","Springer Nature research data policy — Data Availability Statements: standard statement templat","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes"],"scored":false,"signal":null},{"key":"f_discovery_metadata","label":"Description of the dataset as an object","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"Table 1. Metrics for the primary and alternate assemblies of Merriam’s kangaroo rat (Dipodomys merriami) genome.","grounded":false,"rationale":"The paper includes an itemised inventory of assembly metrics in a table, providing structured description of the dataset. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]","anchors":["RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential)","FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability'","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'"],"scored":false,"signal":null},{"key":"f_dataset_cited","label":"Dataset formally cited","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).","grounded":true,"rationale":"The dataset identifiers appear only in the body text (Data Availability Statement), not in a reference-list entry. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first-","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes","FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'"],"scored":true,"signal":null}]},"A":{"name":"Accessible","score":68.75,"criteria":[{"key":"a_data_openly_accessible","label":"Access route free of preconditions","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).","grounded":true,"rationale":"The data are stated to be available in public repositories without any precondition such as embargo, registration, or request. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-A1.1-01D — 'Data is accessible through a free access protocol'","FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'"],"scored":true,"signal":null},{"key":"a_access_conditions_stated","label":"Access level labelled","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).","grounded":true,"rationale":"The paper does not explicitly label the access level, but describes the action of accessing the data via repository identifiers, allowing inference of open access. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","RDA-A1-01M — metadata contains information to enable the user to get access to the data","COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl"],"scored":false,"signal":null},{"key":"a_controlled_access_for_sensitive","label":"Gatekeeper for sensitive data","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":"Raw sequencing data for sample MVZ:Mamm:240054 (NCBI BioSample SAMN29046532) are deposited in the NCBI Short Read Archive (SRA) under accessions SRX17304138 - SRX17304140.","grounded":true,"rationale":"The data are non-sensitive rodent genome sequences, and no gatekeeper is named; the data are openly deposited.","anchors":["NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee","RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and ","NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse"],"scored":false,"signal":null},{"key":"a_timeline_retention","label":"Availability timing & retention","kind":"llm","weight":0.5,"fraction":0.5,"verdict":"partial","evidence":"Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).","grounded":true,"rationale":"The paper states the data are available in the present tense but gives no persistence commitment. [majority verdict 'partial' (3/5 passes agreed)]","anchors":["NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines","NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy'","RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'"],"scored":false,"signal":null}]},"I":{"name":"Interoperable","score":20.0,"criteria":[{"key":"i_open_nonproprietary_format","label":"Open file format","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The paper does not name the file format of the released data (e.g., FASTA, FASTQ).","anchors":["FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co","RDA-R1.3-02D — data is expressed in a machine-understandable community standard","RDA-I1-01D — data uses a knowledge representation expressed in a standardised format"],"scored":true,"signal":null},{"key":"i_community_standard_vocabulary","label":"Community standard / vocabulary","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No data or metadata community standard from the specified list (e.g., MIAME, BIDS, ontologies) is named and applied to the data; BUSCO is a tool, not a standard format or schema.","anchors":["RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential)","RDA-R1.3-01D — 'Data complies with a community standard'","RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'"],"scored":false,"signal":null},{"key":"i_qualified_references","label":"Identifiers for the resources the data depend on","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"We used liftoff ( Shumate and Salzberg 2021 ) to lift over gene coding sequences, exons, and mRNAs from D. spectabilis (NCBI: GCF_019054845.1) to D. merriami","grounded":true,"rationale":"The paper provides an identifier (GCF_019054845.1) for a third-party resource (D. spectabilis genome) used in the analysis. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-I3-01M — '(meta)data include references to other (meta)data'","RDA-I3-03M — 'metadata includes qualified references to other metadata'","FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'"],"scored":false,"signal":null}]},"R":{"name":"Reusable","score":41.67,"criteria":[{"key":"r_reuse_license","label":"Reuse licence","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No reuse licence is stated for the data; the CC BY-NC 4.0 licence applies only to the article text.","anchors":["RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu","RDA-R1.1-02M — 'Metadata refers to a standard reuse licence'","RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'"],"scored":true,"signal":null},{"key":"r_provenance_methods","label":"Provenance of the data","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"High molecular weight (HMW) genomic DNA (gDNA) was extracted from 78 mg of liver tissue (male, MVZ:Mamm:240054, JLP29074) using the Nanobind Tissue Big DNA kit as per the manufacturer’s instructions (Pacific BioSciences—PacBio, Menlo Park, CA).","grounded":true,"rationale":"The paper names specific instruments, kits, and software used to produce the data (e.g., Nanobind Tissue Big DNA kit, PacBio Sequel II, Illumina NovaSeq 6000). [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa","FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati","W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance"],"scored":false,"signal":null},{"key":"r_documentation_codebook","label":"Documentation / codebook","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The paper does not mention a README, data dictionary, or codebook that accompanies the data. [majority verdict 'no' (4/5 passes agreed)]","anchors":["RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'","NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t"],"scored":false,"signal":null},{"key":"r_versioning","label":"Snapshot identified","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"Neither a version token nor a date is provided for the deposited data; the BioProject IDs and DOI are not versioned. [majority verdict 'no' (3/5 passes agreed)]","anchors":["DataCite Metadata Schema 4.6 — the 'Version' property","RDA-R1.2-01M — provenance information (which version was used is provenance)","NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'"],"scored":true,"signal":null},{"key":"x_code_availability","label":"Analysis code available","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"Assembly scripts and other data for the analyses presented can be found at the following GitHub repository: www.github.com/ccgproject/ccgp_assembly .","grounded":true,"rationale":"The paper provides a machine-resolvable code-forge URL for the study's own code.","anchors":["NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code'","FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear","FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)"],"scored":true,"signal":null},{"key":"x_funding_attribution","label":"Funder and award number","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"This work was supported by the California Conservation Genomics Project, with funding provided to the University of California by the State of California, State Budget Act of 2019 [UC Award ID RSI-19-690224].","grounded":true,"rationale":"The paper includes an award/grant number (RSI-19-690224) attached to a named funder.","anchors":["DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award","Crossref Funder Registry — canonical funder identifiers for funding metadata","RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco"],"scored":true,"signal":null}]}},"actions":[{"key":"r_reuse_license","dimension":"R","label":"Reuse licence","action":"Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No reuse licence is stated for the data; the CC BY-NC 4.0 licence applies only to the article text.","gain":16.67,"priority":"essential","scored":true},{"key":"i_open_nonproprietary_format","dimension":"I","label":"Open file format","action":"Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable. Prefer open genomics / sequencing formats such as FASTQ, BAM or VCF.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The paper does not name the file format of the released data (e.g., FASTA, FASTQ).","gain":8.33,"priority":"important","scored":true},{"key":"f_dataset_cited","dimension":"F","label":"Dataset formally cited","action":"Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the genomics / sequencing repository accession (e.g. from GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA)) in the reference list.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).","why":"The dataset identifiers appear only in the body text (Data Availability Statement), not in a reference-list entry. [majority verdict 'partial' (4/5 passes agreed)]","gain":4.17,"priority":"important","scored":true},{"key":"r_versioning","dimension":"R","label":"Snapshot identified","action":"Version the deposit and cite the exact version analysed (a version-specific DOI, or an accession with its version suffix). A reader reproducing your work against 'the current release' is reproducing it against a different dataset.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"Neither a version token nor a date is provided for the deposited data; the BioProject IDs and DOI are not versioned. [majority verdict 'no' (3/5 passes agreed)]","gain":4.17,"priority":"useful","scored":true},{"key":"f_discovery_metadata","dimension":"F","label":"Description of the dataset as an object","action":"Add a 'Data Records' section: itemise every file in the deposit and every variable or sample it holds, with counts and units. Describe the dataset as an object in its own right, not as a by-product of the findings — this is what makes it discoverable to someone who is not looking for your paper.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Table 1. Metrics for the primary and alternate assemblies of Merriam’s kangaroo rat (Dipodomys merriami) genome.","why":"The paper includes an itemised inventory of assembly metrics in a table, providing structured description of the dataset. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]","gain":0.0,"priority":"essential","scored":false},{"key":"a_access_conditions_stated","dimension":"A","label":"Access level labelled","action":"State the access level in words, using the standard vocabulary: 'These data are open access' / 'These data are controlled access'. A reader — and a harvester — should not have to infer the access level from the presence of a download link.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).","why":"The paper does not explicitly label the access level, but describes the action of accessing the data via repository identifiers, allowing inference of open access. [majority verdict 'partial' (4/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"i_community_standard_vocabulary","dimension":"I","label":"Community standard / vocabulary","action":"Adopt and NAME your domain's data standard — the minimum-information checklist, metadata schema, or ontology your community uses (MIAME/MINSEQE, ISA-Tab, BIDS, an OBO ontology, HL7 FHIR/OMOP) — and say which one you followed. A reporting checklist standardises your paper; it does nothing for your data. In genomics / sequencing, describe the data with MIAME, MINSEQE or MIxS.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No data or metadata community standard from the specified list (e.g., MIAME, BIDS, ontologies) is named and applied to the data; BUSCO is a tool, not a standard format or schema.","gain":0.0,"priority":"important","scored":false},{"key":"r_documentation_codebook","dimension":"R","label":"Documentation / codebook","action":"Ship a README and a data dictionary IN the deposit — every file, every variable, its units, its allowed values, its missing-value codes. It is the cheapest single thing that makes a dataset usable by someone who was not in the lab, and a table buried in the article does not travel with the data.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The paper does not mention a README, data dictionary, or codebook that accompanies the data. [majority verdict 'no' (4/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"a_controlled_access_for_sensitive","dimension":"A","label":"Gatekeeper for sensitive data","action":"Route sensitive data through an institutional gatekeeper — deposit in a controlled- access repository (dbGaP, EGA) with a Data Access Committee and a published DUA — rather than through the corresponding author's inbox. An author-gated dataset dies with the author's email address, and 'on reasonable request' has been shown repeatedly not to yield data. For sensitive/human genomics / sequencing data, use a controlled-access repository such as dbGaP or EGA.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"Raw sequencing data for sample MVZ:Mamm:240054 (NCBI BioSample SAMN29046532) are deposited in the NCBI Short Read Archive (SRA) under accessions SRX17304138 - SRX17304140.","why":"The data are non-sensitive rodent genome sequences, and no gatekeeper is named; the data are openly deposited.","gain":0.0,"priority":"useful","scored":false},{"key":"a_timeline_retention","dimension":"A","label":"Availability timing & retention","action":"State when the data become available AND how long they will be retained — cite the repository's preservation policy. NIH DMS Element 4 asks for both; most papers give neither.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Genome assemblies generated for this study are available under NCBI BioProject IDs PRJNA851460 (principal) and PRJNA851459 (alternate).","why":"The paper states the data are available in the present tense but gives no persistence commitment. [majority verdict 'partial' (3/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false}],"suggestions":["Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable. Prefer open genomics / sequencing formats such as FASTQ, BAM or VCF.","Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the genomics / sequencing repository accession (e.g. from GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA)) in the reference list.","Version the deposit and cite the exact version analysed (a version-specific DOI, or an accession with its version suffix). A reader reproducing your work against 'the current release' is reproducing it against a different dataset.","Add a 'Data Records' section: itemise every file in the deposit and every variable or sample it holds, with counts and units. Describe the dataset as an object in its own right, not as a by-product of the findings — this is what makes it discoverable to someone who is not looking for your paper."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v8","fulltext_source":"epmc_xml"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v8","fair_fulltext_source":"epmc_xml","fair_has_llm":true,"fair_computed_at":"2026-07-20T14:03:30.341850Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}