{"doi":"10.12688/f1000research.134798.2","title":"Facilitating accessible, rapid, and appropriate processing of ancient metagenomic data with AMDirT","abstract":"Background: Access to sample-level metadata is important when selecting public metagenomic sequencing datasets for reuse in new biological analyses. The Standards, Precautions, and Advances in Ancient Metagenomics community (SPAAM, https://spaam-community.org) has previously published AncientMetagenomeDir, a collection of curated and standardised sample metadata tables for metagenomic and microbial genome datasets generated from ancient samples. However, while sample-level information is useful for identifying relevant samples for inclusion in new projects, Next Generation Sequencing (NGS) library construction and sequencing metadata are also essential for appropriately reprocessing ancient metagenomic data. Currently, recovering information for downloading and preparing such data is difficult when laboratory and bioinformatic metadata is heterogeneously recorded in prose-based publications. Methods: Through a series of community-based hackathon events, AncientMetagenomeDir was updated to provide standardised library-level metadata of existing and new ancient metagenomic samples. In tandem, the companion tool 'AMDirT' was developed to facilitate rapid data filtering and downloading of ancient metagenomic data, as well as improving automated metadata curation and validation for AncientMetagenomeDir. Results: AncientMetagenomeDir was extended to include standardised metadata of over 6000 ancient metagenomic libraries. The companion tool 'AMDirT' provides both graphical- and command-line interface based access to such metadata for users from a wide range of computational backgrounds. We also report on errors with metadata reporting that appear to commonly occur during data upload and provide suggestions on how to improve the quality of data sharing by the community. Conclusions: Together, both standardised metadata reporting and tooling will help towards easier incorporation and reuse of public ancient metagenomic datasets into future analyses.","journal":"F1000Research","year":2024,"id":454556,"datarank":0.37442627271832657,"base_score":2.0794415416798357,"endowment":2.0794415416798357,"self_citation_contribution":0.31191623125197543,"citation_network_contribution":0.06251004146635115,"self_endowment_contribution":0.31191623125197543,"citer_contribution":0.06251004146635115,"corpus_percentile":51.99195482323818,"corpus_rank":6207,"citation_count":7,"citer_count":4,"citers_with_citation_signal":2,"citers_with_endowment":2,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9396,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2024-01-01","fair_score":100.0,"fair_percentile":99.93885661877101,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1121722,"name":"Adrian Forsythe","orcid":"0000-0003-1966-9946","position":1,"is_corresponding":false},{"id":558851,"name":"Aida Andrades Valtueña","orcid":"0000-0002-1737-2228","position":2,"is_corresponding":false},{"id":326198,"name":"Alexander Hübner","orcid":"0000-0003-3572-9996","position":3,"is_corresponding":false},{"id":1121723,"name":"Anan Ibrahim","orcid":"0000-0003-3719-901X","position":4,"is_corresponding":false},{"id":1121724,"name":"Andrea Quagliariello","orcid":"0000-0002-3941-3684","position":5,"is_corresponding":false},{"id":1121725,"name":"Anna E. White","orcid":"0000-0001-6922-7885","position":6,"is_corresponding":false},{"id":685041,"name":"Arthur Kocher","orcid":"0000-0002-9499-6472","position":7,"is_corresponding":false},{"id":558852,"name":"Åshild J. Vågene‬","orcid":"0000-0002-7478-8297","position":8,"is_corresponding":false},{"id":1121726,"name":"Bjørn Peare Bartholdy","orcid":"0000-0003-3985-1016","position":9,"is_corresponding":false},{"id":1122138,"name":"Diāna Spurīte","orcid":null,"position":10,"is_corresponding":false},{"id":1121727,"name":"Gabriel Yaxal Ponce‐Soto","orcid":"0000-0002-0605-0231","position":11,"is_corresponding":false},{"id":1121728,"name":"Gunnar U. Neumann","orcid":"0000-0003-3825-8536","position":12,"is_corresponding":false},{"id":1121729,"name":"I-Ting Huang","orcid":"0000-0001-9280-4716","position":13,"is_corresponding":false},{"id":1122139,"name":"Ian Light","orcid":null,"position":14,"is_corresponding":false},{"id":558854,"name":"Irina M. Velsko","orcid":"0000-0001-9810-9917","position":15,"is_corresponding":false},{"id":1121730,"name":"Iseult Jackson","orcid":"0000-0001-6572-8244","position":16,"is_corresponding":false},{"id":1121731,"name":"Jasmin Frangenberg","orcid":"0009-0004-5961-4709","position":17,"is_corresponding":false},{"id":1121732,"name":"Javier G. Serrano","orcid":"0000-0001-5993-8536","position":18,"is_corresponding":false},{"id":1121733,"name":"Julien Fumey","orcid":"0000-0002-6272-5619","position":19,"is_corresponding":false},{"id":255459,"name":"Kadir Toykan Özdoğan","orcid":"0000-0002-9508-8193","position":20,"is_corresponding":false},{"id":1121734,"name":"Kelly E. Blevins","orcid":"0000-0002-5740-639X","position":21,"is_corresponding":false},{"id":1121735,"name":"Kevin G. Daly","orcid":"0000-0002-5579-6144","position":22,"is_corresponding":false},{"id":1121736,"name":"Maria Lopopolo","orcid":"0000-0001-8502-9264","position":23,"is_corresponding":false},{"id":1121737,"name":"Markella Moraitou","orcid":"0000-0003-0860-5920","position":24,"is_corresponding":false},{"id":247946,"name":"Megan Michel","orcid":"0000-0002-5484-7974","position":25,"is_corresponding":false},{"id":1121738,"name":"Meriam van Os","orcid":"0009-0008-9835-8874","position":26,"is_corresponding":false},{"id":558855,"name":"Miriam Bravo-Lopez","orcid":"0000-0002-9627-617X","position":27,"is_corresponding":false},{"id":660046,"name":"Mohamed S. Sarhan","orcid":"0000-0003-0904-976X","position":28,"is_corresponding":false},{"id":1121739,"name":"Nihan Dilşad Dağtaş","orcid":"0000-0003-2658-7295","position":29,"is_corresponding":false},{"id":255374,"name":"Nikolay Oskolkov","orcid":"0000-0001-5326-8893","position":30,"is_corresponding":false},{"id":87236,"name":"Olivia S. Smith","orcid":"0000-0001-5435-2982","position":31,"is_corresponding":false},{"id":863663,"name":"Ophélie Lebrasseur","orcid":"0000-0003-0687-8538","position":32,"is_corresponding":false},{"id":1121740,"name":"Piotr Rozwalak","orcid":"0000-0002-4014-6624","position":33,"is_corresponding":false},{"id":1121741,"name":"Raphael Eisenhofer","orcid":"0000-0002-3843-0749","position":34,"is_corresponding":false},{"id":1121742,"name":"Sally Wasef","orcid":"0000-0002-7207-7395","position":35,"is_corresponding":false},{"id":558858,"name":"Shreya L. Ramachandran","orcid":"0000-0001-6760-5053","position":36,"is_corresponding":false},{"id":1121743,"name":"Valentina Vanghi","orcid":"0000-0002-9412-6602","position":37,"is_corresponding":false},{"id":326199,"name":"Christina Warinner","orcid":"0000-0002-4528-5877","position":38,"is_corresponding":false},{"id":558850,"name":"James A. Fellows Yates","orcid":"0000-0001-5585-6277","position":39,"is_corresponding":false},{"id":326183,"name":"Maxime Borry","orcid":"0000-0001-9140-7559","position":0,"is_corresponding":true}],"reference_count":26,"raw_metadata":null,"created_at":"2026-07-19T02:03:12.720997Z","pmid":"39262445","pmcid":"PMC11387932","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":88.8889,"fair_a":75.0,"fair_i":40.0,"fair_r":100.0,"fair_zscore":2.5946,"fair_rationale":{"fair_score":100.0,"has_llm":true,"taxonomy_version":"fair_taxonomy_v5","dimensions":{"F":{"name":"Findable","score":88.89,"criteria":[{"key":"f_dataset_pid","label":"Persistent identifier for the data","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Each release is archived on Zenodo: https://doi.org/10.5281/zenodo.3980833","grounded":true,"rationale":"The paper provides a DOI for the dataset, which is a persistent identifier.","anchors":["RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit","RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier'","FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'"],"scored":true,"signal":null},{"key":"f_repository_named","label":"Named repository","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Each release is archived on Zenodo","grounded":true,"rationale":"Zenodo is a named data repository in the yes list.","anchors":["RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed (","NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived","NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten"],"scored":true,"signal":null},{"key":"f_data_availability_statement","label":"Data-availability statement","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"The existing sample-level and new library-level data is stored on GitHub: https://github.com/SPAAM-community/AncientMetagenomeDir Each release is archived on Zenodo: https://doi.org/10.5281/zenodo.3980833.","grounded":true,"rationale":"The statement points to a repository record with a DOI, qualifying as Colavizza category 3. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li","Springer Nature research data policy — Data Availability Statements: standard statement templat","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes"],"scored":false,"signal":null},{"key":"f_discovery_metadata","label":"Description of the dataset as an object","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"Newly added library information columns include the library name (how data are typically reported in original publications), the aDNA library generation method (e.g., double-stranded or single-stranded libraries), the library indexing polymerase (e.g., proof-reading or non-proofreading), and the library pretreatment method (e.g., non-Uracil-DNA Glycosylase (UDG), full-UDG, or half-UDG treatments).","grounded":true,"rationale":"The dataset content is described in running prose, not in a section, table, or enumerated list.","anchors":["RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential)","FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability'","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'"],"scored":false,"signal":null},{"key":"f_dataset_cited","label":"Dataset formally cited","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"Fellows Yates JA, Andrades Valtueña A, Vågene ÅJ, et al.: SPAAM-community/AncientMetagenomeDir: v23.03.0: Rocky necropolis of pantalica. March 2023. Reference Source","grounded":true,"rationale":"The dataset appears as a reference-list entry in the bibliography. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first-","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes","FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'"],"scored":true,"signal":null}]},"A":{"name":"Accessible","score":75.0,"criteria":[{"key":"a_data_openly_accessible","label":"Access route free of preconditions","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Data are available under the terms of the Creative Commons Attribution 4.0 International license (CC-BY 4.0).","grounded":true,"rationale":"The data are stated to be available under an open license with no precondition.","anchors":["RDA-A1.1-01D — 'Data is accessible through a free access protocol'","FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'"],"scored":true,"signal":null},{"key":"a_access_conditions_stated","label":"Access level labelled","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"Data are available under the terms of the Creative Commons Attribution 4.0 International license (CC-BY 4.0).","grounded":true,"rationale":"The paper explicitly labels the data as available under a CC-BY 4.0 license, which is an open access label. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","RDA-A1-01M — metadata contains information to enable the user to get access to the data","COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl"],"scored":false,"signal":null},{"key":"a_controlled_access_for_sensitive","label":"Gatekeeper for sensitive data","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The data are not sensitive (ancient metagenomic data), so no gatekeeper is named.","anchors":["NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee","RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and ","NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse"],"scored":false,"signal":null},{"key":"a_timeline_retention","label":"Availability timing & retention","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No persistence commitment or availability timing is stated for the data. [majority verdict 'no' (3/5 passes agreed)]","anchors":["NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines","NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy'","RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'"],"scored":false,"signal":null}]},"I":{"name":"Interoperable","score":40.0,"criteria":[{"key":"i_open_nonproprietary_format","label":"Open file format","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"we created new tab-separated value (TSV) tables","grounded":true,"rationale":"TSV is an open, community-standard format.","anchors":["FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co","RDA-R1.3-02D — data is expressed in a machine-understandable community standard","RDA-I1-01D — data uses a knowledge representation expressed in a standardised format"],"scored":true,"signal":null},{"key":"i_community_standard_vocabulary","label":"Community standard / vocabulary","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No data or metadata community standard is named in the text. [majority verdict 'no' (4/5 passes agreed)]","anchors":["RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential)","RDA-R1.3-01D — 'Data complies with a community standard'","RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'"],"scored":false,"signal":null},{"key":"i_qualified_references","label":"Identifiers for the resources the data depend on","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":"Archived source code at time of publication revision (AMDirT v1.6): https://doi.org/10.5281/zenodo.10941007","grounded":false,"rationale":"The paper provides a DOI for the archived source code, a resource other than the study's own dataset. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (4/5 passes agreed)]","anchors":["RDA-I3-01M — '(meta)data include references to other (meta)data'","RDA-I3-03M — 'metadata includes qualified references to other metadata'","FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'"],"scored":false,"signal":null}]},"R":{"name":"Reusable","score":100.0,"criteria":[{"key":"r_reuse_license","label":"Reuse licence","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Data are available under the terms of the Creative Commons Attribution 4.0 International license (CC-BY 4.0).","grounded":true,"rationale":"CC-BY 4.0 is an open standard license.","anchors":["RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu","RDA-R1.1-02M — 'Metadata refers to a standard reuse licence'","RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'"],"scored":true,"signal":null},{"key":"r_provenance_methods","label":"Provenance of the data","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"The autofill command uses the ENA portal API (https://www.ebi.ac.uk/ena/portal/api/)14 to automatically query and return metadata","grounded":true,"rationale":"The paper names specific tools and APIs used to produce the data. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa","FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati","W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance"],"scored":false,"signal":null},{"key":"r_documentation_codebook","label":"Documentation / codebook","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"we created new tab-separated value (TSV) tables and their associated validation checks in the form of JSON schema files","grounded":true,"rationale":"JSON schema files are documentation objects that accompany the data. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'","NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t"],"scored":false,"signal":null},{"key":"r_versioning","label":"Snapshot identified","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"The version of the dataset used for the demonstration of AMDirT, statistics, and figures in this updated manuscript is v24.03 (Monticello).","grounded":true,"rationale":"A version token (v24.03) is given for the dataset.","anchors":["DataCite Metadata Schema 4.6 — the 'Version' property","RDA-R1.2-01M — provenance information (which version was used is provenance)","NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'"],"scored":true,"signal":null},{"key":"x_code_availability","label":"Analysis code available","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"Source code available from: https://github.com/SPAAM-community/AMDirT","grounded":true,"rationale":"A machine-resolvable code repository URL is given. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code'","FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear","FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)"],"scored":true,"signal":null},{"key":"x_funding_attribution","label":"Funder and award number","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"European Research Council under the European Union's Horizon 2020 research and innovation programme (grant agreement number 804884-DAIRYCULTURES awarded to C.W.)","grounded":true,"rationale":"A specific grant number is attached to a named funder.","anchors":["DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award","Crossref Funder Registry — canonical funder identifiers for funding metadata","RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco"],"scored":true,"signal":null}]}},"actions":[{"key":"f_discovery_metadata","dimension":"F","label":"Description of the dataset as an object","action":"Add a 'Data Records' section: itemise every file in the deposit and every variable or sample it holds, with counts and units. Describe the dataset as an object in its own right, not as a by-product of the findings — this is what makes it discoverable to someone who is not looking for your paper.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Newly added library information columns include the library name (how data are typically reported in original publications), the aDNA library generation method (e.g., double-stranded or single-stranded libraries), the library indexing polymerase (e.g., proof-reading or non-proofreading), and the library pretreatment method (e.g., non-Uracil-DNA Glycosylase (UDG), full-UDG, or half-UDG treatments).","why":"The dataset content is described in running prose, not in a section, table, or enumerated list.","gain":0.0,"priority":"essential","scored":false},{"key":"i_community_standard_vocabulary","dimension":"I","label":"Community standard / vocabulary","action":"Adopt and NAME your domain's data standard — the minimum-information checklist, metadata schema, or ontology your community uses (MIAME/MINSEQE, ISA-Tab, BIDS, an OBO ontology, HL7 FHIR/OMOP) — and say which one you followed. A reporting checklist standardises your paper; it does nothing for your data. In neuroimaging, describe the data with BIDS, NIfTI or DICOM.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No data or metadata community standard is named in the text. [majority verdict 'no' (4/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"a_controlled_access_for_sensitive","dimension":"A","label":"Gatekeeper for sensitive data","action":"Route sensitive data through an institutional gatekeeper — deposit in a controlled- access repository (dbGaP, EGA) with a Data Access Committee and a published DUA — rather than through the corresponding author's inbox. An author-gated dataset dies with the author's email address, and 'on reasonable request' has been shown repeatedly not to yield data.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The data are not sensitive (ancient metagenomic data), so no gatekeeper is named.","gain":0.0,"priority":"useful","scored":false},{"key":"i_qualified_references","dimension":"I","label":"Identifiers for the resources the data depend on","action":"Cite by identifier every resource the data depend on — the source datasets' accessions, the reference build (GRCh38 / GCA_000001405.28), the cohort application number, the code DOI — and register those relations on the dataset record (IsDerivedFrom, IsSupplementTo). A name is not a link: it cannot be resolved, versioned, or followed by a machine.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"Archived source code at time of publication revision (AMDirT v1.6): https://doi.org/10.5281/zenodo.10941007","why":"The paper provides a DOI for the archived source code, a resource other than the study's own dataset. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (4/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false},{"key":"a_timeline_retention","dimension":"A","label":"Availability timing & retention","action":"State when the data become available AND how long they will be retained — cite the repository's preservation policy. NIH DMS Element 4 asks for both; most papers give neither.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No persistence commitment or availability timing is stated for the data. [majority verdict 'no' (3/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false}],"suggestions":["Add a 'Data Records' section: itemise every file in the deposit and every variable or sample it holds, with counts and units. Describe the dataset as an object in its own right, not as a by-product of the findings — this is what makes it discoverable to someone who is not looking for your paper.","Adopt and NAME your domain's data standard — the minimum-information checklist, metadata schema, or ontology your community uses (MIAME/MINSEQE, ISA-Tab, BIDS, an OBO ontology, HL7 FHIR/OMOP) — and say which one you followed. A reporting checklist standardises your paper; it does nothing for your data. In neuroimaging, describe the data with BIDS, NIfTI or DICOM.","Route sensitive data through an institutional gatekeeper — deposit in a controlled- access repository (dbGaP, EGA) with a Data Access Committee and a published DUA — rather than through the corresponding author's inbox. An author-gated dataset dies with the author's email address, and 'on reasonable request' has been shown repeatedly not to yield data.","Cite by identifier every resource the data depend on — the source datasets' accessions, the reference build (GRCh38 / GCA_000001405.28), the cohort application number, the code DOI — and register those relations on the dataset record (IsDerivedFrom, IsSupplementTo). A name is not a link: it cannot be resolved, versioned, or followed by a machine.","State when the data become available AND how long they will be retained — cite the repository's preservation policy. NIH DMS Element 4 asks for both; most papers give neither."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v8","fulltext_source":"unpaywall_pdf"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v8","fair_fulltext_source":"unpaywall_pdf","fair_has_llm":true,"fair_computed_at":"2026-07-20T12:41:37.721489Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}