{"doi":"10.1038/s41586-021-03420-7","title":"The structure, function and evolution of a complete human chromosome 8","abstract":". Here we use complementary long-read sequencing technologies to complete the linear assembly of human chromosome 8. Our assembly resolves the sequence of five previously long-standing gaps, including a 2.08-Mb centromeric α-satellite array, a 644-kb copy number polymorphism in the β-defensin gene cluster that is important for disease risk, and an 863-kb variable number tandem repeat at chromosome 8q21.2 that can function as a neocentromere. We show that the centromeric α-satellite array is generally methylated except for a 73-kb hypomethylated region of diverse higher-order α-satellites enriched with CENP-A nucleosomes, consistent with the location of the kinetochore. In addition, we confirm the overall organization and methylation pattern of the centromere in a diploid human genome. Using a dual long-read sequencing approach, we complete high-quality draft assemblies of the orthologous centromere from chromosome 8 in chimpanzee, orangutan and macaque to reconstruct its evolutionary history. Comparative and phylogenetic analyses show that the higher-order α-satellite structure evolved in the great ape ancestor with a layered symmetry, in which more ancient higher-order repeats locate peripherally to monomeric α-satellites. We estimate that the mutation rate of centromeric satellite DNA is accelerated by more than 2.2-fold compared to the unique portions of the genome, and this acceleration extends into the flanking sequence.","journal":"Nature","year":2021,"id":145506,"datarank":5.586047697733786,"base_score":5.963579343618446,"endowment":5.963579343618446,"self_citation_contribution":0.894536901542767,"citation_network_contribution":4.691510796191019,"self_endowment_contribution":0.894536901542767,"citer_contribution":4.691510796191019,"corpus_percentile":96.35646321652355,"corpus_rank":472,"citation_count":388,"citer_count":100,"citers_with_citation_signal":100,"citers_with_endowment":100,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9003,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":66.6667,"fair_percentile":86.48731274839498,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":24561,"name":"Mitchell R. Vollger","orcid":"0000-0002-8651-1615","position":1,"is_corresponding":false},{"id":289887,"name":"PingHsun Hsieh","orcid":"0000-0001-8294-6227","position":2,"is_corresponding":false},{"id":51349,"name":"Yafei Mao","orcid":"0000-0002-9648-4278","position":3,"is_corresponding":false},{"id":388911,"name":"Mikhail Liskovykh","orcid":"0000-0002-6760-8331","position":4,"is_corresponding":false},{"id":2118,"name":"Sergey Koren","orcid":"0000-0002-1472-8962","position":5,"is_corresponding":false},{"id":24611,"name":"Sergey Nurk","orcid":"0000-0003-1301-5749","position":6,"is_corresponding":false},{"id":246457,"name":"Ludovica Mercuri","orcid":"0000-0002-3688-4501","position":7,"is_corresponding":false},{"id":49020,"name":"Philip C. Dishuck","orcid":"0000-0003-2223-9787","position":8,"is_corresponding":false},{"id":21320,"name":"Arang Rhie","orcid":"0000-0002-9809-8127","position":9,"is_corresponding":false},{"id":30860,"name":"Leonardo Gomes de Lima","orcid":"0000-0001-6340-6065","position":10,"is_corresponding":false},{"id":49021,"name":"Tatiana Dvorkina","orcid":"0000-0003-2253-1702","position":11,"is_corresponding":false},{"id":19696,"name":"David Porubsky","orcid":"0000-0001-8414-8966","position":12,"is_corresponding":false},{"id":19702,"name":"William T. Harvey","orcid":"0000-0003-0646-7528","position":13,"is_corresponding":false},{"id":49010,"name":"Alla Mikheenko","orcid":"0000-0003-3400-9719","position":14,"is_corresponding":false},{"id":49009,"name":"Andrey V. Bzikadze","orcid":"0000-0002-7928-7950","position":15,"is_corresponding":false},{"id":2103,"name":"Milinn Kremitzki","orcid":"0000-0001-7980-3153","position":16,"is_corresponding":false},{"id":2128,"name":"Tina A. Graves-Lindsay","orcid":"0000-0002-0409-891X","position":17,"is_corresponding":false},{"id":241559,"name":"Chirag Jain","orcid":"0000-0002-4300-0794","position":18,"is_corresponding":false},{"id":19716,"name":"Kendra Hoekzema","orcid":"0000-0002-8058-0177","position":19,"is_corresponding":false},{"id":246454,"name":"Shwetha C. Murali","orcid":"0009-0008-4414-6614","position":20,"is_corresponding":false},{"id":19706,"name":"Katherine M. Munson","orcid":"0000-0001-8413-6498","position":21,"is_corresponding":false},{"id":327932,"name":"Carl Baker","orcid":null,"position":22,"is_corresponding":false},{"id":109461,"name":"Melanie Sorensen","orcid":"0000-0002-1525-9707","position":23,"is_corresponding":false},{"id":552924,"name":"Alexandra M. Lewis","orcid":null,"position":24,"is_corresponding":false},{"id":49045,"name":"Urvashi Surti","orcid":"0000-0003-4283-9018","position":25,"is_corresponding":false},{"id":108065,"name":"Jennifer L. Gerton","orcid":"0000-0003-0743-3637","position":26,"is_corresponding":false},{"id":330900,"name":"Vladimir Larionov","orcid":"0000-0002-7996-4894","position":27,"is_corresponding":false},{"id":51384,"name":"Mario Ventura","orcid":"0000-0001-7762-8777","position":28,"is_corresponding":false},{"id":108047,"name":"Karen H. Miga","orcid":"0000-0002-3670-4507","position":29,"is_corresponding":false},{"id":2122,"name":"Adam  M. Phillippy","orcid":"0000-0003-2983-8934","position":30,"is_corresponding":false},{"id":2125,"name":"Evan E. Eichler","orcid":"0000-0002-8246-4014","position":31,"is_corresponding":false},{"id":19692,"name":"Glennis A. Logsdon","orcid":"0000-0003-2396-0656","position":0,"is_corresponding":true}],"reference_count":101,"raw_metadata":null,"created_at":"2026-07-18T23:42:12.871665Z","pmid":"33828295","pmcid":"PMC8099727","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":83.3333,"fair_a":75.0,"fair_i":20.0,"fair_r":41.6667,"fair_zscore":1.2752,"fair_rationale":{"fair_score":66.67,"has_llm":true,"taxonomy_version":"fair_taxonomy_v5","dimensions":{"F":{"name":"Findable","score":83.33,"criteria":[{"key":"f_dataset_pid","label":"Persistent identifier for the data","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"complete CHM13 chromosome 8 sequence (PRJNA686384)","grounded":true,"rationale":"The paper provides a BioProject accession (PRJNA686384) which is a persistent identifier scheme.","anchors":["RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit","RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier'","FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'"],"scored":true,"signal":null},{"key":"f_repository_named","label":"Named repository","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"listed in Supplementary Table 9 with their BioProject, accession numbers and/or URL","grounded":true,"rationale":"The paper names BioProject (a repository) as the holder of the data.","anchors":["RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed (","NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived","NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten"],"scored":true,"signal":null},{"key":"f_data_availability_statement","label":"Data-availability statement","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"The complete CHM13 chromosome 8 sequence and all data generated and/or used in this study are publicly available and listed in Supplementary Table 9 with their BioProject, accession numbers and/or URL. For convenience, we also list their BioProjects and/or URLs here: complete CHM13 chromosome 8 sequence (PRJNA686384); CHM13 ONT, Iso-Seq, and CENP-A ChIP-seq data (PRJNA559484); CHM13 Strand-Seq alignments (https://zenodo.org/record/3998125); HG00733 ONT data (PRJNA686388); HG00733 PacBio HiFi data (PRJEB36100); testis and fetal brain Iso-Seq data (PRJNA659539); and NHPs (chimpanzee (Clint; S006007), orangutan (Susie; PR01109), and macaque (AG07107)) ONT and PacBio HiFi data (PRJNA659034). All CHM13 BACs used in this study are listed in Supplementary Table 10 with their accession numbers.","grounded":true,"rationale":"The data availability statement points to repository records with accessions and a persistent link.","anchors":["Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li","Springer Nature research data policy — Data Availability Statements: standard statement templat","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes"],"scored":false,"signal":null},{"key":"f_discovery_metadata","label":"Description of the dataset as an object","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"The complete telomere-to-telomere sequence of human chromosome 8 is 146,259,671 bases long and includes 3,334,256 bases that are missing from the current reference genome (GRCh38).","grounded":true,"rationale":"The dataset's content is described in running prose, not in an itemised inventory. [majority verdict 'partial' (3/5 passes agreed)]","anchors":["RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential)","FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability'","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'"],"scored":false,"signal":null},{"key":"f_dataset_cited","label":"Dataset formally cited","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"complete CHM13 chromosome 8 sequence (PRJNA686384)","grounded":true,"rationale":"The dataset identifiers appear only in the body text (data availability statement), not in the reference list.","anchors":["FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first-","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes","FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'"],"scored":true,"signal":null}]},"A":{"name":"Accessible","score":75.0,"criteria":[{"key":"a_data_openly_accessible","label":"Access route free of preconditions","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"The complete CHM13 chromosome 8 sequence and all data generated and/or used in this study are publicly available","grounded":true,"rationale":"The data are stated to be publicly available with no precondition.","anchors":["RDA-A1.1-01D — 'Data is accessible through a free access protocol'","FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'"],"scored":true,"signal":null},{"key":"a_access_conditions_stated","label":"Access level labelled","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"The complete CHM13 chromosome 8 sequence and all data generated and/or used in this study are publicly available","grounded":true,"rationale":"The paper explicitly labels the data as 'publicly available', which is an access-level label.","anchors":["FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","RDA-A1-01M — metadata contains information to enable the user to get access to the data","COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl"],"scored":false,"signal":null},{"key":"a_controlled_access_for_sensitive","label":"Gatekeeper for sensitive data","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The data are not sensitive and no gatekeeper is named.","anchors":["NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee","RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and ","NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse"],"scored":false,"signal":null},{"key":"a_timeline_retention","label":"Availability timing & retention","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The paper does not state how long the data will persist or when they become available (beyond 'publicly available'). [majority verdict 'no' (4/5 passes agreed)]","anchors":["NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines","NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy'","RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'"],"scored":false,"signal":null}]},"I":{"name":"Interoperable","score":20.0,"criteria":[{"key":"i_open_nonproprietary_format","label":"Open file format","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No file format is named for the released data.","anchors":["FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co","RDA-R1.3-02D — data is expressed in a machine-understandable community standard","RDA-I1-01D — data uses a knowledge representation expressed in a standardised format"],"scored":true,"signal":null},{"key":"i_community_standard_vocabulary","label":"Community standard / vocabulary","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No community data or metadata standard is named for the data.","anchors":["RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential)","RDA-R1.3-01D — 'Data complies with a community standard'","RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'"],"scored":false,"signal":null},{"key":"i_qualified_references","label":"Identifiers for the resources the data depend on","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"CHM13 Illumina data (SRR1997411, SRR3189741, SRR3189742 and SRR3189743)","grounded":true,"rationale":"The paper gives an identifier (SRR3189741) for Illumina data used in the study.","anchors":["RDA-I3-01M — '(meta)data include references to other (meta)data'","RDA-I3-03M — 'metadata includes qualified references to other metadata'","FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'"],"scored":false,"signal":null}]},"R":{"name":"Reusable","score":41.67,"criteria":[{"key":"r_reuse_license","label":"Reuse licence","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No license is stated for the data; the CC-BY license applies to the article only.","anchors":["RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu","RDA-R1.1-02M — 'Metadata refers to a standard reuse licence'","RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'"],"scored":true,"signal":null},{"key":"r_provenance_methods","label":"Provenance of the data","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"Libraries were sequenced on the Sequel II platform (Instrument Control SW v7.1 or v8.0) with three to seven SMRT Cells 8M","grounded":true,"rationale":"The paper names the specific sequencing platform (Sequel II) used to produce the data. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa","FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati","W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance"],"scored":false,"signal":null},{"key":"r_documentation_codebook","label":"Documentation / codebook","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No documentation object (README, codebook) is named as accompanying the data.","anchors":["RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'","NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t"],"scored":false,"signal":null},{"key":"r_versioning","label":"Snapshot identified","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No version token or date is given to identify a specific snapshot of the data.","anchors":["DataCite Metadata Schema 4.6 — the 'Version' property","RDA-R1.2-01M — provenance information (which version was used is provenance)","NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'"],"scored":true,"signal":null},{"key":"x_code_availability","label":"Analysis code available","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"Custom code for the SUNK-based assembly method is available at https://github.com/glogsdon1/sunk-based_assembly","grounded":true,"rationale":"The paper gives a GitHub URL for the custom code, which is a machine-resolvable locator.","anchors":["NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code'","FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear","FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)"],"scored":true,"signal":null},{"key":"x_funding_attribution","label":"Funder and award number","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"HG002385 and HG010169 (E.E.E.)","grounded":true,"rationale":"The paper provides specific grant numbers from the NIH. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award","Crossref Funder Registry — canonical funder identifiers for funding metadata","RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco"],"scored":true,"signal":null}]}},"actions":[{"key":"r_reuse_license","dimension":"R","label":"Reuse licence","action":"Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No license is stated for the data; the CC-BY license applies to the article only.","gain":16.67,"priority":"essential","scored":true},{"key":"i_open_nonproprietary_format","dimension":"I","label":"Open file format","action":"Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable. Prefer open genomics / sequencing formats such as FASTQ, BAM or VCF.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No file format is named for the released data.","gain":8.33,"priority":"important","scored":true},{"key":"f_dataset_cited","dimension":"F","label":"Dataset formally cited","action":"Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the genomics / sequencing repository accession (e.g. from GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA)) in the reference list.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"complete CHM13 chromosome 8 sequence (PRJNA686384)","why":"The dataset identifiers appear only in the body text (data availability statement), not in the reference list.","gain":4.17,"priority":"important","scored":true},{"key":"r_versioning","dimension":"R","label":"Snapshot identified","action":"Version the deposit and cite the exact version analysed (a version-specific DOI, or an accession with its version suffix). A reader reproducing your work against 'the current release' is reproducing it against a different dataset.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No version token or date is given to identify a specific snapshot of the data.","gain":4.17,"priority":"useful","scored":true},{"key":"f_discovery_metadata","dimension":"F","label":"Description of the dataset as an object","action":"Add a 'Data Records' section: itemise every file in the deposit and every variable or sample it holds, with counts and units. Describe the dataset as an object in its own right, not as a by-product of the findings — this is what makes it discoverable to someone who is not looking for your paper.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"The complete telomere-to-telomere sequence of human chromosome 8 is 146,259,671 bases long and includes 3,334,256 bases that are missing from the current reference genome (GRCh38).","why":"The dataset's content is described in running prose, not in an itemised inventory. [majority verdict 'partial' (3/5 passes agreed)]","gain":0.0,"priority":"essential","scored":false},{"key":"i_community_standard_vocabulary","dimension":"I","label":"Community standard / vocabulary","action":"Adopt and NAME your domain's data standard — the minimum-information checklist, metadata schema, or ontology your community uses (MIAME/MINSEQE, ISA-Tab, BIDS, an OBO ontology, HL7 FHIR/OMOP) — and say which one you followed. A reporting checklist standardises your paper; it does nothing for your data. In genomics / sequencing, describe the data with MIAME, MINSEQE or MIxS.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No community data or metadata standard is named for the data.","gain":0.0,"priority":"important","scored":false},{"key":"r_documentation_codebook","dimension":"R","label":"Documentation / codebook","action":"Ship a README and a data dictionary IN the deposit — every file, every variable, its units, its allowed values, its missing-value codes. It is the cheapest single thing that makes a dataset usable by someone who was not in the lab, and a table buried in the article does not travel with the data.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No documentation object (README, codebook) is named as accompanying the data.","gain":0.0,"priority":"important","scored":false},{"key":"a_controlled_access_for_sensitive","dimension":"A","label":"Gatekeeper for sensitive data","action":"Route sensitive data through an institutional gatekeeper — deposit in a controlled- access repository (dbGaP, EGA) with a Data Access Committee and a published DUA — rather than through the corresponding author's inbox. An author-gated dataset dies with the author's email address, and 'on reasonable request' has been shown repeatedly not to yield data. For sensitive/human genomics / sequencing data, use a controlled-access repository such as dbGaP or EGA.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The data are not sensitive and no gatekeeper is named.","gain":0.0,"priority":"useful","scored":false},{"key":"a_timeline_retention","dimension":"A","label":"Availability timing & retention","action":"State when the data become available AND how long they will be retained — cite the repository's preservation policy. NIH DMS Element 4 asks for both; most papers give neither.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The paper does not state how long the data will persist or when they become available (beyond 'publicly available'). [majority verdict 'no' (4/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false}],"suggestions":["Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable. Prefer open genomics / sequencing formats such as FASTQ, BAM or VCF.","Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the genomics / sequencing repository accession (e.g. from GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA)) in the reference list.","Version the deposit and cite the exact version analysed (a version-specific DOI, or an accession with its version suffix). A reader reproducing your work against 'the current release' is reproducing it against a different dataset.","Add a 'Data Records' section: itemise every file in the deposit and every variable or sample it holds, with counts and units. Describe the dataset as an object in its own right, not as a by-product of the findings — this is what makes it discoverable to someone who is not looking for your paper."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v8","fulltext_source":"unpaywall_pdf"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v8","fair_fulltext_source":"unpaywall_pdf","fair_has_llm":true,"fair_computed_at":"2026-07-20T10:49:52.491448Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}