{"doi":"10.1016/j.xgen.2022.100243","title":"The UCLA ATLAS Community Health Initiative: Promoting precision health research in a diverse biobank","abstract":"The UCLA ATLAS Community Health Initiative (ATLAS) has an initial target to recruit 150,000 participants from across the UCLA Health system with the goal of creating a genomic database to accelerate precision medicine efforts in California. This initiative includes a biobank embedded within the UCLA Health system that comprises de-identified genomic data linked to electronic health records (EHRs). The first freeze of data from September 2020 contains 27,987 genotyped samples imputed to 7.9 million SNPs across the genome and is linked with de-identified versions of the EHRs from UCLA Health. Here, we describe a centralized repository of the genotype data and provide tools and pipelines to perform genome- and phenome-wide association studies across a wide range of EHR-derived phenotypes and genetic ancestry groups. We demonstrate the utility of this resource through the analysis of 7 well-studied traits and recapitulate many previous genetic and phenotypic associations.","journal":"Cell Genomics","year":2023,"id":330801,"datarank":0.928490712829832,"base_score":3.58351893845611,"endowment":3.58351893845611,"self_citation_contribution":0.5375278407684165,"citation_network_contribution":0.39096287206141545,"self_endowment_contribution":0.5375278407684165,"citer_contribution":0.39096287206141545,"corpus_percentile":78.21613676800496,"corpus_rank":2817,"citation_count":35,"citer_count":19,"citers_with_citation_signal":15,"citers_with_endowment":15,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9458,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":43.75,"fair_percentile":58.6365025985937,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":551627,"name":"Yi Ding","orcid":"0000-0003-3595-2493","position":1,"is_corresponding":false},{"id":21368,"name":"Arjun Bhattacharya","orcid":"0000-0003-1196-4385","position":2,"is_corresponding":false},{"id":309770,"name":"Sergey Knyazev","orcid":null,"position":3,"is_corresponding":false},{"id":329205,"name":"Alec Chiu","orcid":"0000-0002-1646-1149","position":4,"is_corresponding":false},{"id":37391,"name":"Clara Lajonchere","orcid":"0000-0002-3190-4606","position":5,"is_corresponding":false},{"id":21398,"name":"Daniel H. Geschwind","orcid":"0000-0003-2896-3450","position":6,"is_corresponding":false},{"id":731,"name":"Bogdan Paşaniuc","orcid":"0000-0002-0227-2056","position":7,"is_corresponding":false},{"id":273432,"name":"Ruth Johnson","orcid":"0000-0002-1929-0998","position":0,"is_corresponding":true}],"reference_count":39,"raw_metadata":null,"created_at":"2026-07-19T01:09:14.691473Z","pmid":"36777178","pmcid":"PMC9903668","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":50.0,"fair_a":62.5,"fair_i":40.0,"fair_r":37.5,"fair_zscore":0.3681,"fair_rationale":{"fair_score":43.75,"has_llm":true,"taxonomy_version":"fair_taxonomy_v5","dimensions":{"F":{"name":"Findable","score":50.0,"criteria":[{"key":"f_dataset_pid","label":"Persistent identifier for the data","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"https://atlas-phewas.mednet.ucla.edu/","grounded":true,"rationale":"The paper provides a web address for the summary statistics, which is not a persistent identifier scheme (e.g., DOI, Handle, or repository accession). [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit","RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier'","FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'"],"scored":true,"signal":null},{"key":"f_repository_named","label":"Named repository","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"UCLA ATLAS PheWeb browser","grounded":true,"rationale":"The holder named is a project website (PheWeb browser), not a recognised repository. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed (","NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived","NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten"],"scored":true,"signal":null},{"key":"f_data_availability_statement","label":"Data-availability statement","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"Individual-level genotype and electronic health record data utilized in this study cannot be deposited in a public repository because of privacy regulations. GWAS summary statistics are made available on the UCLA ATLAS PheWeb browser (https://atlas-phewas.mednet.ucla.edu/).","grounded":true,"rationale":"The statement points to a web resource (PheWeb browser) for summary statistics, which is not a repository record, and the primary data are not available. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li","Springer Nature research data policy — Data Availability Statements: standard statement templat","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes"],"scored":false,"signal":null},{"key":"f_discovery_metadata","label":"Description of the dataset as an object","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"The first freeze of data from September 2020 contains 27,987 genotyped samples imputed to 7.9 million SNPs across the genome","grounded":true,"rationale":"The dataset's content is described in a prose sentence, not in an itemised inventory (section, table, or list). [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential)","FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability'","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'"],"scored":false,"signal":null},{"key":"f_dataset_cited","label":"Dataset formally cited","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"https://atlas-phewas.mednet.ucla.edu/","grounded":true,"rationale":"The dataset's identifier (the PheWeb URL) appears only in the body text, not in the reference list. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first-","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes","FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'"],"scored":true,"signal":null}]},"A":{"name":"Accessible","score":62.5,"criteria":[{"key":"a_data_openly_accessible","label":"Access route free of preconditions","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"GWAS summary statistics are made available on the UCLA ATLAS PheWeb browser (https://atlas-phewas.mednet.ucla.edu/).","grounded":true,"rationale":"The text gives a route to the summary statistics with no stated precondition; the primary data are not accessible, but the summary statistics qualify as a data product. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["RDA-A1.1-01D — 'Data is accessible through a free access protocol'","FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'"],"scored":true,"signal":null},{"key":"a_access_conditions_stated","label":"Access level labelled","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"GWAS summary statistics are made available on the UCLA ATLAS PheWeb browser (https://atlas-phewas.mednet.ucla.edu/).","grounded":true,"rationale":"The text describes a URL where summary statistics are available but does not label the access level with a standard vocabulary term. [majority verdict 'partial' (3/5 passes agreed)]","anchors":["FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","RDA-A1-01M — metadata contains information to enable the user to get access to the data","COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl"],"scored":false,"signal":null},{"key":"a_controlled_access_for_sensitive","label":"Gatekeeper for sensitive data","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":"Individual-level genotype and electronic health record data utilized in this study cannot be deposited in a public repository because of privacy regulations.","grounded":true,"rationale":"The paper states that individual-level data cannot be deposited and provides no gatekeeper (neither institutional nor personal) for access.","anchors":["NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee","RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and ","NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse"],"scored":false,"signal":null},{"key":"a_timeline_retention","label":"Availability timing & retention","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No sentence states when the data become available (other than 'made available' without timing) or how long they persist. [majority verdict 'no' (3/5 passes agreed)]","anchors":["NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines","NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy'","RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'"],"scored":false,"signal":null}]},"I":{"name":"Interoperable","score":40.0,"criteria":[{"key":"i_open_nonproprietary_format","label":"Open file format","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The paper does not name any file format for the released data.","anchors":["FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co","RDA-R1.3-02D — data is expressed in a machine-understandable community standard","RDA-I1-01D — data uses a knowledge representation expressed in a standardised format"],"scored":true,"signal":null},{"key":"i_community_standard_vocabulary","label":"Community standard / vocabulary","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"We use phecodes, a coding system that maps diagnosis codes (i.e., ICD-9 and ICD-10 codes) to more clinically meaningful phenotypes","grounded":true,"rationale":"Phecodes are a community-standard vocabulary for EHR phenotyping, registered in FAIRsharing. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential)","RDA-R1.3-01D — 'Data complies with a community standard'","RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'"],"scored":false,"signal":null},{"key":"i_qualified_references","label":"Identifiers for the resources the data depend on","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":"http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/phase3/","grounded":false,"rationale":"The paper provides a URL identifier for the 1000 Genomes Project reference panel, a resource other than the study's own dataset. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (4/5 passes agreed)]","anchors":["RDA-I3-01M — '(meta)data include references to other (meta)data'","RDA-I3-03M — 'metadata includes qualified references to other metadata'","FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'"],"scored":false,"signal":null}]},"R":{"name":"Reusable","score":37.5,"criteria":[{"key":"r_reuse_license","label":"Reuse licence","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":"This is an open access article under the CC BY-NC-ND license","grounded":true,"rationale":"The CC BY-NC-ND license applies to the article, not to the data. No license for the data is named.","anchors":["RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu","RDA-R1.1-02M — 'Metadata refers to a standard reuse licence'","RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'"],"scored":true,"signal":null},{"key":"r_provenance_methods","label":"Provenance of the data","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"Genotyping was performed at the UCLA Neuroscience Genomics Core using a custom genotyping array constructed from the Global Screening Array with the multi-disease drop-in panel 3 under the GRCh37 assembly.","grounded":true,"rationale":"The paper names the specific genotyping array (Illumina Global Screening Array) and other tools, providing provenance. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa","FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati","W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance"],"scored":false,"signal":null},{"key":"r_documentation_codebook","label":"Documentation / codebook","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"Table 1. Summary of UCLA ATLAS demographics","grounded":true,"rationale":"The variable definitions (e.g., demographics) are provided inside the article in a table, not in a separate documentation file shipped with the data. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'","NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t"],"scored":false,"signal":null},{"key":"r_versioning","label":"Snapshot identified","kind":"llm","weight":0.5,"fraction":0.5,"verdict":"partial","evidence":"The first freeze of data from September 2020 contains 27,987 genotyped samples","grounded":true,"rationale":"The paper identifies the snapshot by a date and a descriptive name ('first freeze') but does not provide a formal version token. [majority verdict 'partial' (3/5 passes agreed)]","anchors":["DataCite Metadata Schema 4.6 — the 'Version' property","RDA-R1.2-01M — provenance information (which version was used is provenance)","NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'"],"scored":true,"signal":null},{"key":"x_code_availability","label":"Analysis code available","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":"This paper does not report original code.","grounded":true,"rationale":"The paper states that no original code is reported, so no code location is provided.","anchors":["NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code'","FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear","FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)"],"scored":true,"signal":null},{"key":"x_funding_attribution","label":"Funder and award number","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"UL1TR001881","grounded":true,"rationale":"A grant number (UL1TR001881) from the UCLA Clinical and Translational Science Institute is provided. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award","Crossref Funder Registry — canonical funder identifiers for funding metadata","RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco"],"scored":true,"signal":null}]}},"actions":[{"key":"r_reuse_license","dimension":"R","label":"Reuse licence","action":"Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"This is an open access article under the CC BY-NC-ND license","why":"The CC BY-NC-ND license applies to the article, not to the data. No license for the data is named.","gain":16.67,"priority":"essential","scored":true},{"key":"f_dataset_pid","dimension":"F","label":"Persistent identifier for the data","action":"Mint or cite a persistent identifier for the dataset — a repository DOI or an accession from a registered repository — and print it in the paper. A bare URL is not persistent: it is the single most common cause of a dead data link five years after publication. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"https://atlas-phewas.mednet.ucla.edu/","why":"The paper provides a web address for the summary statistics, which is not a persistent identifier scheme (e.g., DOI, Handle, or repository accession). [majority verdict 'partial' (4/5 passes agreed)]","gain":8.33,"priority":"essential","scored":true},{"key":"f_repository_named","dimension":"F","label":"Named repository","action":"Deposit the data in a repository registered in re3data/FAIRsharing (a domain repository such as GEO, SRA, dbGaP, PRIDE, or a generalist such as Zenodo, Dryad, Dataverse) and name it explicitly in the paper. A lab website is not an archive: it has no retention commitment and no accession. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"UCLA ATLAS PheWeb browser","why":"The holder named is a project website (PheWeb browser), not a recognised repository. [majority verdict 'partial' (4/5 passes agreed)]","gain":8.33,"priority":"essential","scored":true},{"key":"i_open_nonproprietary_format","dimension":"I","label":"Open file format","action":"Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The paper does not name any file format for the released data.","gain":8.33,"priority":"important","scored":true},{"key":"x_code_availability","dimension":"R","label":"Analysis code available","action":"Publish the analysis code in a public forge, archive a tagged release with a DOI (Zenodo/Software Heritage), and cite that DOI in the paper. NIH DMS Element 2 asks for the tools and code, not only the data — and 'available on request' is not a locator. Archive the analysis code in a versioned repository (GitHub + a Zenodo release DOI).","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"This paper does not report original code.","why":"The paper states that no original code is reported, so no code location is provided.","gain":8.33,"priority":"important","scored":true},{"key":"f_dataset_cited","dimension":"F","label":"Dataset formally cited","action":"Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the clinical / human-subjects repository accession (e.g. from dbGaP or the European Genome-phenome Archive (EGA)) in the reference list.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"https://atlas-phewas.mednet.ucla.edu/","why":"The dataset's identifier (the PheWeb URL) appears only in the body text, not in the reference list. [majority verdict 'partial' (4/5 passes agreed)]","gain":4.17,"priority":"important","scored":true},{"key":"r_versioning","dimension":"R","label":"Snapshot identified","action":"Version the deposit and cite the exact version analysed (a version-specific DOI, or an accession with its version suffix). A reader reproducing your work against 'the current release' is reproducing it against a different dataset.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"The first freeze of data from September 2020 contains 27,987 genotyped samples","why":"The paper identifies the snapshot by a date and a descriptive name ('first freeze') but does not provide a formal version token. [majority verdict 'partial' (3/5 passes agreed)]","gain":2.08,"priority":"useful","scored":true},{"key":"f_data_availability_statement","dimension":"F","label":"Data-availability statement","action":"Replace the statement with the repository template: name the repository and give the accession or DOI (Colavizza category 3). This is the only DAS class associated with a measured citation advantage; 'available on reasonable request' and 'within the article' are not.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Individual-level genotype and electronic health record data utilized in this study cannot be deposited in a public repository because of privacy regulations. GWAS summary statistics are made available on the UCLA ATLAS PheWeb browser (https://atlas-phewas.mednet.ucla.edu/).","why":"The statement points to a web resource (PheWeb browser) for summary statistics, which is not a repository record, and the primary data are not available. [majority verdict 'partial' (4/5 passes agreed)]","gain":0.0,"priority":"essential","scored":false},{"key":"f_discovery_metadata","dimension":"F","label":"Description of the dataset as an object","action":"Add a 'Data Records' section: itemise every file in the deposit and every variable or sample it holds, with counts and units. Describe the dataset as an object in its own right, not as a by-product of the findings — this is what makes it discoverable to someone who is not looking for your paper.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"The first freeze of data from September 2020 contains 27,987 genotyped samples imputed to 7.9 million SNPs across the genome","why":"The dataset's content is described in a prose sentence, not in an itemised inventory (section, table, or list). [majority verdict 'partial' (4/5 passes agreed)]","gain":0.0,"priority":"essential","scored":false},{"key":"a_access_conditions_stated","dimension":"A","label":"Access level labelled","action":"State the access level in words, using the standard vocabulary: 'These data are open access' / 'These data are controlled access'. A reader — and a harvester — should not have to infer the access level from the presence of a download link.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"GWAS summary statistics are made available on the UCLA ATLAS PheWeb browser (https://atlas-phewas.mednet.ucla.edu/).","why":"The text describes a URL where summary statistics are available but does not label the access level with a standard vocabulary term. [majority verdict 'partial' (3/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"r_documentation_codebook","dimension":"R","label":"Documentation / codebook","action":"Ship a README and a data dictionary IN the deposit — every file, every variable, its units, its allowed values, its missing-value codes. It is the cheapest single thing that makes a dataset usable by someone who was not in the lab, and a table buried in the article does not travel with the data.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Table 1. Summary of UCLA ATLAS demographics","why":"The variable definitions (e.g., demographics) are provided inside the article in a table, not in a separate documentation file shipped with the data. [majority verdict 'partial' (4/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"a_controlled_access_for_sensitive","dimension":"A","label":"Gatekeeper for sensitive data","action":"Route sensitive data through an institutional gatekeeper — deposit in a controlled- access repository (dbGaP, EGA) with a Data Access Committee and a published DUA — rather than through the corresponding author's inbox. An author-gated dataset dies with the author's email address, and 'on reasonable request' has been shown repeatedly not to yield data. For sensitive/human clinical / human-subjects data, use a controlled-access repository such as dbGaP or EGA.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"Individual-level genotype and electronic health record data utilized in this study cannot be deposited in a public repository because of privacy regulations.","why":"The paper states that individual-level data cannot be deposited and provides no gatekeeper (neither institutional nor personal) for access.","gain":0.0,"priority":"useful","scored":false},{"key":"i_qualified_references","dimension":"I","label":"Identifiers for the resources the data depend on","action":"Cite by identifier every resource the data depend on — the source datasets' accessions, the reference build (GRCh38 / GCA_000001405.28), the cohort application number, the code DOI — and register those relations on the dataset record (IsDerivedFrom, IsSupplementTo). A name is not a link: it cannot be resolved, versioned, or followed by a machine.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/phase3/","why":"The paper provides a URL identifier for the 1000 Genomes Project reference panel, a resource other than the study's own dataset. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (4/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false},{"key":"a_timeline_retention","dimension":"A","label":"Availability timing & retention","action":"State when the data become available AND how long they will be retained — cite the repository's preservation policy. NIH DMS Element 4 asks for both; most papers give neither.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No sentence states when the data become available (other than 'made available' without timing) or how long they persist. [majority verdict 'no' (3/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false}],"suggestions":["Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","Mint or cite a persistent identifier for the dataset — a repository DOI or an accession from a registered repository — and print it in the paper. A bare URL is not persistent: it is the single most common cause of a dead data link five years after publication. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","Deposit the data in a repository registered in re3data/FAIRsharing (a domain repository such as GEO, SRA, dbGaP, PRIDE, or a generalist such as Zenodo, Dryad, Dataverse) and name it explicitly in the paper. A lab website is not an archive: it has no retention commitment and no accession. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable.","Publish the analysis code in a public forge, archive a tagged release with a DOI (Zenodo/Software Heritage), and cite that DOI in the paper. NIH DMS Element 2 asks for the tools and code, not only the data — and 'available on request' is not a locator. Archive the analysis code in a versioned repository (GitHub + a Zenodo release DOI)."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v8","fulltext_source":"unpaywall_pdf"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v8","fair_fulltext_source":"unpaywall_pdf","fair_has_llm":true,"fair_computed_at":"2026-07-20T11:30:04.358216Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}