{"doi":"10.1093/ije/dyae012","title":"Cohort Profile: Dementia Risk Prediction Project (DRPP)","abstract":"The Dementia Risk Prediction Project (DRPP) was established to bring together data from 16 longitudinal cohorts of diverse middle-aged and older adults to provide a resource for research of dementia and its vascular and lifestyle risk factors, and develop a dynamic dementia risk prediction model. Sixteen cohorts with participants from 49 of the 50 US states, France, Iceland, England and the Netherlands with baseline exams ranging from 1948 to 2006 were included. Participants were recruited across multiple sites with multiple in-person assessments of clinical, genetic and behavioural risk factors, follow-up of >10 years and ongoing Alzheimer’s disease and related dementias ascertainment. Collectively, there are 119 061 individuals with data available at baseline. The DRPP provides direct access to harmonized individual-level data for 95 134 individuals from 14 of the 16 cohorts. Median baseline age is 59 years (interquartile range: 49–70) with a maximum age at follow-up of >90 years. Frequency of follow-up varies by cohort, with exams occurring every 2–12 years; the number of in-person exams ranges from 2 to 32. To apply for permission to access our data on our secure portal, please visit our website at drpp.northwestern.edu. Dementia is a major public health problem; despite declines in the age-specific incidence, its prevalence will continue to increase due to ageing of populations.1–3 Worldwide, an estimated 55 million people are living with Alzheimer’s or other dementias and, as the population continues to age, this number is expected to rise to 78 million in 2030 and 139 million in 2050.4 The number of individuals living with dementia is increasing, leading to greater burden of morbidity, caregiving needs and healthcare utilization. Although we have made strides in reducing mortality and morbidity from other diseases such as cardiovascular disease, stroke and cancer, the proportion of deaths due to dementia has increased by >145% in the USA in the past 20 years5 and is presently the seventh leading cause of death among all diseases globally.4 Dementia encompasses a set of complex chronic disorders with neuropathology that is often ‘mixed’, with contributions from vascular and neurodegenerative pathologies. Specifically, >70% of Alzheimer’s disease and related dementias (ADRD) cases are estimated to have mixed vascular and Alzheimer’s disease (AD) pathology.6 Further, many ADRD risk factor profiles are stronger predictors in midlife than in late life. Therefore, the risk factor profiles are also complex and variable. Although non-modifiable risk factors such as age, sex, race/ethnicity and apolipoprotein E (APOE) genotype impact dementia risk, it has been estimated that 40% of all dementia cases could be prevented or delayed by targeting modifiable risk factors.7 Modifiable risk factors that contribute to the largest proportion of dementia cases include midlife hearing status, diabetes, midlife hypertension, midlife obesity, depression, physical inactivity, smoking, low education, excessive alcohol consumption, head injury and air pollution.7 For example, an analysis of 7878 Japanese American men in the Honolulu-Asia Aging Study (HAAS) found that untreated midlife hypertension alone contributes to 27% of dementia risk.8 In a review of English-language systematic reviews and meta-analyses, these modifiable risk factors combined contributed to 54.1% of the population attributable risk for dementia in the USA and 50.7% worldwide.9 Several studies have examined longitudinal trajectories as well as the visit-to-visit variability of risk factors in the long prodrome preceding ADRD diagnosis.10,11 Among them, trajectories of blood pressure, body mass index (BMI) and cholesterol have been studied extensively.10,12,13 Interestingly, most of those studies show an age-dependent pattern in the association between these risk factors and cognitive impairment.12 For example, in a paper using data from Whitehall II participants, different ","journal":"International Journal of Epidemiology","year":2024,"id":449930,"datarank":0.24141568686511508,"base_score":1.6094379124341003,"endowment":1.6094379124341003,"self_citation_contribution":0.24141568686511508,"citation_network_contribution":0.0,"self_endowment_contribution":0.24141568686511508,"citer_contribution":0.0,"corpus_percentile":38.748356153786645,"corpus_rank":7799,"citation_count":4,"citer_count":1,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9379,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2024-01-01","fair_score":33.3333,"fair_percentile":47.93641088352186,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1269727,"name":"John Stephen","orcid":"0000-0001-7309-9193","position":1,"is_corresponding":false},{"id":1270190,"name":"Padraig Carolan","orcid":null,"position":2,"is_corresponding":false},{"id":109446,"name":"Sanaz Sedaghat","orcid":"0000-0002-8296-1811","position":3,"is_corresponding":false},{"id":671144,"name":"Maxwell Mansolf","orcid":"0000-0001-6861-8657","position":4,"is_corresponding":false},{"id":245380,"name":"Aïcha Soumaré","orcid":"0000-0002-0178-6506","position":5,"is_corresponding":false},{"id":322305,"name":"Alden L. Gross","orcid":"0000-0002-4678-0562","position":6,"is_corresponding":false},{"id":109254,"name":"Allison E. Aiello","orcid":"0000-0001-7029-2537","position":7,"is_corresponding":false},{"id":108685,"name":"Archana Singh‐Manoux","orcid":"0000-0002-1244-5037","position":8,"is_corresponding":false},{"id":38764,"name":"M Arfan Ikram","orcid":"0000-0003-0372-8585","position":9,"is_corresponding":false},{"id":305418,"name":"Catherine Helmer","orcid":"0000-0002-5169-7421","position":10,"is_corresponding":false},{"id":17482,"name":"Christophe Tzourio","orcid":"0000-0002-6517-2984","position":11,"is_corresponding":false},{"id":245362,"name":"Claudia L. Satizábal","orcid":"0000-0002-1115-4430","position":12,"is_corresponding":false},{"id":399917,"name":"Deborah A. Levine","orcid":"0000-0002-6864-8001","position":13,"is_corresponding":false},{"id":73090,"name":"Donald M. Lloyd‐Jones","orcid":"0000-0003-0847-6110","position":14,"is_corresponding":false},{"id":426134,"name":"Emily M. Briceño","orcid":"0000-0002-9360-8917","position":15,"is_corresponding":false},{"id":424355,"name":"Farzaneh A. Sorond","orcid":"0000-0002-6127-0739","position":16,"is_corresponding":false},{"id":483828,"name":"Frank J. Wolters","orcid":null,"position":17,"is_corresponding":false},{"id":276095,"name":"Jayandra J. Himali","orcid":null,"position":18,"is_corresponding":false},{"id":17636,"name":"Lenore J. Launer","orcid":"0000-0002-3238-7612","position":19,"is_corresponding":false},{"id":441256,"name":"Lihui Zhao","orcid":"0000-0002-6877-9490","position":20,"is_corresponding":false},{"id":258458,"name":"Mary N. Haan","orcid":"0000-0001-9312-4501","position":21,"is_corresponding":false},{"id":245369,"name":"Oscar L. López","orcid":"0000-0002-8546-8256","position":22,"is_corresponding":false},{"id":17645,"name":"Stéphanie Debette","orcid":"0000-0001-8675-7968","position":23,"is_corresponding":false},{"id":17639,"name":"Sudha Seshadri","orcid":"0000-0001-6135-2622","position":24,"is_corresponding":false},{"id":74972,"name":"Suzanne E. Judd","orcid":"0000-0001-7594-6587","position":25,"is_corresponding":false},{"id":325846,"name":"Timothy M. Hughes","orcid":"0000-0002-2919-7199","position":26,"is_corresponding":false},{"id":22041,"name":"Vilmundur Guðnason","orcid":"0000-0001-5696-0084","position":27,"is_corresponding":false},{"id":331145,"name":"Denise Scholtens","orcid":"0000-0002-8252-7863","position":28,"is_corresponding":false},{"id":278983,"name":"Norrina B. Allen","orcid":"0000-0002-6993-4384","position":29,"is_corresponding":false},{"id":735156,"name":"Amy E. Krefman","orcid":"0000-0002-6692-0104","position":0,"is_corresponding":true}],"reference_count":38,"raw_metadata":null,"created_at":"2026-07-19T02:02:20.585759Z","pmid":"38339864","pmcid":"PMC10858348","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":61.1111,"fair_a":43.75,"fair_i":20.0,"fair_r":33.3333,"fair_zscore":-0.0442,"fair_rationale":{"fair_score":33.33,"has_llm":true,"taxonomy_version":"fair_taxonomy_v5","dimensions":{"F":{"name":"Findable","score":61.11,"criteria":[{"key":"f_dataset_pid","label":"Persistent identifier for the data","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"please visit our website at drpp.northwestern.edu","grounded":true,"rationale":"The paper gives a URL for the data location, not a persistent identifier scheme. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit","RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier'","FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'"],"scored":true,"signal":null},{"key":"f_repository_named","label":"Named repository","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"please visit our website at drpp.northwestern.edu","grounded":true,"rationale":"The data are held on a project website, not a named repository.","anchors":["RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed (","NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived","NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten"],"scored":true,"signal":null},{"key":"f_data_availability_statement","label":"Data-availability statement","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"Investigators can apply for use of the DRPP data or review the data documentation at drpp.northwestern.edu.","grounded":true,"rationale":"The statement points to a project website with an application process, not a repository record. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li","Springer Nature research data policy — Data Availability Statements: standard statement templat","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes"],"scored":false,"signal":null},{"key":"f_discovery_metadata","label":"Description of the dataset as an object","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Supplementary Table S11. DRPP harmonized variables and domains.","grounded":true,"rationale":"The paper includes a table listing the variables in the harmonized dataset, which is an itemised inventory.","anchors":["RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential)","FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability'","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'"],"scored":false,"signal":null},{"key":"f_dataset_cited","label":"Dataset formally cited","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"To apply for permission to access our data on our secure portal, please visit our website at drpp.northwestern.edu.","grounded":true,"rationale":"The dataset identifier (the website URL) appears only in the body text, not in the reference list. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first-","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes","FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'"],"scored":true,"signal":null}]},"A":{"name":"Accessible","score":43.75,"criteria":[{"key":"a_data_openly_accessible","label":"Access route free of preconditions","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"To apply for permission to access our data on our secure portal, please visit our website at drpp.northwestern.edu.","grounded":true,"rationale":"Access requires an application process, which is a specified precondition.","anchors":["RDA-A1.1-01D — 'Data is accessible through a free access protocol'","FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'"],"scored":true,"signal":null},{"key":"a_access_conditions_stated","label":"Access level labelled","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"Approved collaborators will have access to a data set on the virtual computing platform built specifically for the DRPP.","grounded":true,"rationale":"The paper describes the access procedure but does not label the access level with a standard term.","anchors":["FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","RDA-A1-01M — metadata contains information to enable the user to get access to the data","COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl"],"scored":false,"signal":null},{"key":"a_controlled_access_for_sensitive","label":"Gatekeeper for sensitive data","kind":"llm","weight":0.5,"fraction":0.5,"verdict":"partial","evidence":"Output is then reviewed by the Northwestern DRPP team to confirm that no individual-level data leave our server.","grounded":false,"rationale":"The paper names an institutional gatekeeper (the Northwestern DRPP team) that reviews output before release, constituting a controlled-access procedure. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (2/5 passes agreed)]","anchors":["NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee","RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and ","NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse"],"scored":false,"signal":null},{"key":"a_timeline_retention","label":"Availability timing & retention","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No sentence states when the data become available or how long they persist.","anchors":["NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines","NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy'","RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'"],"scored":false,"signal":null}]},"I":{"name":"Interoperable","score":20.0,"criteria":[{"key":"i_open_nonproprietary_format","label":"Open file format","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No file format is named for the released data.","anchors":["FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co","RDA-R1.3-02D — data is expressed in a machine-understandable community standard","RDA-I1-01D — data uses a knowledge representation expressed in a standardised format"],"scored":true,"signal":null},{"key":"i_community_standard_vocabulary","label":"Community standard / vocabulary","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"diagnosed according to the Diagnostic and Statistical Manual of Mental Disorders, fourth edition (DSM-IV) criteria","grounded":false,"rationale":"DSM-IV is a community-standard diagnostic vocabulary used for the data. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (2/5 passes agreed)]","anchors":["RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential)","RDA-R1.3-01D — 'Data complies with a community standard'","RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'"],"scored":false,"signal":null},{"key":"i_qualified_references","label":"Identifiers for the resources the data depend on","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No identifier for any external resource (e.g., another dataset, database, or code) is provided.","anchors":["RDA-I3-01M — '(meta)data include references to other (meta)data'","RDA-I3-03M — 'metadata includes qualified references to other metadata'","FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'"],"scored":false,"signal":null}]},"R":{"name":"Reusable","score":33.33,"criteria":[{"key":"r_reuse_license","label":"Reuse licence","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No licence or terms of use are stated for the data.","anchors":["RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu","RDA-R1.1-02M — 'Metadata refers to a standard reuse licence'","RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'"],"scored":true,"signal":null},{"key":"r_provenance_methods","label":"Provenance of the data","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"our team developed an R package titled ‘psHarmonize’.","grounded":true,"rationale":"The paper names the specific software tool used for data harmonization. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa","FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati","W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance"],"scored":false,"signal":null},{"key":"r_documentation_codebook","label":"Documentation / codebook","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"Supplementary Table S11. DRPP harmonized variables and domains.","grounded":true,"rationale":"The variable definitions are provided inside the article as a supplementary table, not as a separate documentation object shipped with the data.","anchors":["RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'","NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t"],"scored":false,"signal":null},{"key":"r_versioning","label":"Snapshot identified","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No version token or date is given to identify the snapshot of the data.","anchors":["DataCite Metadata Schema 4.6 — the 'Version' property","RDA-R1.2-01M — provenance information (which version was used is provenance)","NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'"],"scored":true,"signal":null},{"key":"x_code_availability","label":"Analysis code available","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The paper names the R package but provides no locator (URL, DOI, or repository) for it.","anchors":["NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code'","FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear","FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)"],"scored":true,"signal":null},{"key":"x_funding_attribution","label":"Funder and award number","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"1R61NS120245-01/R33NS120245","grounded":true,"rationale":"The paper provides specific grant numbers for the funding.","anchors":["DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award","Crossref Funder Registry — canonical funder identifiers for funding metadata","RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco"],"scored":true,"signal":null}]}},"actions":[{"key":"r_reuse_license","dimension":"R","label":"Reuse licence","action":"Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No licence or terms of use are stated for the data.","gain":16.67,"priority":"essential","scored":true},{"key":"f_dataset_pid","dimension":"F","label":"Persistent identifier for the data","action":"Mint or cite a persistent identifier for the dataset — a repository DOI or an accession from a registered repository — and print it in the paper. A bare URL is not persistent: it is the single most common cause of a dead data link five years after publication. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"please visit our website at drpp.northwestern.edu","why":"The paper gives a URL for the data location, not a persistent identifier scheme. [majority verdict 'partial' (4/5 passes agreed)]","gain":8.33,"priority":"essential","scored":true},{"key":"f_repository_named","dimension":"F","label":"Named repository","action":"Deposit the data in a repository registered in re3data/FAIRsharing (a domain repository such as GEO, SRA, dbGaP, PRIDE, or a generalist such as Zenodo, Dryad, Dataverse) and name it explicitly in the paper. A lab website is not an archive: it has no retention commitment and no accession. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"please visit our website at drpp.northwestern.edu","why":"The data are held on a project website, not a named repository.","gain":8.33,"priority":"essential","scored":true},{"key":"a_data_openly_accessible","dimension":"A","label":"Access route free of preconditions","action":"Remove the precondition or justify it. Release the data at publication with no embargo, no registration wall, and no approval step — NIH's zero-embargo public- access rule (NOT-OD-25-101) has already made 'available at publication' the federal baseline for the article; the data should not lag behind it. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"To apply for permission to access our data on our secure portal, please visit our website at drpp.northwestern.edu.","why":"Access requires an application process, which is a specified precondition.","gain":8.33,"priority":"essential","scored":true},{"key":"i_open_nonproprietary_format","dimension":"I","label":"Open file format","action":"Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No file format is named for the released data.","gain":8.33,"priority":"important","scored":true},{"key":"x_code_availability","dimension":"R","label":"Analysis code available","action":"Publish the analysis code in a public forge, archive a tagged release with a DOI (Zenodo/Software Heritage), and cite that DOI in the paper. NIH DMS Element 2 asks for the tools and code, not only the data — and 'available on request' is not a locator. Archive the analysis code in a versioned repository (GitHub + a Zenodo release DOI).","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The paper names the R package but provides no locator (URL, DOI, or repository) for it.","gain":8.33,"priority":"important","scored":true},{"key":"f_dataset_cited","dimension":"F","label":"Dataset formally cited","action":"Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the clinical / human-subjects repository accession (e.g. from dbGaP or the European Genome-phenome Archive (EGA)) in the reference list.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"To apply for permission to access our data on our secure portal, please visit our website at drpp.northwestern.edu.","why":"The dataset identifier (the website URL) appears only in the body text, not in the reference list. [majority verdict 'partial' (4/5 passes agreed)]","gain":4.17,"priority":"important","scored":true},{"key":"r_versioning","dimension":"R","label":"Snapshot identified","action":"Version the deposit and cite the exact version analysed (a version-specific DOI, or an accession with its version suffix). A reader reproducing your work against 'the current release' is reproducing it against a different dataset.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No version token or date is given to identify the snapshot of the data.","gain":4.17,"priority":"useful","scored":true},{"key":"f_data_availability_statement","dimension":"F","label":"Data-availability statement","action":"Replace the statement with the repository template: name the repository and give the accession or DOI (Colavizza category 3). This is the only DAS class associated with a measured citation advantage; 'available on reasonable request' and 'within the article' are not.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Investigators can apply for use of the DRPP data or review the data documentation at drpp.northwestern.edu.","why":"The statement points to a project website with an application process, not a repository record. [majority verdict 'partial' (4/5 passes agreed)]","gain":0.0,"priority":"essential","scored":false},{"key":"a_access_conditions_stated","dimension":"A","label":"Access level labelled","action":"State the access level in words, using the standard vocabulary: 'These data are open access' / 'These data are controlled access'. A reader — and a harvester — should not have to infer the access level from the presence of a download link.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Approved collaborators will have access to a data set on the virtual computing platform built specifically for the DRPP.","why":"The paper describes the access procedure but does not label the access level with a standard term.","gain":0.0,"priority":"important","scored":false},{"key":"i_community_standard_vocabulary","dimension":"I","label":"Community standard / vocabulary","action":"Adopt and NAME your domain's data standard — the minimum-information checklist, metadata schema, or ontology your community uses (MIAME/MINSEQE, ISA-Tab, BIDS, an OBO ontology, HL7 FHIR/OMOP) — and say which one you followed. A reporting checklist standardises your paper; it does nothing for your data. In clinical / human-subjects, describe the data with OMOP CDM, CDISC SDTM or HL7 FHIR.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"diagnosed according to the Diagnostic and Statistical Manual of Mental Disorders, fourth edition (DSM-IV) criteria","why":"DSM-IV is a community-standard diagnostic vocabulary used for the data. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (2/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"r_documentation_codebook","dimension":"R","label":"Documentation / codebook","action":"Ship a README and a data dictionary IN the deposit — every file, every variable, its units, its allowed values, its missing-value codes. It is the cheapest single thing that makes a dataset usable by someone who was not in the lab, and a table buried in the article does not travel with the data.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Supplementary Table S11. DRPP harmonized variables and domains.","why":"The variable definitions are provided inside the article as a supplementary table, not as a separate documentation object shipped with the data.","gain":0.0,"priority":"important","scored":false},{"key":"a_controlled_access_for_sensitive","dimension":"A","label":"Gatekeeper for sensitive data","action":"Route sensitive data through an institutional gatekeeper — deposit in a controlled- access repository (dbGaP, EGA) with a Data Access Committee and a published DUA — rather than through the corresponding author's inbox. An author-gated dataset dies with the author's email address, and 'on reasonable request' has been shown repeatedly not to yield data. For sensitive/human clinical / human-subjects data, use a controlled-access repository such as dbGaP or EGA.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Output is then reviewed by the Northwestern DRPP team to confirm that no individual-level data leave our server.","why":"The paper names an institutional gatekeeper (the Northwestern DRPP team) that reviews output before release, constituting a controlled-access procedure. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (2/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false},{"key":"i_qualified_references","dimension":"I","label":"Identifiers for the resources the data depend on","action":"Cite by identifier every resource the data depend on — the source datasets' accessions, the reference build (GRCh38 / GCA_000001405.28), the cohort application number, the code DOI — and register those relations on the dataset record (IsDerivedFrom, IsSupplementTo). A name is not a link: it cannot be resolved, versioned, or followed by a machine.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No identifier for any external resource (e.g., another dataset, database, or code) is provided.","gain":0.0,"priority":"useful","scored":false},{"key":"a_timeline_retention","dimension":"A","label":"Availability timing & retention","action":"State when the data become available AND how long they will be retained — cite the repository's preservation policy. NIH DMS Element 4 asks for both; most papers give neither.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No sentence states when the data become available or how long they persist.","gain":0.0,"priority":"useful","scored":false}],"suggestions":["Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","Mint or cite a persistent identifier for the dataset — a repository DOI or an accession from a registered repository — and print it in the paper. A bare URL is not persistent: it is the single most common cause of a dead data link five years after publication. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","Deposit the data in a repository registered in re3data/FAIRsharing (a domain repository such as GEO, SRA, dbGaP, PRIDE, or a generalist such as Zenodo, Dryad, Dataverse) and name it explicitly in the paper. A lab website is not an archive: it has no retention commitment and no accession. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","Remove the precondition or justify it. Release the data at publication with no embargo, no registration wall, and no approval step — NIH's zero-embargo public- access rule (NOT-OD-25-101) has already made 'available at publication' the federal baseline for the article; the data should not lag behind it. For clinical / human-subjects data, deposit in dbGaP or the European Genome-phenome Archive (EGA).","Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v8","fulltext_source":"unpaywall_pdf"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v8","fair_fulltext_source":"unpaywall_pdf","fair_has_llm":true,"fair_computed_at":"2026-07-20T13:09:58.302901Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}