{"doi":"10.1093/nar/gkaa921","title":"RNAcentral 2021: secondary structure integration, improved sequence search and new member databases","abstract":"RNAcentral is a comprehensive database of non-coding RNA (ncRNA) sequences that provides a single access point to 44 RNA resources and >18 million ncRNA sequences from a wide range of organisms and RNA types. RNAcentral now also includes secondary (2D) structure information for >13 million sequences, making RNAcentral the world's largest RNA 2D structure database. The 2D diagrams are displayed using R2DT, a new 2D structure visualization method that uses consistent, reproducible and recognizable layouts for related RNAs. The sequence similarity search has been updated with a faster interface featuring facets for filtering search results by RNA type, organism, source database or any keyword. This sequence search tool is available as a reusable web component, and has been integrated into several RNAcentral member databases, including Rfam, miRBase and snoDB. To allow for a more fine-grained assignment of RNA types and subtypes, all RNAcentral sequences have been annotated with Sequence Ontology terms. The RNAcentral database continues to grow and provide a central data resource for the RNA community. RNAcentral is freely available at https://rnacentral.org.","journal":"Nucleic Acids Research","year":2020,"id":48441,"datarank":5.06725453195081,"base_score":6.040254711277414,"endowment":6.040254711277414,"self_citation_contribution":0.9060382066916122,"citation_network_contribution":4.161216325259198,"self_endowment_contribution":0.9060382066916122,"citer_contribution":4.161216325259198,"corpus_percentile":95.81496093447822,"corpus_rank":542,"citation_count":419,"citer_count":100,"citers_with_citation_signal":100,"citers_with_endowment":100,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9482,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2020-01-01","fair_score":41.6667,"fair_percentile":54.173035768878016,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":106800,"name":"Anton I. Petrov","orcid":"0000-0001-7279-2682","position":1,"is_corresponding":false},{"id":228791,"name":"Carlos Eduardo Ribas","orcid":"0000-0002-9572-273X","position":2,"is_corresponding":false},{"id":103841,"name":"ROBERT FINN","orcid":"0000-0001-8626-2148","position":3,"is_corresponding":false},{"id":228792,"name":"Alex Bateman","orcid":"0000-0002-6982-4660","position":4,"is_corresponding":false},{"id":228793,"name":"M. Szymański","orcid":"0000-0001-7694-6706","position":5,"is_corresponding":false},{"id":86711,"name":"Wojciech M. Karłowski","orcid":"0000-0002-8086-5404","position":6,"is_corresponding":false},{"id":228794,"name":"Stefan E. Seemann","orcid":"0000-0002-2359-4927","position":7,"is_corresponding":false},{"id":51099,"name":"Jan Gorodkin","orcid":"0000-0001-5823-4000","position":8,"is_corresponding":false},{"id":228795,"name":"Jamie J. Cannone","orcid":"0000-0002-0060-7887","position":9,"is_corresponding":false},{"id":228796,"name":"Robin R. Gutell","orcid":"0000-0002-5919-0801","position":10,"is_corresponding":false},{"id":56293,"name":"Simon Kay","orcid":"0000-0001-7690-2711","position":11,"is_corresponding":false},{"id":103818,"name":"Steven J Marygold","orcid":"0000-0003-2759-266X","position":12,"is_corresponding":false},{"id":103820,"name":"Gil dos Santos","orcid":"0000-0003-3507-8273","position":13,"is_corresponding":false},{"id":105364,"name":"Adam Frankish","orcid":"0000-0002-4333-628X","position":14,"is_corresponding":false},{"id":105367,"name":"Jonathan M. Mudge","orcid":"0000-0003-4789-7495","position":15,"is_corresponding":false},{"id":228797,"name":"Ruth Barshir","orcid":"0000-0003-0776-9587","position":16,"is_corresponding":false},{"id":228798,"name":"Simon Fishilevich","orcid":"0000-0003-1167-1638","position":17,"is_corresponding":false},{"id":228799,"name":"Patricia P. Chan","orcid":"0000-0003-0810-1642","position":18,"is_corresponding":false},{"id":19953,"name":"Todd M. Lowe","orcid":"0000-0003-3253-6021","position":19,"is_corresponding":false},{"id":228800,"name":"Ruth L. Seal","orcid":"0000-0002-7545-6817","position":20,"is_corresponding":false},{"id":78836,"name":"Elspeth A. Bruford","orcid":"0000-0002-8380-5247","position":21,"is_corresponding":false},{"id":228801,"name":"Simona Panni","orcid":"0000-0002-7500-4028","position":22,"is_corresponding":false},{"id":103839,"name":"Pablo Porras","orcid":"0000-0002-8429-8793","position":23,"is_corresponding":false},{"id":79586,"name":"Dimitra Karagkouni","orcid":"0009-0007-3961-3399","position":24,"is_corresponding":false},{"id":228802,"name":"Artemis G. Hatzigeorgiou","orcid":"0000-0003-1414-5668","position":25,"is_corresponding":false},{"id":228803,"name":"Lina Ma","orcid":"0000-0001-6390-6289","position":26,"is_corresponding":false},{"id":50353,"name":"Zhang Zhang","orcid":"0000-0001-6603-5060","position":27,"is_corresponding":false},{"id":228804,"name":"Pieter‐Jan Volders","orcid":"0000-0002-2685-2637","position":28,"is_corresponding":false},{"id":228805,"name":"Pieter Mestdagh","orcid":"0000-0001-7821-9684","position":29,"is_corresponding":false},{"id":57002,"name":"Sam Griffiths-Jones","orcid":"0000-0001-6043-807X","position":30,"is_corresponding":false},{"id":178244,"name":"Bastian Fromm","orcid":"0000-0003-0352-3037","position":31,"is_corresponding":false},{"id":228806,"name":"Kevin J. Peterson","orcid":"0000-0002-6780-0964","position":32,"is_corresponding":false},{"id":106793,"name":"Ioanna Kalvari","orcid":"0000-0001-9424-9197","position":33,"is_corresponding":false},{"id":17117,"name":"Eric P. Nawrocki","orcid":"0000-0002-2497-3427","position":34,"is_corresponding":false},{"id":228807,"name":"Anton S. Petrov","orcid":"0000-0003-3359-7299","position":35,"is_corresponding":false},{"id":86953,"name":"Shuai Weng","orcid":"0000-0003-4233-0772","position":36,"is_corresponding":false},{"id":229919,"name":"Philia Bouchard-Bourelle","orcid":null,"position":37,"is_corresponding":false},{"id":228808,"name":"Michelle S. Scott","orcid":"0000-0001-6231-7714","position":38,"is_corresponding":false},{"id":228809,"name":"Lauren Michelle Lui","orcid":"0000-0001-8720-5268","position":39,"is_corresponding":false},{"id":228810,"name":"David Hoksza","orcid":"0000-0003-4679-0557","position":40,"is_corresponding":false},{"id":103829,"name":"Ruth C. Lovering","orcid":"0000-0002-9791-0064","position":41,"is_corresponding":false},{"id":103830,"name":"Barbara Kramarz","orcid":"0000-0002-3898-1727","position":42,"is_corresponding":false},{"id":229920,"name":"Prita Mani","orcid":null,"position":43,"is_corresponding":false},{"id":228811,"name":"Sridhar Ramachandran","orcid":"0000-0002-2246-3722","position":44,"is_corresponding":false},{"id":106798,"name":"Zasha Weinberg","orcid":"0000-0002-6681-3624","position":45,"is_corresponding":false},{"id":228790,"name":"Blake Sweeney","orcid":"0000-0002-6497-2883","position":0,"is_corresponding":true}],"reference_count":51,"raw_metadata":null,"created_at":"2026-07-18T20:34:12.647743Z","pmid":"33106848","pmcid":"PMC7779037","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":55.5556,"fair_a":37.5,"fair_i":40.0,"fair_r":25.0,"fair_zscore":0.2857,"fair_rationale":{"fair_score":41.67,"has_llm":true,"taxonomy_version":"fair_taxonomy_v5","dimensions":{"F":{"name":"Findable","score":55.56,"criteria":[{"key":"f_dataset_pid","label":"Persistent identifier for the data","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"RNAcentral is freely available at https://rnacentral.org.","grounded":true,"rationale":"The identifier given is a web URL starting with https, not a persistent identifier scheme (DOI, Handle, ARK, etc.). [majority verdict 'partial' (3/5 passes agreed)]","anchors":["RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit","RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier'","FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'"],"scored":true,"signal":null},{"key":"f_repository_named","label":"Named repository","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"RNAcentral is a comprehensive database of non-coding RNA (ncRNA) sequences","grounded":true,"rationale":"RNAcentral is named as the database holding the data. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed (","NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived","NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten"],"scored":true,"signal":null},{"key":"f_data_availability_statement","label":"Data-availability statement","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"All data are freely available at https://rnacentral.org.","grounded":false,"rationale":"The data availability statement provides a link to a public repository (RNAcentral), which corresponds to Colavizza category 3. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (4/5 passes agreed)]","anchors":["Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li","Springer Nature research data policy — Data Availability Statements: standard statement templat","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes"],"scored":false,"signal":null},{"key":"f_discovery_metadata","label":"Description of the dataset as an object","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"RNAcentral is a comprehensive database of non-coding RNA (ncRNA) sequences that provides a single access point to 44 RNA resources and >18 million ncRNA sequences from a wide range of organisms and RNA types.","grounded":true,"rationale":"The dataset's content is described in running prose, not in an itemised inventory (section, table, or list). [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential)","FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability'","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'"],"scored":false,"signal":null},{"key":"f_dataset_cited","label":"Dataset formally cited","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":"All data are freely available at https://rnacentral.org.","grounded":false,"rationale":"The dataset's identifier (URL) appears only in the body text, not as a reference-list entry. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (3/5 passes agreed)]","anchors":["FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first-","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes","FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'"],"scored":true,"signal":null}]},"A":{"name":"Accessible","score":37.5,"criteria":[{"key":"a_data_openly_accessible","label":"Access route free of preconditions","kind":"llm","weight":2.0,"fraction":0.5,"verdict":"partial","evidence":"All data are freely available at https://rnacentral.org.","grounded":false,"rationale":"The text gives a route to the data with no stated precondition; the data are stated to be freely available now. [downgraded to 'partial' — no verifiable quote from the paper]","anchors":["RDA-A1.1-01D — 'Data is accessible through a free access protocol'","FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'"],"scored":true,"signal":null},{"key":"a_access_conditions_stated","label":"Access level labelled","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"All data are freely available at https://rnacentral.org.","grounded":false,"rationale":"The paper states 'All data are freely available' which is a natural-language synonym for 'open access'. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (4/5 passes agreed)]","anchors":["FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","RDA-A1-01M — metadata contains information to enable the user to get access to the data","COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl"],"scored":false,"signal":null},{"key":"a_controlled_access_for_sensitive","label":"Gatekeeper for sensitive data","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":"All data are freely available at https://rnacentral.org.","grounded":false,"rationale":"The data are non-sensitive RNA sequences, and no gatekeeper is named; the freely available statement implies no access control.","anchors":["NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee","RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and ","NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse"],"scored":false,"signal":null},{"key":"a_timeline_retention","label":"Availability timing & retention","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":"All data are freely available at https://rnacentral.org.","grounded":false,"rationale":"The paper states that the data are currently available but provides no persistence commitment or retention period. [downgraded to 'no' — no verifiable quote from the paper]","anchors":["NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines","NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy'","RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'"],"scored":false,"signal":null}]},"I":{"name":"Interoperable","score":40.0,"criteria":[{"key":"i_open_nonproprietary_format","label":"Open file format","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":"The data can be accessed in the FTP archive, as well as through an API and a public Postgres database","grounded":false,"rationale":"No specific file format (e.g., FASTA, CSV) is named for the released data; only access methods are mentioned.","anchors":["FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co","RDA-R1.3-02D — data is expressed in a machine-understandable community standard","RDA-I1-01D — data uses a knowledge representation expressed in a standardised format"],"scored":true,"signal":null},{"key":"i_community_standard_vocabulary","label":"Community standard / vocabulary","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"all RNAcentral sequences have been annotated with Sequence Ontology terms.","grounded":true,"rationale":"The paper names Sequence Ontology (SO), a community standard ontology for RNA types.","anchors":["RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential)","RDA-R1.3-01D — 'Data complies with a community standard'","RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'"],"scored":false,"signal":null},{"key":"i_qualified_references","label":"Identifiers for the resources the data depend on","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":"The code is available at https://github.com/rnacentral under the Apache 2.0 license.","grounded":false,"rationale":"The paper provides a GitHub URL for the code, which is a machine-resolvable locator for a resource other than the dataset. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (4/5 passes agreed)]","anchors":["RDA-I3-01M — '(meta)data include references to other (meta)data'","RDA-I3-03M — 'metadata includes qualified references to other metadata'","FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'"],"scored":false,"signal":null}]},"R":{"name":"Reusable","score":25.0,"criteria":[{"key":"r_reuse_license","label":"Reuse licence","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No license is named for the data; the CC BY license applies to the article, not the data.","anchors":["RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu","RDA-R1.1-02M — 'Metadata refers to a standard reuse licence'","RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'"],"scored":true,"signal":null},{"key":"r_provenance_methods","label":"Provenance of the data","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"The R2DT software automatically selects the best-matching template from a library of 36322D templates","grounded":false,"rationale":"The paper names specific software (R2DT, nhmmer, etc.) used to produce the data. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa","FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati","W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance"],"scored":false,"signal":null},{"key":"r_documentation_codebook","label":"Documentation / codebook","kind":"llm","weight":1.0,"fraction":0.0,"verdict":"no","evidence":"Table 1. Sixteen new member databases incorporated into RNAcentral in releases 11–16","grounded":false,"rationale":"Variable-level definitions are given inside the article (Table 1), not as a separate documentation object shipped with the data. [downgraded to 'no' — no verifiable quote from the paper]","anchors":["RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'","NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t"],"scored":false,"signal":null},{"key":"r_versioning","label":"Snapshot identified","kind":"llm","weight":0.5,"fraction":0.5,"verdict":"partial","evidence":"In the most recent release (version 16), we generated >13 million 2D structure diagrams.","grounded":false,"rationale":"A version token ('version 16') is stated for the dataset. [downgraded to 'partial' — no verifiable quote from the paper]","anchors":["DataCite Metadata Schema 4.6 — the 'Version' property","RDA-R1.2-01M — provenance information (which version was used is provenance)","NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'"],"scored":true,"signal":null},{"key":"x_code_availability","label":"Analysis code available","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"The code is available at https://github.com/rnacentral under the Apache 2.0 license.","grounded":false,"rationale":"A machine-resolvable code repository URL is given for the study's own code. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (4/5 passes agreed)]","anchors":["NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code'","FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear","FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)"],"scored":true,"signal":null},{"key":"x_funding_attribution","label":"Funder and award number","kind":"llm","weight":0.5,"fraction":0.5,"verdict":"partial","evidence":"Biotechnology and Biological Sciences Research Council (BBSRC) [BB/N019199/1]; Wellcome Trust [218302/Z/19/Z, 208349/Z/17/Z]; National Institutes of Health [U24HG003345, U41HG000739]; Charles University [SVV 260588].","grounded":false,"rationale":"Award/grant numbers are provided for each named funder. [downgraded to 'partial' — no verifiable quote from the paper]","anchors":["DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award","Crossref Funder Registry — canonical funder identifiers for funding metadata","RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco"],"scored":true,"signal":null}]}},"actions":[{"key":"r_reuse_license","dimension":"R","label":"Reuse licence","action":"Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No license is named for the data; the CC BY license applies to the article, not the data.","gain":16.67,"priority":"essential","scored":true},{"key":"f_dataset_pid","dimension":"F","label":"Persistent identifier for the data","action":"Mint or cite a persistent identifier for the dataset — a repository DOI or an accession from a registered repository — and print it in the paper. A bare URL is not persistent: it is the single most common cause of a dead data link five years after publication. For genomics / sequencing data, deposit in GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"RNAcentral is freely available at https://rnacentral.org.","why":"The identifier given is a web URL starting with https, not a persistent identifier scheme (DOI, Handle, ARK, etc.). [majority verdict 'partial' (3/5 passes agreed)]","gain":8.33,"priority":"essential","scored":true},{"key":"a_data_openly_accessible","dimension":"A","label":"Access route free of preconditions","action":"Remove the precondition or justify it. Release the data at publication with no embargo, no registration wall, and no approval step — NIH's zero-embargo public- access rule (NOT-OD-25-101) has already made 'available at publication' the federal baseline for the article; the data should not lag behind it. For genomics / sequencing data, deposit in GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"All data are freely available at https://rnacentral.org.","why":"The text gives a route to the data with no stated precondition; the data are stated to be freely available now. [downgraded to 'partial' — no verifiable quote from the paper]","gain":8.33,"priority":"essential","scored":true},{"key":"f_dataset_cited","dimension":"F","label":"Dataset formally cited","action":"Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the genomics / sequencing repository accession (e.g. from GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA)) in the reference list.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"All data are freely available at https://rnacentral.org.","why":"The dataset's identifier (URL) appears only in the body text, not as a reference-list entry. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (3/5 passes agreed)]","gain":8.33,"priority":"important","scored":true},{"key":"i_open_nonproprietary_format","dimension":"I","label":"Open file format","action":"Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable. Prefer open genomics / sequencing formats such as FASTQ, BAM or VCF.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"The data can be accessed in the FTP archive, as well as through an API and a public Postgres database","why":"No specific file format (e.g., FASTA, CSV) is named for the released data; only access methods are mentioned.","gain":8.33,"priority":"important","scored":true},{"key":"x_code_availability","dimension":"R","label":"Analysis code available","action":"Publish the analysis code in a public forge, archive a tagged release with a DOI (Zenodo/Software Heritage), and cite that DOI in the paper. NIH DMS Element 2 asks for the tools and code, not only the data — and 'available on request' is not a locator. Archive the analysis code in a versioned repository (GitHub + a Zenodo release DOI).","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"The code is available at https://github.com/rnacentral under the Apache 2.0 license.","why":"A machine-resolvable code repository URL is given for the study's own code. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (4/5 passes agreed)]","gain":4.17,"priority":"important","scored":true},{"key":"r_versioning","dimension":"R","label":"Snapshot identified","action":"Version the deposit and cite the exact version analysed (a version-specific DOI, or an accession with its version suffix). A reader reproducing your work against 'the current release' is reproducing it against a different dataset.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"In the most recent release (version 16), we generated >13 million 2D structure diagrams.","why":"A version token ('version 16') is stated for the dataset. [downgraded to 'partial' — no verifiable quote from the paper]","gain":2.08,"priority":"useful","scored":true},{"key":"x_funding_attribution","dimension":"R","label":"Funder and award number","action":"State the funder AND the award number in the paper, and put them in the dataset's FundingReference metadata. A funder name alone cannot be linked back to the award, so the funding provenance of the data is lost the moment the paper is indexed.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Biotechnology and Biological Sciences Research Council (BBSRC) [BB/N019199/1]; Wellcome Trust [218302/Z/19/Z, 208349/Z/17/Z]; National Institutes of Health [U24HG003345, U41HG000739]; Charles University [SVV 260588].","why":"Award/grant numbers are provided for each named funder. [downgraded to 'partial' — no verifiable quote from the paper]","gain":2.08,"priority":"useful","scored":true},{"key":"f_data_availability_statement","dimension":"F","label":"Data-availability statement","action":"Replace the statement with the repository template: name the repository and give the accession or DOI (Colavizza category 3). This is the only DAS class associated with a measured citation advantage; 'available on reasonable request' and 'within the article' are not.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"All data are freely available at https://rnacentral.org.","why":"The data availability statement provides a link to a public repository (RNAcentral), which corresponds to Colavizza category 3. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (4/5 passes agreed)]","gain":0.0,"priority":"essential","scored":false},{"key":"f_discovery_metadata","dimension":"F","label":"Description of the dataset as an object","action":"Add a 'Data Records' section: itemise every file in the deposit and every variable or sample it holds, with counts and units. Describe the dataset as an object in its own right, not as a by-product of the findings — this is what makes it discoverable to someone who is not looking for your paper.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"RNAcentral is a comprehensive database of non-coding RNA (ncRNA) sequences that provides a single access point to 44 RNA resources and >18 million ncRNA sequences from a wide range of organisms and RNA types.","why":"The dataset's content is described in running prose, not in an itemised inventory (section, table, or list). [majority verdict 'partial' (4/5 passes agreed)]","gain":0.0,"priority":"essential","scored":false},{"key":"a_access_conditions_stated","dimension":"A","label":"Access level labelled","action":"State the access level in words, using the standard vocabulary: 'These data are open access' / 'These data are controlled access'. A reader — and a harvester — should not have to infer the access level from the presence of a download link.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"All data are freely available at https://rnacentral.org.","why":"The paper states 'All data are freely available' which is a natural-language synonym for 'open access'. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (4/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"r_provenance_methods","dimension":"R","label":"Provenance of the data","action":"Name the instruments, kits, and software — with versions — that produced the data, not just the verbs. 'Reads were aligned' is not provenance; 'aligned with STAR v2.7.9a to GRCh38' is, because someone else can rerun it.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"The R2DT software automatically selects the best-matching template from a library of 36322D templates","why":"The paper names specific software (R2DT, nhmmer, etc.) used to produce the data. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (4/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"r_documentation_codebook","dimension":"R","label":"Documentation / codebook","action":"Ship a README and a data dictionary IN the deposit — every file, every variable, its units, its allowed values, its missing-value codes. It is the cheapest single thing that makes a dataset usable by someone who was not in the lab, and a table buried in the article does not travel with the data.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"Table 1. Sixteen new member databases incorporated into RNAcentral in releases 11–16","why":"Variable-level definitions are given inside the article (Table 1), not as a separate documentation object shipped with the data. [downgraded to 'no' — no verifiable quote from the paper]","gain":0.0,"priority":"important","scored":false},{"key":"a_controlled_access_for_sensitive","dimension":"A","label":"Gatekeeper for sensitive data","action":"Route sensitive data through an institutional gatekeeper — deposit in a controlled- access repository (dbGaP, EGA) with a Data Access Committee and a published DUA — rather than through the corresponding author's inbox. An author-gated dataset dies with the author's email address, and 'on reasonable request' has been shown repeatedly not to yield data. For sensitive/human genomics / sequencing data, use a controlled-access repository such as dbGaP or EGA.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"All data are freely available at https://rnacentral.org.","why":"The data are non-sensitive RNA sequences, and no gatekeeper is named; the freely available statement implies no access control.","gain":0.0,"priority":"useful","scored":false},{"key":"i_qualified_references","dimension":"I","label":"Identifiers for the resources the data depend on","action":"Cite by identifier every resource the data depend on — the source datasets' accessions, the reference build (GRCh38 / GCA_000001405.28), the cohort application number, the code DOI — and register those relations on the dataset record (IsDerivedFrom, IsSupplementTo). A name is not a link: it cannot be resolved, versioned, or followed by a machine.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"The code is available at https://github.com/rnacentral under the Apache 2.0 license.","why":"The paper provides a GitHub URL for the code, which is a machine-resolvable locator for a resource other than the dataset. [downgraded to 'no' — no verifiable quote from the paper] [majority verdict 'no' (4/5 passes agreed)]","gain":0.0,"priority":"useful","scored":false},{"key":"a_timeline_retention","dimension":"A","label":"Availability timing & retention","action":"State when the data become available AND how long they will be retained — cite the repository's preservation policy. NIH DMS Element 4 asks for both; most papers give neither.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":"All data are freely available at https://rnacentral.org.","why":"The paper states that the data are currently available but provides no persistence commitment or retention period. [downgraded to 'no' — no verifiable quote from the paper]","gain":0.0,"priority":"useful","scored":false}],"suggestions":["Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","Mint or cite a persistent identifier for the dataset — a repository DOI or an accession from a registered repository — and print it in the paper. A bare URL is not persistent: it is the single most common cause of a dead data link five years after publication. For genomics / sequencing data, deposit in GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA).","Remove the precondition or justify it. Release the data at publication with no embargo, no registration wall, and no approval step — NIH's zero-embargo public- access rule (NOT-OD-25-101) has already made 'available at publication' the federal baseline for the article; the data should not lag behind it. For genomics / sequencing data, deposit in GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA).","Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the genomics / sequencing repository accession (e.g. from GEO (GSE accession), SRA (SRP/SRR) or ENA/BioProject (PRJEB/PRJNA)) in the reference list.","Release the data in an open, community-standard format (CSV/TSV, JSON, HDF5, NetCDF, FASTQ, VCF, NIfTI…) instead of — or alongside — any proprietary or instrument-native format, and name the format in the paper. A dataset that needs a €2,000 licence to open is not reusable. Prefer open genomics / sequencing formats such as FASTQ, BAM or VCF."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v8","fulltext_source":"unpaywall_pdf"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v8","fair_fulltext_source":"unpaywall_pdf","fair_has_llm":true,"fair_computed_at":"2026-07-20T10:49:14.560052Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}