{"doi":"10.1016/j.crmeth.2024.100964","title":"Unifying community whole-brain imaging datasets enables robust neuron identification and reveals determinants of neuron position in C. elegans","abstract":"We develop a data harmonization approach for C. elegans volumetric microscopy data, consisting of a standardized format, pre-processing techniques, and human-in-the-loop machine-learning-based analysis tools. Using this approach, we unify a diverse collection of 118 whole-brain neural activity imaging datasets from five labs, storing these and accompanying tools in an online repository WormID (wormid.org). With this repository, we train three existing automated cell-identification algorithms, CPD, StatAtlas, and CRF_ID, to enable accuracy that generalizes across labs, recovering all human-labeled neurons in some cases. We mine this repository to identify factors that influence the developmental positioning of neurons. This growing resource of data, code, apps, and tutorials enables users to (1) study neuroanatomical organization and neural activity across diverse experimental paradigms, (2) develop and benchmark algorithms for automated neuron detection, segmentation, cell identification, tracking, and activity extraction, and (3) share data with the community and comply with data-sharing policies.","journal":"Cell Reports Methods","year":2025,"id":538517,"datarank":0.21949660717731603,"base_score":1.3862943611198906,"endowment":1.3862943611198906,"self_citation_contribution":0.20794415416798362,"citation_network_contribution":0.011552453009332421,"self_endowment_contribution":0.20794415416798362,"citer_contribution":0.011552453009332421,"corpus_percentile":37.24762125783244,"corpus_rank":8112,"citation_count":3,"citer_count":2,"citers_with_citation_signal":1,"citers_with_endowment":1,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9381,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":79.1667,"fair_percentile":97.67655151329869,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1329580,"name":"Kevin Rusch","orcid":"0000-0001-9897-9242","position":1,"is_corresponding":false},{"id":572482,"name":"Raymond L. Dunn","orcid":"0000-0003-4443-5519","position":2,"is_corresponding":false},{"id":572481,"name":"Jackson Borchardt","orcid":"0000-0001-9033-6365","position":3,"is_corresponding":false},{"id":1329922,"name":"Steven Ban","orcid":null,"position":4,"is_corresponding":false},{"id":1329581,"name":"Greg Bubnis","orcid":"0000-0002-9397-8977","position":5,"is_corresponding":false},{"id":1329923,"name":"Grace C. Chiu","orcid":null,"position":6,"is_corresponding":false},{"id":630407,"name":"Chentao Wen","orcid":"0000-0002-8609-476X","position":7,"is_corresponding":false},{"id":1090985,"name":"Ryoga Suzuki","orcid":null,"position":8,"is_corresponding":false},{"id":564368,"name":"Shivesh Chaudhary","orcid":"0000-0002-1928-0933","position":9,"is_corresponding":false},{"id":1024083,"name":"Hyun Jee Lee","orcid":"0000-0001-9662-2063","position":10,"is_corresponding":false},{"id":1176393,"name":"Zikai Yu","orcid":"0000-0001-9377-9196","position":11,"is_corresponding":false},{"id":59111,"name":"Ben Dichter","orcid":"0000-0001-5725-6910","position":12,"is_corresponding":false},{"id":805314,"name":"Ryan Ly","orcid":"0000-0001-9238-0642","position":13,"is_corresponding":false},{"id":52926,"name":"Shuichi Onami","orcid":"0000-0002-8255-1724","position":14,"is_corresponding":false},{"id":302031,"name":"Hang Lu","orcid":"0000-0002-6881-660X","position":15,"is_corresponding":false},{"id":630416,"name":"Koutarou D. Kimura","orcid":"0000-0002-3359-1578","position":16,"is_corresponding":false},{"id":233150,"name":"Eviatar Yemini","orcid":"0000-0003-1977-0761","position":17,"is_corresponding":false},{"id":293047,"name":"Saul Kato","orcid":"0000-0003-2990-8306","position":18,"is_corresponding":false},{"id":295467,"name":"Daniel Sprague","orcid":"0009-0001-7773-2280","position":0,"is_corresponding":true}],"reference_count":57,"raw_metadata":null,"created_at":"2026-07-19T02:52:21.389196Z","pmid":"39826553","pmcid":"PMC11840940","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":94.4444,"fair_a":87.5,"fair_i":80.0,"fair_r":58.3333,"fair_zscore":1.77,"fair_rationale":{"fair_score":79.17,"has_llm":true,"taxonomy_version":"fair_taxonomy_v5","dimensions":{"F":{"name":"Findable","score":94.44,"criteria":[{"key":"f_dataset_pid","label":"Persistent identifier for the data","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"https://doi.org/10.48324/dandi.000715/0.241009.1514","grounded":true,"rationale":"The paper provides DOIs for the datasets, which are persistent identifiers. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-F1-01D — FAIR Data Maturity Model: 'Data is identified by a persistent identifier' (priorit","RDA-F1-02D — FAIR Data Maturity Model: 'Data is identified by a globally unique identifier'","FsF-F1-02D — F-UJI/FAIRsFAIR: 'Data is assigned a persistent identifier'"],"scored":true,"signal":null},{"key":"f_repository_named","label":"Named repository","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"This corpus is stored in a popular community archive called the Distributed Archives for Neurophysiology Data Integration (DANDI), which serves as a free, persistent, open access repository for experimental neuroscience data from a variety of model organisms.","grounded":true,"rationale":"The paper names DANDI as the repository holding the data.","anchors":["RDA-F4-01M — FAIR Data Maturity Model: metadata is offered so it can be harvested and indexed (","NIH DMS Policy Element 4 (NOT-OD-21-014) — name the repository where data will be archived","NSTC Desirable Characteristics of Data Repositories (2022) — 'Long-Term Sustainability', 'Reten"],"scored":true,"signal":null},{"key":"f_data_availability_statement","label":"Data-availability statement","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"All datasets are available on the DANDI archive, with URLs provided in STAR Methods.","grounded":true,"rationale":"The statement points to a repository record (DANDI archive with URLs).","anchors":["Colavizza, Hrynaszkiewicz, Staden, Whitaker & McGillivray (2020), 'The citation advantage of li","Springer Nature research data policy — Data Availability Statements: standard statement templat","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes"],"scored":false,"signal":null},{"key":"f_discovery_metadata","label":"Description of the dataset as an object","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"Table 1. Summary of aggregated dataset characteristics","grounded":true,"rationale":"The paper includes a table that itemises the dataset characteristics, providing an inventory of the data. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-F2-01M — 'Rich metadata is provided to allow discovery' (priority Essential)","FsF-F2-01M — F-UJI: 'Metadata includes descriptive core elements to support data findability'","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'"],"scored":false,"signal":null},{"key":"f_dataset_cited","label":"Dataset formally cited","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"https://doi.org/10.48324/dandi.000715/0.241009.1514","grounded":true,"rationale":"The dataset identifier appears in the body text (STAR Methods table), not in the reference list.","anchors":["FORCE11 Joint Declaration of Data Citation Principles (2014) — data should be cited as a first-","RDA-F3-01M — metadata clearly and explicitly includes the identifier of the data it describes","FsF-F3-01M — F-UJI: 'Metadata includes the identifier of the data it describes'"],"scored":true,"signal":null}]},"A":{"name":"Accessible","score":87.5,"criteria":[{"key":"a_data_openly_accessible","label":"Access route free of preconditions","kind":"llm","weight":2.0,"fraction":1.0,"verdict":"yes","evidence":"free, persistent, open access repository","grounded":true,"rationale":"The paper states the data are stored in an open access repository with no precondition.","anchors":["RDA-A1.1-01D — 'Data is accessible through a free access protocol'","FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","NSTC Desirable Characteristics of Data Repositories (2022) — 'Free and Easy Access'"],"scored":true,"signal":null},{"key":"a_access_conditions_stated","label":"Access level labelled","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"free, persistent, open access repository","grounded":true,"rationale":"The paper describes the DANDI archive as an open access repository, which is an explicit access-level label. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["FsF-A1-01M — F-UJI: 'Metadata contains access level and access conditions of the data'","RDA-A1-01M — metadata contains information to enable the user to get access to the data","COAR Controlled Vocabularies — Access Rights v1.0 (open / embargoed / restricted / metadata-onl"],"scored":false,"signal":null},{"key":"a_controlled_access_for_sensitive","label":"Gatekeeper for sensitive data","kind":"llm","weight":0.5,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"The data are from C. elegans, not human subjects, so no gatekeeper is required or named.","anchors":["NIH Genomic Data Sharing Policy (NOT-OD-14-124) — controlled-access via a Data Access Committee","RDA-A1.2-01D — 'Data is accessible through an access protocol that supports authentication and ","NIH DMS Policy Element 5 (NOT-OD-21-014) — Access, Distribution, or Reuse Considerations (conse"],"scored":false,"signal":null},{"key":"a_timeline_retention","label":"Availability timing & retention","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"free, persistent, open access repository","grounded":true,"rationale":"The repository is described as persistent, which implies a retention commitment. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["NIH DMS Plan Element 4 (NOT-OD-21-014) — Data Preservation, Access, and Associated Timelines","NSTC Desirable Characteristics (2022), Organizational Infrastructure: 'Retention Policy'","RDA-A2-01M — 'Metadata is guaranteed to remain available after data is no longer available'"],"scored":false,"signal":null}]},"I":{"name":"Interoperable","score":80.0,"criteria":[{"key":"i_open_nonproprietary_format","label":"Open file format","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"All associated raw data and metadata is stored in the standardized NWB 41 file format","grounded":true,"rationale":"NWB is an open, community-standard format based on HDF5, which is listed as an open format. [majority verdict 'yes' (3/5 passes agreed)]","anchors":["FsF-R1.3-02D — F-UJI: 'Data is available in a file format recommended by the target research co","RDA-R1.3-02D — data is expressed in a machine-understandable community standard","RDA-I1-01D — data uses a knowledge representation expressed in a standardised format"],"scored":true,"signal":null},{"key":"i_community_standard_vocabulary","label":"Community standard / vocabulary","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"standardized NWB41 file format","grounded":false,"rationale":"NWB is a community standard for neurophysiology data, registered in FAIRsharing. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]","anchors":["RDA-R1.3-01M — 'Metadata complies with a community standard' (priority Essential)","RDA-R1.3-01D — 'Data complies with a community standard'","RDA-I2-01M — '(Meta)data use vocabularies that follow FAIR principles'"],"scored":false,"signal":null},{"key":"i_qualified_references","label":"Identifiers for the resources the data depend on","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"https://doi.org/10.1016/j.cub.2023.04.071","grounded":true,"rationale":"The paper provides a DOI for an external synaptic connectivity dataset used in the study.","anchors":["RDA-I3-01M — '(meta)data include references to other (meta)data'","RDA-I3-03M — 'metadata includes qualified references to other metadata'","FsF-I3-01M — F-UJI: 'Metadata includes links between the data and its related entities'"],"scored":false,"signal":null}]},"R":{"name":"Reusable","score":58.33,"criteria":[{"key":"r_reuse_license","label":"Reuse licence","kind":"llm","weight":2.0,"fraction":0.0,"verdict":"no","evidence":null,"grounded":false,"rationale":"No license for the data is stated in the paper.","anchors":["RDA-R1.1-01M — 'Metadata includes information about the licence under which the data can be reu","RDA-R1.1-02M — 'Metadata refers to a standard reuse licence'","RDA-R1.1-03M — 'Metadata refers to a machine-understandable reuse licence'"],"scored":true,"signal":null},{"key":"r_provenance_methods","label":"Provenance of the data","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"Zeiss LSM880 spinning disk confocal","grounded":true,"rationale":"The paper names specific instruments (e.g., Zeiss LSM880) used to produce the data. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["RDA-R1.2-01M — 'Metadata includes provenance information according to community- specific standa","FsF-R1.2-01M — F-UJI: 'Metadata includes provenance information about data creation or generati","W3C PROV-O (W3C Recommendation, 2013) — the entity/activity/agent model of provenance"],"scored":false,"signal":null},{"key":"r_documentation_codebook","label":"Documentation / codebook","kind":"llm","weight":1.0,"fraction":0.5,"verdict":"partial","evidence":"Table 1. Summary of aggregated dataset characteristics","grounded":true,"rationale":"The table defines the data characteristics within the article, but no separate documentation object is shipped with the data. [majority verdict 'partial' (4/5 passes agreed)]","anchors":["RDA-R1-01M — '(Meta)data are richly described with a plurality of accurate and relevant attribu","FsF-R1-01MD — F-UJI: 'Metadata specifies the content of the data'","NIH DMS Policy Element 3 (NOT-OD-21-014) — Standards (documentation and metadata to accompany t"],"scored":false,"signal":null},{"key":"r_versioning","label":"Snapshot identified","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"0.241009.1514","grounded":true,"rationale":"The dataset DOI includes a version string, indicating the snapshot.","anchors":["DataCite Metadata Schema 4.6 — the 'Version' property","RDA-R1.2-01M — provenance information (which version was used is provenance)","NSTC Desirable Characteristics of Data Repositories (2022) — 'Provenance', 'Retention Policy'"],"scored":true,"signal":null},{"key":"x_code_availability","label":"Analysis code available","kind":"llm","weight":1.0,"fraction":1.0,"verdict":"yes","evidence":"https://doi.org/10.5281/zenodo.13910335","grounded":true,"rationale":"The paper provides a DOI for the code repository. [majority verdict 'yes' (4/5 passes agreed)]","anchors":["NIH DMS Policy Element 2 (NOT-OD-21-014) — 'Related Tools, Software and/or Code'","FAIR4RS Principles v1.0 (Chue Hong et al., 2022; RDA/FORCE11/ReSA) — FAIR Principles for Resear","FORCE11 Software Citation Principles (Smith, Katz & Niemeyer, 2016, PeerJ CS 2:e86)"],"scored":true,"signal":null},{"key":"x_funding_attribution","label":"Funder and award number","kind":"llm","weight":0.5,"fraction":1.0,"verdict":"yes","evidence":"R35GM124735","grounded":true,"rationale":"The paper includes a specific grant number (R35GM124735) from NIH.","anchors":["DataCite Metadata Schema 4.6 — 'FundingReference' property (funderName, funderIdentifier, award","Crossref Funder Registry — canonical funder identifiers for funding metadata","RDA-F2-01M — rich metadata provided to allow discovery (funding is part of the descriptive reco"],"scored":true,"signal":null}]}},"actions":[{"key":"r_reuse_license","dimension":"R","label":"Reuse licence","action":"Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"No license for the data is stated in the paper.","gain":16.67,"priority":"essential","scored":true},{"key":"f_dataset_cited","dimension":"F","label":"Dataset formally cited","action":"Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the neuroimaging repository accession (e.g. from OpenNeuro or NeuroVault) in the reference list.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"https://doi.org/10.48324/dandi.000715/0.241009.1514","why":"The dataset identifier appears in the body text (STAR Methods table), not in the reference list.","gain":4.17,"priority":"important","scored":true},{"key":"i_community_standard_vocabulary","dimension":"I","label":"Community standard / vocabulary","action":"Adopt and NAME your domain's data standard — the minimum-information checklist, metadata schema, or ontology your community uses (MIAME/MINSEQE, ISA-Tab, BIDS, an OBO ontology, HL7 FHIR/OMOP) — and say which one you followed. A reporting checklist standardises your paper; it does nothing for your data. In neuroimaging, describe the data with BIDS, NIfTI or DICOM.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"standardized NWB41 file format","why":"NWB is a community standard for neurophysiology data, registered in FAIRsharing. [downgraded to 'partial' — no verifiable quote from the paper] [majority verdict 'partial' (3/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"r_documentation_codebook","dimension":"R","label":"Documentation / codebook","action":"Ship a README and a data dictionary IN the deposit — every file, every variable, its units, its allowed values, its missing-value codes. It is the cheapest single thing that makes a dataset usable by someone who was not in the lab, and a table buried in the article does not travel with the data.","anchors":["yes","partial","no"],"verdict":"partial","current":0.5,"evidence":"Table 1. Summary of aggregated dataset characteristics","why":"The table defines the data characteristics within the article, but no separate documentation object is shipped with the data. [majority verdict 'partial' (4/5 passes agreed)]","gain":0.0,"priority":"important","scored":false},{"key":"a_controlled_access_for_sensitive","dimension":"A","label":"Gatekeeper for sensitive data","action":"Route sensitive data through an institutional gatekeeper — deposit in a controlled- access repository (dbGaP, EGA) with a Data Access Committee and a published DUA — rather than through the corresponding author's inbox. An author-gated dataset dies with the author's email address, and 'on reasonable request' has been shown repeatedly not to yield data.","anchors":["yes","partial","no"],"verdict":"no","current":0.0,"evidence":null,"why":"The data are from C. elegans, not human subjects, so no gatekeeper is required or named.","gain":0.0,"priority":"useful","scored":false}],"suggestions":["Attach a standard, machine-readable open licence to the deposit — CC0 or CC BY, which is what Horizon Europe and most funders expect — and print the licence identifier in the paper. 'Free to use' is not a licence: it grants nothing a reuser's institution can rely on.","Cite the dataset in the reference list like a publication — creator, year, title, repository, DOI/accession — and cite it in-text where it is used. Only a reference- list entry is machine-readable to Crossref/DataCite, and only a citation lets the data earn credit. Cite the neuroimaging repository accession (e.g. from OpenNeuro or NeuroVault) in the reference list.","Adopt and NAME your domain's data standard — the minimum-information checklist, metadata schema, or ontology your community uses (MIAME/MINSEQE, ISA-Tab, BIDS, an OBO ontology, HL7 FHIR/OMOP) — and say which one you followed. A reporting checklist standardises your paper; it does nothing for your data. In neuroimaging, describe the data with BIDS, NIfTI or DICOM.","Ship a README and a data dictionary IN the deposit — every file, every variable, its units, its allowed values, its missing-value codes. It is the cheapest single thing that makes a dataset usable by someone who was not in the lab, and a table buried in the article does not travel with the data.","Route sensitive data through an institutional gatekeeper — deposit in a controlled- access repository (dbGaP, EGA) with a Data Access Committee and a published DUA — rather than through the corresponding author's inbox. An author-gated dataset dies with the author's email address, and 'on reasonable request' has been shown repeatedly not to yield data."],"model":"deepseek/deepseek-v4-flash","agent_version":"fair_agent_v8","fulltext_source":"unpaywall_pdf"},"fair_model":"deepseek/deepseek-v4-flash","fair_agent_version":"fair_agent_v8","fair_fulltext_source":"unpaywall_pdf","fair_has_llm":true,"fair_computed_at":"2026-07-20T13:17:45.142960Z","clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}