{"doi":"10.1093/nar/gkae1142","title":"CZ CELLxGENE Discover: a single-cell data platform for scalable exploration, analysis and modeling of aggregated data","abstract":"<jats:title>Abstract</jats:title>\n               <jats:p>Hundreds of millions of single cells have been analyzed using high-throughput transcriptomic methods. The cumulative knowledge within these datasets provides an exciting opportunity for unlocking insights into health and disease at the level of single cells. Meta-analyses that span diverse datasets building on recent advances in large language models and other machine-learning approaches pose exciting new directions to model and extract insight from single-cell data. Despite the promise of these and emerging analytical tools for analyzing large amounts of data, the sheer number of datasets, data models and accessibility remains a challenge. Here, we present CZ CELLxGENE Discover (cellxgene.cziscience.com), a data platform that provides curated and interoperable single-cell data. Available via a free-to-use online data portal, CZ CELLxGENE hosts a growing corpus of community-contributed data of over 93 million unique cells. Curated, standardized and associated with consistent cell-level metadata, this collection of single-cell transcriptomic data is the largest of its kind and growing rapidly via community contributions. A suite of tools and features enables accessibility and reusability of the data via both computational and visual interfaces to allow researchers to explore individual datasets, perform cross-corpus analysis, and run meta-analyses of tens of millions of cells across studies and tissues at the resolution of single cells.</jats:p>","journal":"Nucleic Acids Research","year":2025,"id":588212,"datarank":3.6326118434156127,"base_score":5.8664680569332965,"endowment":5.8664680569332965,"self_citation_contribution":0.8799702085399946,"citation_network_contribution":2.752641634875618,"self_endowment_contribution":0.8799702085399946,"citer_contribution":2.752641634875618,"corpus_percentile":93.83460973156959,"corpus_rank":798,"citation_count":352,"citer_count":100,"citers_with_citation_signal":100,"citers_with_endowment":100,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1504772,"name":"Shibla Abdulla","orcid":null,"position":1,"is_corresponding":false},{"id":1504773,"name":"Brian Aevermann","orcid":null,"position":2,"is_corresponding":false},{"id":1016446,"name":"Pedro Assis","orcid":"0000-0001-9289-0677","position":3,"is_corresponding":false},{"id":1504774,"name":"Seve Badajoz","orcid":null,"position":4,"is_corresponding":false},{"id":1504775,"name":"Sidney M Bell","orcid":null,"position":5,"is_corresponding":false},{"id":1504776,"name":"Emanuele Bezzi","orcid":null,"position":6,"is_corresponding":false},{"id":27445,"name":"Batuhan Cakir","orcid":null,"position":7,"is_corresponding":false},{"id":1504777,"name":"Jim Chaffer","orcid":null,"position":8,"is_corresponding":false},{"id":1504778,"name":"Signe Chambers","orcid":null,"position":9,"is_corresponding":false},{"id":11685,"name":"J. Michael Cherry","orcid":"0000-0001-9163-5180","position":10,"is_corresponding":false},{"id":1504779,"name":"Tiffany Chi","orcid":null,"position":11,"is_corresponding":false},{"id":515077,"name":"Jennifer Chien","orcid":"0000-0003-4389-9821","position":12,"is_corresponding":false},{"id":1504780,"name":"Leah Dorman","orcid":null,"position":13,"is_corresponding":false},{"id":1504781,"name":"Pablo Garcia-Nieto","orcid":null,"position":14,"is_corresponding":false},{"id":1504782,"name":"Nayib Gloria","orcid":null,"position":15,"is_corresponding":false},{"id":847281,"name":"Mim Hastie","orcid":null,"position":16,"is_corresponding":false},{"id":1504783,"name":"Daniel Hegeman","orcid":null,"position":17,"is_corresponding":false},{"id":1363723,"name":"Jason Hilton","orcid":"0000-0003-0286-8236","position":18,"is_corresponding":false},{"id":1504784,"name":"Timmy Huang","orcid":null,"position":19,"is_corresponding":false},{"id":1504785,"name":"Amanda Infeld","orcid":null,"position":20,"is_corresponding":false},{"id":1504786,"name":"Ana-Maria Istrate","orcid":null,"position":21,"is_corresponding":false},{"id":1504787,"name":"Ivana Jelic","orcid":null,"position":22,"is_corresponding":false},{"id":1504788,"name":"Kuni Katsuya","orcid":null,"position":23,"is_corresponding":false},{"id":277063,"name":"Yang Joon Kim","orcid":"0000-0003-1742-5657","position":24,"is_corresponding":false},{"id":1504789,"name":"Karen Liang","orcid":null,"position":25,"is_corresponding":false},{"id":16825,"name":"Mike Lin","orcid":null,"position":26,"is_corresponding":false},{"id":1326270,"name":"Maximilian Lombardo","orcid":null,"position":27,"is_corresponding":false},{"id":1254757,"name":"Bailey Marshall","orcid":"0009-0005-2344-2260","position":28,"is_corresponding":false},{"id":1161930,"name":"Bruce Martin","orcid":null,"position":29,"is_corresponding":false},{"id":1504790,"name":"Fran McDade","orcid":null,"position":30,"is_corresponding":false},{"id":1504791,"name":"Colin Megill","orcid":null,"position":31,"is_corresponding":false},{"id":390535,"name":"Nikhil Patel","orcid":"0000-0001-9335-0082","position":32,"is_corresponding":false},{"id":1504792,"name":"Alexander Predeus","orcid":null,"position":33,"is_corresponding":false},{"id":1504793,"name":"Brian Raymor","orcid":null,"position":34,"is_corresponding":false},{"id":1504794,"name":"Behnam Robatmili","orcid":null,"position":35,"is_corresponding":false},{"id":847279,"name":"Dave Rogers","orcid":null,"position":36,"is_corresponding":false},{"id":482351,"name":"Erica Rutherford","orcid":"0000-0001-8134-3037","position":37,"is_corresponding":false},{"id":1504795,"name":"Dana Sadgat","orcid":null,"position":38,"is_corresponding":false},{"id":815293,"name":"Andrew Shin","orcid":"0000-0002-3903-1432","position":39,"is_corresponding":false},{"id":1363722,"name":"Corinn Small","orcid":"0000-0002-2336-2552","position":40,"is_corresponding":false},{"id":1504796,"name":"Trent Smith","orcid":null,"position":41,"is_corresponding":false},{"id":1504797,"name":"Prathap Sridharan","orcid":null,"position":42,"is_corresponding":false},{"id":1504798,"name":"Alexander Tarashansky","orcid":null,"position":43,"is_corresponding":false},{"id":1504799,"name":"Norbert Tavares","orcid":null,"position":44,"is_corresponding":false},{"id":1504800,"name":"Harley Thomas","orcid":null,"position":45,"is_corresponding":false},{"id":1209359,"name":"Andrew Tolopko","orcid":null,"position":46,"is_corresponding":false},{"id":1504801,"name":"Meghan Urisko","orcid":null,"position":47,"is_corresponding":false},{"id":688848,"name":"Joyce Yan","orcid":null,"position":48,"is_corresponding":false},{"id":953196,"name":"Garabet Yeretssian","orcid":"0009-0006-2059-3810","position":49,"is_corresponding":false},{"id":1504802,"name":"Jennifer Zamanian","orcid":null,"position":50,"is_corresponding":false},{"id":1504803,"name":"Arathi Mani","orcid":null,"position":51,"is_corresponding":false},{"id":986668,"name":"Jonah Cool","orcid":"0000-0002-0287-011X","position":52,"is_corresponding":false},{"id":86683,"name":"Ambrose Carr","orcid":"0000-0002-8457-2836","position":53,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"CZ CELLxGENE Discover: a single-cell data platform for scalable exploration, analysis and modeling of aggregated data","abstract":"<jats:title>Abstract</jats:title>\n               <jats:p>Hundreds of millions of single cells have been analyzed using high-throughput transcriptomic methods. The cumulative knowledge within these datasets provides an exciting opportunity for unlocking insights into health and disease at the level of single cells. Meta-analyses that span diverse datasets building on recent advances in large language models and other machine-learning approaches pose exciting new directions to model and extract insight from single-cell data. Despite the promise of these and emerging analytical tools for analyzing large amounts of data, the sheer number of datasets, data models and accessibility remains a challenge. Here, we present CZ CELLxGENE Discover (cellxgene.cziscience.com), a data platform that provides curated and interoperable single-cell data. Available via a free-to-use online data portal, CZ CELLxGENE hosts a growing corpus of community-contributed data of over 93 million unique cells. Curated, standardized and associated with consistent cell-level metadata, this collection of single-cell transcriptomic data is the largest of its kind and growing rapidly via community contributions. A suite of tools and features enables accessibility and reusability of the data via both computational and visual interfaces to allow researchers to explore individual datasets, perform cross-corpus analysis, and run meta-analyses of tens of millions of cells across studies and tissues at the resolution of single cells.</jats:p>","is_dataset_classified":null,"base_score":5.820082930352362,"endowment":5.820082930352362,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"39607691","pmcid":"PMC11701654","openalex_id":"https://openalex.org/W4404793799","authors":[],"funders":[],"total_grants":0,"fwci":55.276,"citation_percentile":0.99964391,"influential_citations":0,"citation_trend":[{"year":2023,"count":2},{"year":2024,"count":20},{"year":2025,"count":179},{"year":2026,"count":135}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://academic.oup.com/nar/advance-article-pdf/doi/10.1093/nar/gkae1142/60882786/gkae1142.pdf","host_type":"journal"},{"url":"https://academic.oup.com/nar/advance-article-pdf/doi/10.1093/nar/gkae1142/60882786/gkae1142.pdf","host_type":"publisher"},{"url":"https://academic.oup.com/nar/article-pdf/53/D1/D886/60882786/gkae1142.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/nar/gkae1142","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/39607691","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/11701654","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC11701654/pdf/gkae1142.pdf","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC11701654","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC11701654?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Single-cell and spatial transcriptomics","Cell Image Analysis Techniques","Gene Regulatory Network Analysis","Single-Cell Analysis","Humans","Software","Transcriptome","Gene Expression Profiling","Machine Learning","Computational Biology"],"mesh_terms":["Machine Learning","Humans","Software","Computational Biology","Gene Expression Profiling","Single-Cell Analysis","Transcriptome"],"keywords":["Biology","Computational biology","Scalability","Data exploration","Visualization","Data science","Bioinformatics","Data mining","Computer science","Database"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[{"name":"efo"}],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-19T18:15:55.627690Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}