{"doi":"10.1101/2025.01.13.632775","title":"SwarmMAP: Swarm Learning for Decentralized Cell Type Annotation in Single Cell Sequencing Data","abstract":"Abstract Rapid technological advancements have made it possible to generate single-cell data at a large scale. Several laboratories around the world can now generate single-cell transcriptomic data from different tissues. Unsupervised clustering, followed by annotation of the cell type of the identified clusters, is a crucial step in single-cell analyses. However, there is no consensus on the marker genes to use for annotation, and celltype annotation is currently mostly done by manual inspection of marker genes, which is irreproducible, and poorly scalable. Additionally, patient-privacy is also a critical issue with human datasets. There is a critical need to standardize and automate celltype annotation across datasets in a privacy-preserving manner. Here, we developed SwarmMAP that uses Swarm Learning to train machine learning models for cell-type classification based on single-cell sequencing data in a decentralized way. SwarmMAP does not require any exchange of raw data between data centers. SwarmMAP has a F1-score of 0.93, 0.98, and 0.88 for cell type classification in human heart, lung, and breast datasets, respectively. Swarm Learning-based models yield an average performance of 0.907 which is on par with the performance achieved by models trained on centralized data ( p -val= 0.937 , Mann-Whitney U Test). We also find that increasing the number of datasets increases cell-type prediction accuracy and enables handling higher cell-type diversity. Together, these findings demonstrate that Swarm Learning is a viable approach to automate cell-type annotation. SwarmMAP is available at https://github.com/hayatlab/SwarmMAP .","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2025,"id":557081,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":1,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9516,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1457184,"name":"Vivien Goepp","orcid":null,"position":1,"is_corresponding":false},{"id":1376164,"name":"Kathy Pfeiffer","orcid":"0009-0000-7643-4284","position":2,"is_corresponding":false},{"id":858595,"name":"Hyojin Kim","orcid":"0000-0003-1553-4702","position":3,"is_corresponding":false},{"id":1456697,"name":"Jie Zhu","orcid":"0000-0003-0330-6382","position":4,"is_corresponding":false},{"id":350123,"name":"Rafael Kramann","orcid":"0000-0003-4048-6351","position":5,"is_corresponding":false},{"id":839687,"name":"Sikander Hayat","orcid":"0000-0001-5919-8371","position":6,"is_corresponding":false},{"id":110000,"name":"Jakob Nikolas Kather","orcid":"0000-0002-3730-5348","position":7,"is_corresponding":false},{"id":848638,"name":"Oliver Lester Saldanha","orcid":"0000-0002-3594-7590","position":0,"is_corresponding":true}],"reference_count":48,"raw_metadata":null,"created_at":"2026-07-19T02:55:13.130091Z","pmid":"39868099","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}