{"doi":"10.1093/bioinformatics/btaf655","title":"CeLLTra: aligning cell names with gene expression via a pathway-informed transformer","abstract":"MOTIVATION: Single-cell RNA sequencing (scRNA-Seq) technology enables detailed exploration of gene expression at the individual cell level, crucial for annotating cell types and understanding cellular diversity. Traditional methods for cell type annotation often rely on marker genes and manual labeling, posing challenges due to low data quality and incomplete reference datasets. RESULTS: We developed CeLLTra, a novel contrastive learning framework that leverages a Transformer-based model integrating biological pathway information to group genes into super tokens, effectively capturing comprehensive gene expression from scRNA-Seq data. By combining this pathway-informed Transformer with a pretrained domain-specific language model, CeLLTra accurately aligns cell-type annotations with gene expression profiles. Evaluations on a large-scale human scRNA-Seq dataset showed that CeLLTra significantly outperformed state-of-the-art methods in supervised and zero-shot cell-type prediction. Additionally, CeLLTra generalized well to external datasets, improving clustering performance and enabling better characterization of cancerous cell states in tumor-infiltrating myeloid cells from non-small cell lung cancer patients. AVAILABILITY AND IMPLEMENTATION: CeLLTra is freely available on GitHub (https://github.com/WJZheng-group/CeLLTra) and Zenodo (https://doi.org/10.5281/zenodo.17666735). The datasets underlying this article are the following: GSE201333 and GSE127465. All these datasets are publicly available and can be freely accessed on the Gene Expression Omnibus repository.","journal":"Bioinformatics","year":2025,"id":587830,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9411,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1503359,"name":"Zaiyi Zheng","orcid":"0009-0003-0685-0057","position":1,"is_corresponding":false},{"id":829115,"name":"Rongbin Li","orcid":"0000-0002-3605-6570","position":2,"is_corresponding":false},{"id":514608,"name":"Wenbo Chen","orcid":"0000-0002-1414-6540","position":3,"is_corresponding":false},{"id":441641,"name":"Yuntao Yang","orcid":"0009-0009-2703-5872","position":4,"is_corresponding":false},{"id":1504038,"name":"Meer Asif Ali","orcid":null,"position":5,"is_corresponding":false},{"id":1166060,"name":"Jundong Li","orcid":"0000-0002-1878-817X","position":6,"is_corresponding":false},{"id":1504039,"name":"W Jim Zheng","orcid":null,"position":7,"is_corresponding":false},{"id":927017,"name":"Li Zhao","orcid":"0000-0002-2544-0876","position":0,"is_corresponding":true}],"reference_count":32,"raw_metadata":null,"created_at":"2026-07-19T02:59:39.958043Z","pmid":"41652996","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}