{"doi":"10.1093/bib/bbag119","title":"Kun-peng enables scalable and accurate pan-domain metagenomic classification","abstract":"<jats:title>Abstract</jats:title>\n                  <jats:p>Comprehensive pan-domain metagenomic classification is increasingly constrained by the memory and runtime costs of building and querying the rapidly expanding reference genome space. We introduce Kun-peng, a taxonomic classifier powered by an intelligent block-partitioned database structure and optimized search strategies, enabling ultra-scalable, memory-efficient pan-domain profiling. Using the Critical Assessment of Metagenome Interpretation II benchmark, Kun-peng substantially reduces the memory usage of database-building and querying by up to 24-fold, and accelerates sample classification by up to 4.73-fold compared with Kraken2. Kun-peng achieves competitive accuracy with fewer false positives than Kraken2, Centrifuger, and even KrakenUniq, while maintaining consistently high sensitivity across diverse datasets. In a real-world evaluation of 586 metagenomic samples spanning air, water, soil, and human-associated environments, we performed classification using a 4.3 TB pan-domain database comprising 204,477 genomes, which was built by Kun-peng with only 4.1 GB peak memory. Kun-peng processed each sample in 0.2–11.2 min with 4.0–35.4 GB peak memory, corresponding to a 54–473-fold reduction in memory usage relative to Kraken2. Compared with Sylph, Kun-peng achieved up to a 46-fold speedup while requiring 21-fold less memory. Kun-peng classified 69.8%–94.3% of reads, improving coverage by 20%–60% over the standard Kraken2 database with 62,026 genomes. This improvement reflects expanded reference coverage, although a small fraction of false positives is inherent to k-mer-based methods. Overall, Kun-peng effectively eliminates the long-standing memory bottleneck in pan-domain database building and classification, enabling rapid and scalable pan-domain taxonomic analysis of complex environmental, ecological, and exposomic sequencing datasets.</jats:p>","journal":"Briefings in Bioinformatics","year":2026,"id":650676,"datarank":0.20794415416798362,"base_score":1.3862943611198906,"endowment":1.3862943611198906,"self_citation_contribution":0.20794415416798362,"citation_network_contribution":0.0,"self_endowment_contribution":0.20794415416798362,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1696647,"name":"Boliang Zhang","orcid":null,"position":1,"is_corresponding":false},{"id":331278,"name":"Chen Peng","orcid":"0000-0001-6141-8775","position":2,"is_corresponding":false},{"id":1661200,"name":"Jiajun Huang","orcid":null,"position":3,"is_corresponding":false},{"id":275170,"name":"Zhen Liu","orcid":"0000-0002-1119-0693","position":4,"is_corresponding":false},{"id":808048,"name":"Xiaotao Shen","orcid":"0000-0002-9608-9964","position":5,"is_corresponding":false},{"id":374792,"name":"Chao Jiang","orcid":"0000-0003-0260-7271","position":6,"is_corresponding":false},{"id":495376,"name":"Qiong Chen","orcid":"0000-0003-2401-0046","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Kun-peng enables scalable and accurate pan-domain metagenomic classification","abstract":"<jats:title>Abstract</jats:title>\n                  <jats:p>Comprehensive pan-domain metagenomic classification is increasingly constrained by the memory and runtime costs of building and querying the rapidly expanding reference genome space. We introduce Kun-peng, a taxonomic classifier powered by an intelligent block-partitioned database structure and optimized search strategies, enabling ultra-scalable, memory-efficient pan-domain profiling. Using the Critical Assessment of Metagenome Interpretation II benchmark, Kun-peng substantially reduces the memory usage of database-building and querying by up to 24-fold, and accelerates sample classification by up to 4.73-fold compared with Kraken2. Kun-peng achieves competitive accuracy with fewer false positives than Kraken2, Centrifuger, and even KrakenUniq, while maintaining consistently high sensitivity across diverse datasets. In a real-world evaluation of 586 metagenomic samples spanning air, water, soil, and human-associated environments, we performed classification using a 4.3 TB pan-domain database comprising 204,477 genomes, which was built by Kun-peng with only 4.1 GB peak memory. Kun-peng processed each sample in 0.2–11.2 min with 4.0–35.4 GB peak memory, corresponding to a 54–473-fold reduction in memory usage relative to Kraken2. Compared with Sylph, Kun-peng achieved up to a 46-fold speedup while requiring 21-fold less memory. Kun-peng classified 69.8%–94.3% of reads, improving coverage by 20%–60% over the standard Kraken2 database with 62,026 genomes. This improvement reflects expanded reference coverage, although a small fraction of false positives is inherent to k-mer-based methods. Overall, Kun-peng effectively eliminates the long-standing memory bottleneck in pan-domain database building and classification, enabling rapid and scalable pan-domain taxonomic analysis of complex environmental, ecological, and exposomic sequencing datasets.</jats:p>","is_dataset_classified":null,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"41838875","pmcid":"PMC12991049","openalex_id":"https://openalex.org/W7137288452","authors":[],"funders":[{"funder_name":"National Natural Science Foundation of China","grant_id":"82341109","title":null},{"funder_name":"National Natural Science Foundation of China","grant_id":"82173645","title":null}],"total_grants":2,"fwci":7.1874,"citation_percentile":0.96680906,"influential_citations":0,"citation_trend":[{"year":2026,"count":2}],"oa_status":"gold","license":"cc-by-nc","oa_locations":[{"url":"https://doi.org/10.1093/bib/bbag119","host_type":"journal"},{"url":"https://doi.org/10.1093/bib/bbag119","host_type":"publisher"},{"url":"https://academic.oup.com/bib/article-pdf/27/2/bbag119/67371377/bbag119.pdf","host_type":"publisher"},{"url":"https://pubmed.ncbi.nlm.nih.gov/41838875","host_type":"repository"},{"url":"https://hdl.handle.net/10356/211181","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12991049/","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC12991049","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC12991049?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Genomics and Phylogenetic Studies","Machine Learning in Bioinformatics","Epigenetics and DNA Methylation","Metagenomics","Metagenome","Algorithms","Humans","Software","Databases, Genetic","Computational Biology"],"mesh_terms":["Algorithms","Humans","Software","Computational Biology","Databases, Genetic","Metagenome","Metagenomics"],"keywords":["Metagenomics","False positive paradox","Bottleneck","Scalability","Classifier (UML)","Reference database","Speedup","Sample (material)","Metagenomic Classification","Block-Partitioned Database","Memory-Efficient Algorithms","Minimizer-Based Indexing","Pan-Domain Microbiome Analysis"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-10T05:55:13.322379Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}