{"doi":"10.3389/frsip.2022.842513","title":"How Scalable Are Clade-Specific Marker K-Mer Based Hash Methods for Metagenomic Taxonomic Classification?","abstract":"<jats:p>Efficiently and accurately identifying which microbes are present in a biological sample is important to medicine and biology. For example, in medicine, microbe identification allows doctors to better diagnose diseases. Two questions are essential to metagenomic analysis (the analysis of a random sampling of DNA in a patient/environment sample): How to accurately identify the microbes in samples and how to efficiently update the taxonomic classifier as new microbe genomes are sequenced and added to the reference database. To investigate how classifiers change as they train on more knowledge, we made sub-databases composed of genomes that existed in past years that served as “snapshots in time” (1999–2020) of the NCBI reference genome database. We evaluated two classification methods, Kraken 2 and CLARK with these snapshots using a real, experimental metagenomic sample from a human gut. This allowed us to measure how much of a real sample could confidently classify using these methods and as the database grows. Despite not knowing the ground truth, we could measure the concordance between methods and between years of the database within each method using a Bray-Curtis distance. In addition, we also recorded the training times of the classifiers for each snapshot. For all data for Kraken 2, we observed that as more genomes were added, more microbes from the sample were classified. CLARK had a similar trend, but in the final year, this trend reversed with the microbial variation and less unique k-mers. Also, both classifiers, while having different ways of training, generally are linear in time - but Kraken 2 has a significantly lower slope in scaling to more data.</jats:p>","journal":"Frontiers in Signal Processing","year":2022,"id":650782,"datarank":0.24141568686511508,"base_score":1.6094379124341003,"endowment":1.6094379124341003,"self_citation_contribution":0.24141568686511508,"citation_network_contribution":0.0,"self_endowment_contribution":0.24141568686511508,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":4,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":247399,"name":"Zhengqiao Zhao","orcid":"0000-0001-6873-6098","position":1,"is_corresponding":false},{"id":85496,"name":"Gail L. Rosen","orcid":"0000-0003-1763-5750","position":2,"is_corresponding":false},{"id":962836,"name":"Melissa Gray","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"How Scalable Are Clade-Specific Marker K-Mer Based Hash Methods for Metagenomic Taxonomic Classification?","abstract":"<jats:p>Efficiently and accurately identifying which microbes are present in a biological sample is important to medicine and biology. For example, in medicine, microbe identification allows doctors to better diagnose diseases. Two questions are essential to metagenomic analysis (the analysis of a random sampling of DNA in a patient/environment sample): How to accurately identify the microbes in samples and how to efficiently update the taxonomic classifier as new microbe genomes are sequenced and added to the reference database. To investigate how classifiers change as they train on more knowledge, we made sub-databases composed of genomes that existed in past years that served as “snapshots in time” (1999–2020) of the NCBI reference genome database. We evaluated two classification methods, Kraken 2 and CLARK with these snapshots using a real, experimental metagenomic sample from a human gut. This allowed us to measure how much of a real sample could confidently classify using these methods and as the database grows. Despite not knowing the ground truth, we could measure the concordance between methods and between years of the database within each method using a Bray-Curtis distance. In addition, we also recorded the training times of the classifiers for each snapshot. For all data for Kraken 2, we observed that as more genomes were added, more microbes from the sample were classified. CLARK had a similar trend, but in the final year, this trend reversed with the microbial variation and less unique k-mers. Also, both classifiers, while having different ways of training, generally are linear in time - but Kraken 2 has a significantly lower slope in scaling to more data.</jats:p>","is_dataset_classified":null,"base_score":1.6094379124341003,"endowment":1.6094379124341003,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19965766","pmcid":null,"openalex_id":"https://openalex.org/W4283818660","authors":[],"funders":[{"funder_name":"National Science Foundation","grant_id":"2107108","title":"III: Small: Learning Multi-scale Sequence Features for Predicting Gene to Microbiome Function"},{"funder_name":"National Science Foundation","grant_id":"1919691","title":"MRI: Proteus++: Enabling Data-Intensive Computing at Drexel University"},{"funder_name":"National Science Foundation","grant_id":"1936791","title":"Collaborative Research: IIBR Informatics: Keeping up with the genomes - Continual Learning of Metagenomic Data"}],"total_grants":3,"fwci":0.1571,"citation_percentile":0.42556909,"influential_citations":0,"citation_trend":[{"year":2024,"count":2},{"year":2026,"count":2}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://doi.org/10.3389/frsip.2022.842513","host_type":"journal"},{"url":"https://doi.org/10.3389/frsip.2022.842513","host_type":"publisher"},{"url":"https://www.frontiersin.org/articles/10.3389/frsip.2022.842513/full","host_type":"publisher"},{"url":"https://doaj.org/article/39777fcd8e354b24b260be7cbe3960ed","host_type":"repository"}],"fields_of_study":["Gut microbiota and health","Genomics and Phylogenetic Studies","Metabolomics and Mass Spectrometry Studies","0301 basic medicine","03 medical and health sciences","0206 medical engineering","02 engineering and technology"],"mesh_terms":[],"keywords":["Metagenomics","Biological classification","Genome","Concordance","Classifier (UML)","1000 Genomes Project","Computer science","Sample (material)","Computational biology","Snapshot (computer storage)","Artificial intelligence","Biology","Data mining","Bioinformatics","Genetics","Evolutionary biology","Database","Gene","Single-nucleotide polymorphism","incremental learning","supervised classification","taxonomic classification","Electrical engineering. Electronics. Nuclear engineering","hash-based indexing","algorithm scalability","TK1-9971"],"sdg_mappings":[{"sdg_number":3,"sdg_label":"3. Good health"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-10T06:06:25.048748Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}