{"doi":"10.3389/fbinf.2022.861505","title":"Metascan: METabolic Analysis, SCreening and ANnotation of Metagenomes","abstract":"<jats:p>Large scale next generation metagenomic sequencing of complex environmental samples paves the way for detailed analysis of nutrient cycles in ecosystems. For such an analysis, large scale unequivocal annotation is a prerequisite, which however is increasingly hampered by growing databases and analysis time. Hereto, we created a hidden Markov model (HMM) database by clustering proteins according to their KEGG indexing. HMM profiles for key genes of specific metabolic pathways and nutrient cycles were organized in subsets to be able to analyze each important elemental cycle separately. An important motivation behind the clustered database was to enable a high degree of resolution for annotation, while decreasing database size and analysis time. Here, we present Metascan, a new tool that can fully annotate and analyze deeply sequenced samples with an average analysis time of 11 min per genome for a publicly available dataset containing 2,537 genomes, and 1.1 min per genome for nutrient cycle analysis of the same sample. Metascan easily detected general proteins like cytochromes and ferredoxins, and additional <jats:italic>pmoCAB</jats:italic> operons were identified that were overlooked in previous analyses. For a mock community, the BEACON (F1) score was 0.72–0.93 compared to the information in NCBI GenBank. In combination with the accompanying database, Metascan provides a fast and useful annotation and analysis tool, as demonstrated by our proof-of-principle analysis of a complex mock community metagenome.</jats:p>","journal":"Frontiers in Bioinformatics","year":2022,"id":607690,"datarank":0.4493598410330987,"base_score":2.995732273553991,"endowment":2.995732273553991,"self_citation_contribution":0.4493598410330987,"citation_network_contribution":0.0,"self_endowment_contribution":0.4493598410330987,"citer_contribution":0.0,"corpus_percentile":58.0,"corpus_rank":5527,"citation_count":19,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1560476,"name":"Mike S. M. Jetten","orcid":null,"position":1,"is_corresponding":false},{"id":1560477,"name":"Huub J. M. Op den Camp","orcid":null,"position":2,"is_corresponding":false},{"id":273601,"name":"Sebastian Lücker","orcid":"0000-0003-2935-4454","position":3,"is_corresponding":false},{"id":1560474,"name":"Geert Cremers","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Metascan: METabolic Analysis, SCreening and ANnotation of Metagenomes","abstract":"<jats:p>Large scale next generation metagenomic sequencing of complex environmental samples paves the way for detailed analysis of nutrient cycles in ecosystems. For such an analysis, large scale unequivocal annotation is a prerequisite, which however is increasingly hampered by growing databases and analysis time. Hereto, we created a hidden Markov model (HMM) database by clustering proteins according to their KEGG indexing. HMM profiles for key genes of specific metabolic pathways and nutrient cycles were organized in subsets to be able to analyze each important elemental cycle separately. An important motivation behind the clustered database was to enable a high degree of resolution for annotation, while decreasing database size and analysis time. Here, we present Metascan, a new tool that can fully annotate and analyze deeply sequenced samples with an average analysis time of 11 min per genome for a publicly available dataset containing 2,537 genomes, and 1.1 min per genome for nutrient cycle analysis of the same sample. Metascan easily detected general proteins like cytochromes and ferredoxins, and additional <jats:italic>pmoCAB</jats:italic> operons were identified that were overlooked in previous analyses. For a mock community, the BEACON (F1) score was 0.72–0.93 compared to the information in NCBI GenBank. In combination with the accompanying database, Metascan provides a fast and useful annotation and analysis tool, as demonstrated by our proof-of-principle analysis of a complex mock community metagenome.</jats:p>","is_dataset_classified":null,"base_score":2.995732273553991,"endowment":2.995732273553991,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"36304333","pmcid":"PMC9580885","openalex_id":"https://openalex.org/W4283371129","authors":[],"funders":[{"funder_name":"Nederlandse Organisatie voor Wetenschappelijk Onderzoek","grant_id":"024.002.002 016.Vidi.189.050","title":null},{"funder_name":"H2020 European Research Council","grant_id":"669371 854088","title":null},{"funder_name":"Dutch Research Council (NWO)","grant_id":"024.002.002","title":"Microbes for health and environment"},{"funder_name":"European Research Council","grant_id":"854088","title":"Methane and Ammonium Removal In redoX transition zones"},{"funder_name":"European Research Council","grant_id":"669371","title":"Microbiology of extremely acidic terrestrial volcanic ecosystems"}],"total_grants":5,"fwci":1.0969,"citation_percentile":0.75460489,"influential_citations":0,"citation_trend":[{"year":2022,"count":2},{"year":2023,"count":1},{"year":2024,"count":3},{"year":2025,"count":8},{"year":2026,"count":5}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://doi.org/10.3389/fbinf.2022.861505","host_type":"journal"},{"url":"https://doi.org/10.3389/fbinf.2022.861505","host_type":"publisher"},{"url":"https://www.frontiersin.org/articles/10.3389/fbinf.2022.861505/full","host_type":"publisher"},{"url":"https://pubmed.ncbi.nlm.nih.gov/36304333","host_type":"repository"},{"url":"https://doaj.org/article/d18c29c7fc3043b9a873696a55bc53dc","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/9580885","host_type":"repository"},{"url":"https://repository.ubn.ru.nl/handle/2066/252558","host_type":"repository"},{"url":"http://hdl.handle.net/2066/252558","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC9580885","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC9580885?pdf=render","host_type":"Europe_PMC"},{"url":"http://dx.doi.org/10.3389/fbinf.2022.861505","host_type":""},{"url":"https://hdl.handle.net/https://repository.ubn.ru.nl/handle/2066/252558","host_type":""},{"url":"https://repository.ubn.ru.nl//bitstream/handle/2066/252558/252558.pdf","host_type":""},{"url":"https://hdl.handle.net/2066/252558","host_type":""}],"fields_of_study":["Genomics and Phylogenetic Studies","Advanced Proteomics Techniques and Applications","Microbial Community Ecology and Physiology","0301 basic medicine","03 medical and health sciences","0206 medical engineering","02 engineering and technology"],"mesh_terms":[],"keywords":["Annotation","GenBank","Metagenomics","KEGG","Genome","Computer science","Genome project","Computational biology","Biology","Gene","Genetics","Artificial intelligence","Gene ontology","Ecology","Metabolism","Microbiology","Bioinformatics","Ecological Microbiology","Computer applications to medicine. Medical informatics","R858-859.7"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Life in Land"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-30T06:50:20.797205Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}