{"doi":"10.1101/2025.05.15.654203","title":"Fast, flexible gene cluster family delineation with IGUA","abstract":"<jats:title>ABSTRACT</jats:title>\n                <jats:p>\n                  Prokaryotic genomes harbor a variety of functional elements encoded as contiguous multi-gene clusters, with biosynthetic gene clusters (BGCs, genetic determinants of secondary metabolite biosynthesis) serving as a notable example. In a typical workflow, BGCs are clustered into Gene Cluster Families (GCFs), units that group BGCs encoding similar biosynthetic pathways together. However, existing methods cannot readily scale to massive datasets and cannot be used for GCF delineation tasks beyond BGC clustering. Here, we present IGUA (Iterative Gene clUster Analysis;\n                  <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://github.com/zellerlab/IGUA\">https://github.com/zellerlab/IGUA</jats:ext-link>\n                  ), a scalable, flexible GCF delineation method for genomic segments with multi-gene architectures. On a BGC clustering task, IGUA is ≥10x faster than the state-of-the-art (BiG-SCAPE/BiG-SLiCE), without sacrificing accuracy. To highlight its scalability, we use IGUA to cluster &gt;2.8 million BGCs from ≈1 million prokaryotic genomes in &lt;18 hours (\n                  <jats:italic>n</jats:italic>\n                  = 2,829,071 BGCs to 56,960 GCFs). To showcase its utility beyond BGC clustering, we use IGUA to cluster (i) secretion systems and (ii) prophages into GCFs (\n                  <jats:italic>n</jats:italic>\n                  = 10,576 and 356,776 gene clusters to 2,744 and 213,699 GCFs, respectively). Overall, IGUA represents a versatile GCF delineation tool with unmatched computational efficiency and flexibility, enabling (meta)genomic mining applications at unprecedented scales.\n                </jats:p>\n                <jats:sec>\n                  <jats:title>GRAPHICAL ABSTRACT</jats:title>\n                  <jats:fig id=\"ufig1\" position=\"float\" orientation=\"portrait\" fig-type=\"figure\">\n                    <jats:graphic xmlns:xlink=\"http://www.w3.org/1999/xlink\" xlink:href=\"654203v1_ufig1\" position=\"float\" orientation=\"portrait\"/>\n                  </jats:fig>\n                </jats:sec>","journal":null,"year":null,"id":655642,"datarank":0.26876392038420827,"base_score":1.791759469228055,"endowment":1.791759469228055,"self_citation_contribution":0.26876392038420827,"citation_network_contribution":0.0,"self_endowment_contribution":0.26876392038420827,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":5,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1481842,"name":"J. Blom","orcid":"0009-0000-9910-7374","position":1,"is_corresponding":false},{"id":1711533,"name":"Hadrien Gourlé","orcid":"0000-0001-9807-1082","position":2,"is_corresponding":false},{"id":829286,"name":"Laura M. Carroll","orcid":"0000-0002-3677-0192","position":3,"is_corresponding":false},{"id":36360,"name":"Georg Zeller","orcid":"0000-0003-1429-7485","position":4,"is_corresponding":false},{"id":7598,"name":"Martin Larralde","orcid":"0000-0002-3947-4444","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Fast, flexible gene cluster family delineation with IGUA","abstract":"<jats:title>ABSTRACT</jats:title>\n                <jats:p>\n                  Prokaryotic genomes harbor a variety of functional elements encoded as contiguous multi-gene clusters, with biosynthetic gene clusters (BGCs, genetic determinants of secondary metabolite biosynthesis) serving as a notable example. In a typical workflow, BGCs are clustered into Gene Cluster Families (GCFs), units that group BGCs encoding similar biosynthetic pathways together. However, existing methods cannot readily scale to massive datasets and cannot be used for GCF delineation tasks beyond BGC clustering. Here, we present IGUA (Iterative Gene clUster Analysis;\n                  <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://github.com/zellerlab/IGUA\">https://github.com/zellerlab/IGUA</jats:ext-link>\n                  ), a scalable, flexible GCF delineation method for genomic segments with multi-gene architectures. On a BGC clustering task, IGUA is ≥10x faster than the state-of-the-art (BiG-SCAPE/BiG-SLiCE), without sacrificing accuracy. To highlight its scalability, we use IGUA to cluster &gt;2.8 million BGCs from ≈1 million prokaryotic genomes in &lt;18 hours (\n                  <jats:italic>n</jats:italic>\n                  = 2,829,071 BGCs to 56,960 GCFs). To showcase its utility beyond BGC clustering, we use IGUA to cluster (i) secretion systems and (ii) prophages into GCFs (\n                  <jats:italic>n</jats:italic>\n                  = 10,576 and 356,776 gene clusters to 2,744 and 213,699 GCFs, respectively). Overall, IGUA represents a versatile GCF delineation tool with unmatched computational efficiency and flexibility, enabling (meta)genomic mining applications at unprecedented scales.\n                </jats:p>\n                <jats:sec>\n                  <jats:title>GRAPHICAL ABSTRACT</jats:title>\n                  <jats:fig id=\"ufig1\" position=\"float\" orientation=\"portrait\" fig-type=\"figure\">\n                    <jats:graphic xmlns:xlink=\"http://www.w3.org/1999/xlink\" xlink:href=\"654203v1_ufig1\" position=\"float\" orientation=\"portrait\"/>\n                  </jats:fig>\n                </jats:sec>","is_dataset_classified":null,"base_score":1.791759469228055,"endowment":1.791759469228055,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19162232","pmcid":null,"openalex_id":"https://openalex.org/W4410524657","authors":[],"funders":[{"funder_name":"Knut and Alice Wallenberg Foundation","grant_id":"KAW 2020.0239","title":null},{"funder_name":"Swedish Research Council","grant_id":"2023-05212","title":null},{"funder_name":"Deutsche Forschungsgemeinschaft","grant_id":"395357507","title":null},{"funder_name":"European Research Council","grant_id":"101118531","title":"Functional cartography of intestinal host-microbiome interactions"},{"funder_name":"Deutsche Forschungsgemeinschaft","grant_id":"395357507/SFB 1371","title":"Microbiome Signatures -- Functional Relevance in the Digestive Tract"}],"total_grants":5,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2025,"count":4},{"year":2026,"count":1}],"oa_status":"green","license":"cc-by","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2025/05/19/2025.05.15.654203.full.pdf","host_type":"repository"},{"url":"https://www.biorxiv.org/content/biorxiv/early/2025/05/19/2025.05.15.654203.full.pdf","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.1101/2025.05.15.654203","host_type":"publisher"},{"url":"https://doi.org/10.1101/2025.05.15.654203","host_type":"repository"},{"url":"https://github.com/zellerlab/IGUA/tree/v0.2.1","host_type":"repository"},{"url":"https://github.com/zellerlab/IGUA/tree/v0.2.0","host_type":"repository"},{"url":"https://doi.org/10.5281/zenodo.20540570","host_type":"repository"},{"url":"https://europepmc.org/article/PPR/PPR1023339","host_type":"Europe_PMC"},{"url":"https://europepmc.org/api/fulltextRepo?pprId=PPR1023339&type=FILE&fileName=EMS205739-pdf.pdf&mimeType=application/pdf","host_type":"Europe_PMC"}],"fields_of_study":["Gene expression and cancer classification","Machine Learning in Bioinformatics","0206 medical engineering","02 engineering and technology"],"mesh_terms":[],"keywords":["Cluster (spacecraft)","Gene cluster","Gene","Computational biology","Computer science","Genetics","Biology","Programming language"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[{"name":"doi"}],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-11T18:24:26.603346Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}