{"doi":"10.1101/2025.08.11.669102","title":"PanBGC: A Pangenome-inspired framework for comparative analysis of biosynthetic gene clusters","abstract":"<jats:title>\n                  A\n                  <jats:sc>bstract</jats:sc>\n                </jats:title>\n                <jats:p>\n                  Bacterial secondary metabolites are a major source of therapeutics and play key roles in microbial ecology. These compounds are encoded by biosynthetic gene clusters (BGCs), which show extensive genetic diversity across microbial genomes. While recent advances have enabled clustering of BGCs into gene cluster families (GCFs), there is still a lack of frameworks for systematically analysing their internal diversity at a population scale. Here, we introduce\n                  <jats:bold>PanBGC</jats:bold>\n                  , a pangenome-inspired framework that treats each GCF as a population of related BGCs. This enables classification of biosynthetic genes into core, accessory, and unique categories and provides openness metrics to quantify compositional diversity. Applied to over 250 000 BGCs from more than 35 000 genomes, PanBGC maps biosynthetic diversity of more than 80 000 GCFs. To facilitate exploration, we present\n                  <jats:bold>PanBGC-DB</jats:bold>\n                  (\n                  <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://panbgc-db.cs.uni-tuebingen.de\">https://panbgc-db.cs.uni-tuebingen.de</jats:ext-link>\n                  ), an interactive web platform for comparative BGC analysis. PanBGC-DB offers gene- and domain-level visualizations, phylogenetic tools, openness metrics, and custom query integration. Together, PanBGC and PanBGC-DB provide a scalable framework for exploring biosynthetic gene clusters at population resolution and for contextualizing newly discovered BGCs within the global landscape of secondary metabolism.\n                </jats:p>","journal":null,"year":null,"id":619517,"datarank":0.16479184330021646,"base_score":1.0986122886681096,"endowment":1.0986122886681096,"self_citation_contribution":0.16479184330021646,"citation_network_contribution":0.0,"self_endowment_contribution":0.16479184330021646,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":2,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1598783,"name":"Caner Bagci","orcid":null,"position":1,"is_corresponding":false},{"id":839432,"name":"Athina Gavriilidou","orcid":"0000-0002-6805-3868","position":2,"is_corresponding":false},{"id":305308,"name":"Nadine Ziemert","orcid":"0000-0002-7264-1857","position":3,"is_corresponding":false},{"id":1598781,"name":"Davide Paccagnella","orcid":"0000-0001-9746-0430","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"PanBGC: A Pangenome-inspired framework for comparative analysis of biosynthetic gene clusters","abstract":"<jats:title>\n                  A\n                  <jats:sc>bstract</jats:sc>\n                </jats:title>\n                <jats:p>\n                  Bacterial secondary metabolites are a major source of therapeutics and play key roles in microbial ecology. These compounds are encoded by biosynthetic gene clusters (BGCs), which show extensive genetic diversity across microbial genomes. While recent advances have enabled clustering of BGCs into gene cluster families (GCFs), there is still a lack of frameworks for systematically analysing their internal diversity at a population scale. Here, we introduce\n                  <jats:bold>PanBGC</jats:bold>\n                  , a pangenome-inspired framework that treats each GCF as a population of related BGCs. This enables classification of biosynthetic genes into core, accessory, and unique categories and provides openness metrics to quantify compositional diversity. Applied to over 250 000 BGCs from more than 35 000 genomes, PanBGC maps biosynthetic diversity of more than 80 000 GCFs. To facilitate exploration, we present\n                  <jats:bold>PanBGC-DB</jats:bold>\n                  (\n                  <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://panbgc-db.cs.uni-tuebingen.de\">https://panbgc-db.cs.uni-tuebingen.de</jats:ext-link>\n                  ), an interactive web platform for comparative BGC analysis. PanBGC-DB offers gene- and domain-level visualizations, phylogenetic tools, openness metrics, and custom query integration. Together, PanBGC and PanBGC-DB provide a scalable framework for exploring biosynthetic gene clusters at population resolution and for contextualizing newly discovered BGCs within the global landscape of secondary metabolism.\n                </jats:p>","is_dataset_classified":null,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19910364","pmcid":null,"openalex_id":"https://openalex.org/W4413128422","authors":[],"funders":[],"total_grants":0,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2026,"count":1}],"oa_status":"green","license":"cc-by","oa_locations":[{"url":"https://www.biorxiv.org/content/biorxiv/early/2025/08/11/2025.08.11.669102.full.pdf","host_type":"repository"},{"url":"https://www.biorxiv.org/content/biorxiv/early/2025/08/11/2025.08.11.669102.full.pdf","host_type":"repository"},{"url":"https://syndication.highwire.org/content/doi/10.1101/2025.08.11.669102","host_type":"publisher"},{"url":"https://doi.org/10.1101/2025.08.11.669102","host_type":"repository"}],"fields_of_study":["Microbial Metabolic Engineering and Bioproduction","Microbial Natural Products and Biosynthesis","Genomics and Phylogenetic Studies"],"mesh_terms":[],"keywords":["Gene","Computational biology","Biology","Evolutionary biology","Genetics"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-03T07:28:38.425544Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}