{"doi":"10.1093/bioinformatics/btab516","title":"Founder reconstruction enables scalable and seamless pangenomic analysis","abstract":"<jats:title>Abstract</jats:title>\n               <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>Variant calling workflows that utilize a single reference sequence are the de facto standard elementary genomic analysis routine for resequencing projects. Various ways to enhance the reference with pangenomic information have been proposed, but scalability combined with seamless integration to existing workflows remains a challenge.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>We present PanVC with founder sequences, a scalable and accurate variant calling workflow based on a multiple alignment of reference sequences. Scalability is achieved by removing duplicate parts up to a limit into a founder multiple alignment, that is then indexed using a hybrid scheme that exploits general purpose read aligners. Our implemented workflow uses GATK or BCFtools for variant calling, but the various steps of our workflow (e.g. vcf2multialign tool, founder reconstruction) can be of independent interest as a basis for creating novel pangenome analysis workflows beyond variant calling.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation</jats:title>\n                  <jats:p>Our open access tools and instructions how to reproduce our experiments are available at the following address: https://github.com/algbio/panvc-founders.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information</jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.</jats:p>\n               </jats:sec>","journal":"Bioinformatics","year":2021,"id":620510,"datarank":0.42498200160843247,"base_score":2.833213344056216,"endowment":2.833213344056216,"self_citation_contribution":0.42498200160843247,"citation_network_contribution":0.0,"self_endowment_contribution":0.42498200160843247,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":16,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":15,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1601893,"name":"Bastien Cazaux","orcid":null,"position":1,"is_corresponding":false},{"id":1601894,"name":"Saska Dönges","orcid":null,"position":2,"is_corresponding":false},{"id":1601895,"name":"Daniel Valenzuela","orcid":null,"position":3,"is_corresponding":false},{"id":1601896,"name":"Veli Mäkinen","orcid":"0000-0003-4454-1493","position":4,"is_corresponding":false},{"id":1601892,"name":"Tuukka Norri","orcid":"0000-0002-8276-0585","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Founder reconstruction enables scalable and seamless pangenomic analysis","abstract":"<jats:title>Abstract</jats:title>\n               <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>Variant calling workflows that utilize a single reference sequence are the de facto standard elementary genomic analysis routine for resequencing projects. Various ways to enhance the reference with pangenomic information have been proposed, but scalability combined with seamless integration to existing workflows remains a challenge.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>We present PanVC with founder sequences, a scalable and accurate variant calling workflow based on a multiple alignment of reference sequences. Scalability is achieved by removing duplicate parts up to a limit into a founder multiple alignment, that is then indexed using a hybrid scheme that exploits general purpose read aligners. Our implemented workflow uses GATK or BCFtools for variant calling, but the various steps of our workflow (e.g. vcf2multialign tool, founder reconstruction) can be of independent interest as a basis for creating novel pangenome analysis workflows beyond variant calling.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation</jats:title>\n                  <jats:p>Our open access tools and instructions how to reproduce our experiments are available at the following address: https://github.com/algbio/panvc-founders.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information</jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.</jats:p>\n               </jats:sec>","is_dataset_classified":null,"base_score":2.833213344056216,"endowment":2.833213344056216,"datacite_reuse_total":15,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"34260702","pmcid":"PMC8665761","openalex_id":"https://openalex.org/W3181407718","authors":[],"funders":[{"funder_name":"Academy of Finland","grant_id":"309048","title":"Sequence analysis revisited and extended"},{"funder_name":"Helsinki Institute for Information Technology","grant_id":"","title":null}],"total_grants":2,"fwci":1.7446,"citation_percentile":0.85162366,"influential_citations":0,"citation_trend":[{"year":2021,"count":1},{"year":2022,"count":2},{"year":2023,"count":3},{"year":2024,"count":8},{"year":2025,"count":2}],"oa_status":"hybrid","license":"cc-by","oa_locations":[{"url":"https://academic.oup.com/bioinformatics/article-pdf/37/24/4611/41726965/btab516.pdf","host_type":"journal"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/37/24/4611/41726965/btab516.pdf","host_type":"publisher"},{"url":"http://academic.oup.com/bioinformatics/advance-article-pdf/doi/10.1093/bioinformatics/btab516/40391795/btab516.pdf","host_type":"publisher"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/37/24/4611/50334896/btab516.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/bioinformatics/btab516","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/34260702","host_type":"repository"},{"url":"http://hdl.handle.net/10138/338304","host_type":"journal"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/8665761","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC8665761","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC8665761?pdf=render","host_type":"Europe_PMC"},{"url":"http://dx.doi.org/10.1093/bioinformatics/btab516","host_type":""},{"url":"https://dx.doi.org/10.1093/bioinformatics/btab516","host_type":""}],"fields_of_study":["Genomics and Phylogenetic Studies","Scientific Computing and Data Management","Biomedical Text Mining and Ontologies","0301 basic medicine","0206 medical engineering","02 engineering and technology","03 medical and health sciences","Software","Genomics","Sequence Analysis, DNA","Genome","Workflow"],"mesh_terms":["Software","Genome","Sequence Analysis, DNA","Genomics","Workflow"],"keywords":["Computer science","Scalability","Software","Programming language","Operating system","Computer and information sciences","Genome","GENOMES","Genomics","Sequence Analysis, DNA","Original Papers","READ ALIGNMENT","GRAPHS","Workflow","SET"],"sdg_mappings":[],"linked_datasets":[{"doi":"10.6084/m9.figshare.26596085.v1","title":"Additional file 1 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.26596085","title":"Additional file 1 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.26596100.v1","title":"Additional file 6 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.26596100","title":"Additional file 6 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.26621083.v1","title":"Additional file 1 of Constructing founder sets under allelic and non-allelic homologous recombination","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.6084/m9.figshare.26621083","title":"Additional file 1 of Constructing founder sets under allelic and non-allelic homologous recombination","publisher":"figshare","resource_type":"JournalArticle"},{"doi":"10.4230/lipics.cpm.2022.19","title":"Indexable Elastic Founder Graphs of Minimum Height","publisher":"Schloss Dagstuhl – Leibniz-Zentrum für Informatik","resource_type":"ConferencePaper"},{"doi":"10.6084/m9.figshare.26596097.v1","title":"Additional file 5 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"Dataset"},{"doi":"10.6084/m9.figshare.26596088","title":"Additional file 2 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"Dataset"},{"doi":"10.6084/m9.figshare.26596097","title":"Additional file 5 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"Dataset"},{"doi":"10.6084/m9.figshare.26596094.v1","title":"Additional file 4 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"Dataset"},{"doi":"10.6084/m9.figshare.26596088.v1","title":"Additional file 2 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"Dataset"},{"doi":"10.6084/m9.figshare.26596094","title":"Additional file 4 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"Dataset"},{"doi":"10.6084/m9.figshare.26596091.v1","title":"Additional file 3 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"Dataset"},{"doi":"10.6084/m9.figshare.26596091","title":"Additional file 3 of Matchtigs: minimum plain text representation of k-mer sets","publisher":"figshare","resource_type":"Dataset"}],"clinical_trials":[],"software_tools":[],"database_accessions":[{"name":"doi"}],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-03T11:18:28.880993Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}