{"doi":"10.1093/bioinformatics/btx261","title":"deBGR: an efficient and near-exact representation of the weighted de Bruijn graph","abstract":"<jats:title>Abstract</jats:title>\n               <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>Almost all de novo short-read genome and transcriptome assemblers start by building a representation of the de Bruijn Graph of the reads they are given as input. Even when other approaches are used for subsequent assembly (e.g. when one is using ‘long read’ technologies like those offered by PacBio or Oxford Nanopore), efficient k-mer processing is still crucial for accurate assembly, and state-of-the-art long-read error-correction methods use de Bruijn Graphs. Because of the centrality of de Bruijn Graphs, researchers have proposed numerous methods for representing de Bruijn Graphs compactly. Some of these proposals sacrifice accuracy to save space. Further, none of these methods store abundance information, i.e. the number of times that each k-mer occurs, which is key in transcriptome assemblers.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>We present a method for compactly representing the weighted de Bruijn Graph (i.e. with abundance information) with essentially no errors. Our representation yields zero errors while increasing the space requirements by less than 18–28% compared to the approximate de Bruijn graph representation in Squeakr. Our technique is based on a simple invariant that all weighted de Bruijn Graphs must satisfy, and hence is likely to be of general interest and applicable in most weighted de Bruijn Graph-based systems.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation</jats:title>\n                  <jats:p>https://github.com/splatlab/debgr.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information</jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.</jats:p>\n               </jats:sec>","journal":"Bioinformatics","year":2017,"id":596730,"datarank":0.5641800173540344,"base_score":3.7612001156935624,"endowment":3.7612001156935624,"self_citation_contribution":0.5641800173540344,"citation_network_contribution":0.0,"self_endowment_contribution":0.5641800173540344,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":42,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1528389,"name":"Michael A Bender","orcid":null,"position":1,"is_corresponding":false},{"id":1304436,"name":"Rob Johnson","orcid":"0000-0002-7365-0042","position":2,"is_corresponding":false},{"id":87821,"name":"Rob Patro","orcid":"0000-0001-8463-1675","position":3,"is_corresponding":false},{"id":556058,"name":"Prashant Pandey","orcid":"0000-0001-5576-0320","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"deBGR: an efficient and near-exact representation of the weighted de Bruijn graph","abstract":"<jats:title>Abstract</jats:title>\n               <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>Almost all de novo short-read genome and transcriptome assemblers start by building a representation of the de Bruijn Graph of the reads they are given as input. Even when other approaches are used for subsequent assembly (e.g. when one is using ‘long read’ technologies like those offered by PacBio or Oxford Nanopore), efficient k-mer processing is still crucial for accurate assembly, and state-of-the-art long-read error-correction methods use de Bruijn Graphs. Because of the centrality of de Bruijn Graphs, researchers have proposed numerous methods for representing de Bruijn Graphs compactly. Some of these proposals sacrifice accuracy to save space. Further, none of these methods store abundance information, i.e. the number of times that each k-mer occurs, which is key in transcriptome assemblers.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>We present a method for compactly representing the weighted de Bruijn Graph (i.e. with abundance information) with essentially no errors. Our representation yields zero errors while increasing the space requirements by less than 18–28% compared to the approximate de Bruijn graph representation in Squeakr. Our technique is based on a simple invariant that all weighted de Bruijn Graphs must satisfy, and hence is likely to be of general interest and applicable in most weighted de Bruijn Graph-based systems.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation</jats:title>\n                  <jats:p>https://github.com/splatlab/debgr.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information</jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.</jats:p>\n               </jats:sec>","is_dataset_classified":null,"base_score":3.7612001156935624,"endowment":3.7612001156935624,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"28881995","pmcid":"PMC5870571","openalex_id":"https://openalex.org/W2735897904","authors":[],"funders":[{"funder_name":"National Science Foundation","grant_id":"BBSRC-NSF/BIO-1564917","title":null},{"funder_name":"National Science Foundation","grant_id":"IIS-1247726","title":null},{"funder_name":"National Science Foundation","grant_id":"IIS-1251137","title":null},{"funder_name":"National Science Foundation","grant_id":"CNS-1408695","title":null},{"funder_name":"National Science Foundation","grant_id":"CCF-1439084","title":null},{"funder_name":"National Science Foundation","grant_id":"CCF-1617618","title":null},{"funder_name":"National Science Foundation","grant_id":"1251137","title":"BIGDATA: Small: DCM: Collaborative Research: An efficient, versatile, scalable, and portable storage system for scientific data containers"},{"funder_name":"National Science Foundation","grant_id":"1247726","title":"BIGDATA: Mid-Scale: DCM: Collaborative Research: Eliminating the Data Ingestion Bottleneck in Big Data Applications"},{"funder_name":"National Science Foundation","grant_id":"1408695","title":"CSR: Medium: Collaborative Research: FTFS: A Read/Write-Optimized Fractal Tree File System"},{"funder_name":"National Science Foundation","grant_id":"1439084","title":"XPS: FULL: CCA: Collaborative Research: Cache-Adaptive Algorithms: How to Share Core among Many Cores"},{"funder_name":"Sandia National Laboratories","grant_id":"","title":null}],"total_grants":11,"fwci":2.2954,"citation_percentile":0.88875959,"influential_citations":0,"citation_trend":[{"year":2017,"count":3},{"year":2018,"count":4},{"year":2019,"count":9},{"year":2020,"count":5},{"year":2021,"count":6},{"year":2022,"count":4},{"year":2023,"count":6},{"year":2024,"count":3},{"year":2025,"count":1},{"year":2026,"count":1}],"oa_status":"bronze","license":"CC BY","oa_locations":[{"url":"https://academic.oup.com/bioinformatics/article-pdf/33/14/i133/25157228/btx261.pdf","host_type":"journal"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/33/14/i133/25157228/btx261.pdf","host_type":"publisher"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/33/14/i133/50315096/bioinformatics_33_14_i133.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/bioinformatics/btx261","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/28881995","host_type":"repository"},{"url":"https://www.osti.gov/biblio/1430372","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/5870571","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC5870571","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC5870571?pdf=render","host_type":"Europe_PMC"},{"url":"https://doi.org/10.7490/f1000research.1114778.1","host_type":""},{"url":"https://dx.doi.org/10.7490/f1000research.1114778.1","host_type":""},{"url":"http://dx.doi.org/10.1093/bioinformatics/btx261","host_type":""},{"url":"https://dx.doi.org/10.1093/bioinformatics/btx261","host_type":""}],"fields_of_study":["Genomics and Phylogenetic Studies","Bioinformatics and Genomic Networks","Genome Rearrangement Algorithms","0301 basic medicine","03 medical and health sciences","0206 medical engineering","02 engineering and technology"],"mesh_terms":["Algorithms","Software","Sequence Analysis, RNA","Computational Biology","Gene Expression Profiling"],"keywords":["De Bruijn graph","De Bruijn sequence","Computer science","Graph","Theoretical computer science","Algorithm","Combinatorics","Mathematics","Sequence Analysis, RNA","Gene Expression Profiling","Computational Biology","Ismb/Eccb 2017: The 25th Annual Conference Intelligent Systems for Molecular Biology Held Jointly with the 16th Annual European Conference on Computational Biology, Prague, Czech Republic, July 21–25, 2017","Algorithms","Software"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[{"name":"gen"}],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-28T11:07:31.325981Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}