{"doi":"10.1093/bioinformatics/btad264","title":"isONform: reference-free transcriptome reconstruction from Oxford Nanopore data","abstract":"<jats:title>Abstract</jats:title>\n               <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>With advances in long-read transcriptome sequencing, we can now fully sequence transcripts, which greatly improves our ability to study transcription processes. A popular long-read transcriptome sequencing technique is Oxford Nanopore Technologies (ONT), which through its cost-effective sequencing and high throughput, has the potential to characterize the transcriptome in a cell. However, due to transcript variability and sequencing errors, long cDNA reads need substantial bioinformatic processing to produce a set of isoform predictions from the reads. Several genome and annotation-based methods exist to produce transcript predictions. However, such methods require high-quality genomes and annotations and are limited by the accuracy of long-read splice aligners. In addition, gene families with high heterogeneity may not be well represented by a reference genome and would benefit from reference-free analysis. Reference-free methods to predict transcripts from ONT, such as RATTLE, exist, but their sensitivity is not comparable to reference-based approaches.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>We present isONform, a high-sensitivity algorithm to construct isoforms from ONT cDNA sequencing data. The algorithm is based on iterative bubble popping on gene graphs built from fuzzy seeds from the reads. Using simulated, synthetic, and biological ONT cDNA data, we show that isONform has substantially higher sensitivity than RATTLE albeit with some loss in precision. On biological data, we show that isONform’s predictions have substantially higher consistency with the annotation-based method StringTie2 compared with RATTLE. We believe isONform can be used both for isoform construction for organisms without well-annotated genomes and as an orthogonal method to verify predictions of reference-based methods.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation</jats:title>\n                  <jats:p>https://github.com/aljpetri/isONform</jats:p>\n               </jats:sec>","journal":"Bioinformatics","year":2023,"id":634867,"datarank":0.4566783656585135,"base_score":3.044522437723423,"endowment":3.044522437723423,"self_citation_contribution":0.4566783656585135,"citation_network_contribution":0.0,"self_endowment_contribution":0.4566783656585135,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":20,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":553654,"name":"Kristoffer Sahlin","orcid":"0000-0001-7378-2320","position":1,"is_corresponding":false},{"id":1646801,"name":"Alexander J Petri","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"isONform: reference-free transcriptome reconstruction from Oxford Nanopore data","abstract":"<jats:title>Abstract</jats:title>\n               <jats:sec>\n                  <jats:title>Motivation</jats:title>\n                  <jats:p>With advances in long-read transcriptome sequencing, we can now fully sequence transcripts, which greatly improves our ability to study transcription processes. A popular long-read transcriptome sequencing technique is Oxford Nanopore Technologies (ONT), which through its cost-effective sequencing and high throughput, has the potential to characterize the transcriptome in a cell. However, due to transcript variability and sequencing errors, long cDNA reads need substantial bioinformatic processing to produce a set of isoform predictions from the reads. Several genome and annotation-based methods exist to produce transcript predictions. However, such methods require high-quality genomes and annotations and are limited by the accuracy of long-read splice aligners. In addition, gene families with high heterogeneity may not be well represented by a reference genome and would benefit from reference-free analysis. Reference-free methods to predict transcripts from ONT, such as RATTLE, exist, but their sensitivity is not comparable to reference-based approaches.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Results</jats:title>\n                  <jats:p>We present isONform, a high-sensitivity algorithm to construct isoforms from ONT cDNA sequencing data. The algorithm is based on iterative bubble popping on gene graphs built from fuzzy seeds from the reads. Using simulated, synthetic, and biological ONT cDNA data, we show that isONform has substantially higher sensitivity than RATTLE albeit with some loss in precision. On biological data, we show that isONform’s predictions have substantially higher consistency with the annotation-based method StringTie2 compared with RATTLE. We believe isONform can be used both for isoform construction for organisms without well-annotated genomes and as an orthogonal method to verify predictions of reference-based methods.</jats:p>\n               </jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation</jats:title>\n                  <jats:p>https://github.com/aljpetri/isONform</jats:p>\n               </jats:sec>","is_dataset_classified":null,"base_score":3.044522437723423,"endowment":3.044522437723423,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"37387174","pmcid":"PMC10311309","openalex_id":"https://openalex.org/W4382631187","authors":[],"funders":[{"funder_name":"Swedish Research Council","grant_id":"2021-04000","title":null},{"funder_name":"Swedish Research Council","grant_id":"unidentified","title":"unidentified"}],"total_grants":2,"fwci":1.8548,"citation_percentile":0.85884521,"influential_citations":0,"citation_trend":[{"year":2023,"count":1},{"year":2024,"count":5},{"year":2025,"count":9},{"year":2026,"count":5}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://academic.oup.com/bioinformatics/article-pdf/39/Supplement_1/i222/50741745/btad264.pdf","host_type":"journal"},{"url":"https://academic.oup.com/bioinformatics/article-pdf/39/Supplement_1/i222/50741745/btad264.pdf","host_type":"publisher"},{"url":"https://doi.org/10.1093/bioinformatics/btad264","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/37387174","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/10311309","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC10311309/pdf/btad264.pdf","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC10311309","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC10311309?pdf=render","host_type":"Europe_PMC"},{"url":"http://dx.doi.org/10.1093/bioinformatics/btad264","host_type":""},{"url":"https://publications-affiliated.scilifelab.se/publication/7755fb9e6e24401e8ce62f85bdee5f8d","host_type":""}],"fields_of_study":["Genomics and Phylogenetic Studies","Genomics and Chromatin Dynamics","Single-cell and spatial transcriptomics","0206 medical engineering","02 engineering and technology","DNA, Complementary","Nanopores","Transcriptome","Algorithms","Computational Biology"],"mesh_terms":["Algorithms","DNA, Complementary","Computational Biology","Nanopores","Transcriptome"],"keywords":["Computational biology","Nanopore sequencing","Genome","Annotation","Sequence assembly","Reference genome","Computer science","Genomics","Biology","DNA sequencing","Transcriptome","De novo transcriptome assembly","Data mining","Gene","Genetics","Gene expression","Genome Sequence Analysis","Nanopores","DNA, Complementary","Algorithms"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-06T14:17:42.086174Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}