{"doi":"10.1101/2020.12.22.423964","title":"LISA: A Case For Learned Index based Acceleration of Biological Sequence Analysis","abstract":"<h4>ABSTRACT</h4> Next Generation Sequencing (NGS) is transforming fields like genomics, transcriptomics, and epigenetics with rapidly increasing throughput at reduced cost. This also demands overcoming performance bottlenecks in the downstream analysis of the sequencing data. A key performance bottleneck is searching for exact matches of entire or substrings of short DNA/RNA sequence queries in a long reference sequence database. This task is typically performed by using an index of the reference - such as FM-index, suffix arrays, suffix trees, hash tables, or lookup tables. In this paper, we propose accelerating this sequence search by substituting or enhancing the indexes with machine learning based indexes - called learned indexes - and present LISA (Learned Indexes for Sequence Analysis). We evaluate LISA through a number of case studies – that cover widely used software tools; short and long reads; human, animal, and plant genome datasets; DNA and RNA sequences; various traditional indexing techniques (FM-indexes, hash tables and suffix arrays) – and demonstrate significant performance benefits in a majority of them. For example, our experiments on real datasets show that LISA achieves speedups of up to 2.2 fold and 4.7 fold over the state-of-the-art FM-index based implementations for exact sequence search modules in popular tools bowtie2 and BWA-MEM2, respectively. <h4>Code availability</h4> LISA-based FM-index: https://github.com/IntelLabs/Trans-Omics-Acceleration-Library/tree/master/src/LISA-FMI LISA-based hash-table: https://github.com/IntelLabs/Trans-Omics-Acceleration-Library/tree/master/src/LISA-hash LISA applied to BWA-MEM2: https://github.com/bwa-mem2/bwa-mem2/tree/bwa-mem2-lisa .","journal":"bioRxiv (Cold Spring Harbor Laboratory)","year":2020,"id":7632,"datarank":0.24141568686511508,"base_score":1.6094379124341003,"endowment":1.6094379124341003,"self_citation_contribution":0.24141568686511508,"citation_network_contribution":0.0,"self_endowment_contribution":0.24141568686511508,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":4,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.0428,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2020-12-22","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":68777,"name":"Saurabh Kalikar","orcid":null,"position":1,"is_corresponding":false},{"id":45197,"name":"Sanchit Misra","orcid":"0000-0001-7863-858X","position":2,"is_corresponding":false},{"id":68778,"name":"Jialin Ding","orcid":null,"position":3,"is_corresponding":false},{"id":68779,"name":"Vasimuddin Md","orcid":"0000-0001-7615-0388","position":4,"is_corresponding":false},{"id":68780,"name":"Nesime Tatbul","orcid":"0000-0002-0416-7022","position":5,"is_corresponding":false},{"id":30887,"name":"Alexandra P. Lewis","orcid":"0000-0002-6195-4786","position":6,"is_corresponding":false},{"id":68781,"name":"Tim Kraska","orcid":"0009-0003-2414-2759","position":7,"is_corresponding":false},{"id":68776,"name":"Darryl Ho","orcid":null,"position":0,"is_corresponding":true}],"reference_count":63,"raw_metadata":null,"created_at":"2026-03-01T18:20:47.508186Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}