{"doi":"10.3390/bdcc7020098","title":"Massive Parallel Alignment of RNA-seq Reads in Serverless Computing","abstract":"<jats:p>In recent years, the use of Cloud infrastructures for data processing has proven useful, with a computing potential that is not affected by the limitations of a local infrastructure. In this context, Serverless computing is the fastest-growing Cloud service model due to its auto-scaling methodologies, reliability, and fault tolerance. We present a solution based on in-house Serverless infrastructure, which is able to perform large-scale RNA-seq data analysis focused on the mapping of sequencing reads to a reference genome. The main contribution was bringing the computation of genomic data into serverless computing, focusing on RNA-seq read-mapping to a reference genome, as this is the most time-consuming task for some pipelines. The proposed solution handles massive parallel instances to maximize the efficiency in terms of running time. We evaluated the performance of our solution by performing two main tests, both based on the mapping of RNA-seq reads to Human GRCh38. Our experiments demonstrated a reduction of 79.838%, 90.079%, and 96.382%, compared to the local environments with 16, 8, and 4 virtual cores, respectively. Furthermore, serverless limitations were investigated.</jats:p>","journal":"Big Data and Cognitive Computing","year":2023,"id":590016,"datarank":0.5250539407029695,"base_score":2.3978952727983707,"endowment":2.3978952727983707,"self_citation_contribution":0.3596842909197557,"citation_network_contribution":0.1653696497832139,"self_endowment_contribution":0.3596842909197557,"citer_contribution":0.1653696497832139,"corpus_percentile":null,"corpus_rank":null,"citation_count":10,"citer_count":4,"citers_with_citation_signal":2,"citers_with_endowment":2,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1509632,"name":"José Luis Vázquez-Poletti","orcid":"0000-0002-6241-8141","position":1,"is_corresponding":false},{"id":554573,"name":"Mario Cannataro","orcid":"0000-0003-1502-2387","position":2,"is_corresponding":false},{"id":1509631,"name":"Pietro Cinaglia","orcid":"0000-0003-2237-6984","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Massive Parallel Alignment of RNA-seq Reads in Serverless Computing","abstract":"<jats:p>In recent years, the use of Cloud infrastructures for data processing has proven useful, with a computing potential that is not affected by the limitations of a local infrastructure. In this context, Serverless computing is the fastest-growing Cloud service model due to its auto-scaling methodologies, reliability, and fault tolerance. We present a solution based on in-house Serverless infrastructure, which is able to perform large-scale RNA-seq data analysis focused on the mapping of sequencing reads to a reference genome. The main contribution was bringing the computation of genomic data into serverless computing, focusing on RNA-seq read-mapping to a reference genome, as this is the most time-consuming task for some pipelines. The proposed solution handles massive parallel instances to maximize the efficiency in terms of running time. We evaluated the performance of our solution by performing two main tests, both based on the mapping of RNA-seq reads to Human GRCh38. Our experiments demonstrated a reduction of 79.838%, 90.079%, and 96.382%, compared to the local environments with 16, 8, and 4 virtual cores, respectively. Furthermore, serverless limitations were investigated.</jats:p>","is_dataset_classified":null,"base_score":2.3978952727983707,"endowment":2.3978952727983707,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"26657633","pmcid":null,"openalex_id":"https://openalex.org/W4376614239","authors":[],"funders":[],"total_grants":0,"fwci":1.3326,"citation_percentile":0.79747429,"influential_citations":0,"citation_trend":[{"year":2023,"count":3},{"year":2024,"count":2},{"year":2025,"count":5}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://www.mdpi.com/2504-2289/7/2/98/pdf?version=1684130944","host_type":"journal"},{"url":"https://www.mdpi.com/2504-2289/7/2/98/pdf?version=1684130944","host_type":"publisher"},{"url":"https://www.mdpi.com/2504-2289/7/2/98/pdf","host_type":"publisher"},{"url":"https://doi.org/10.3390/bdcc7020098","host_type":"journal"},{"url":"https://doaj.org/article/bba83cacae9941489119cf7923eef018","host_type":"repository"},{"url":"https://dx.doi.org/10.3390/bdcc7020098","host_type":"repository"}],"fields_of_study":["Caching and Content Delivery","Advanced Data Storage Technologies","Genomics and Phylogenetic Studies"],"mesh_terms":[],"keywords":["Computer science","Cloud computing","Distributed computing","Context (archaeology)","Fault tolerance","Computation","Service (business)","Reliability (semiconductor)","Task (project management)","Reference genome","Data mining","Genome","Operating system","Biology","Algorithm"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Industry, innovation and infrastructure"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-24T12:19:24.443654Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}