{"doi":"10.5753/sbbd.2017.174092","title":"Spark Scalability Analysis in a Scientific Workflow","abstract":"<jats:p>Spark is being successfully used for big data parallel processing in many business domains (social media, finance, retail). Spark’s scalability, usability, and large user community have motivated developers from scientific domains (bioinformatics, oil and gas, astronomy) to try it. However, scientific applications’ profile, e.g., black-box programs and intense file writes, differs from traditional business workflows, which may affect its scalability. We present a scalability analysis of Spark in a real case-study in Oil and Gas domain. We explore workloads on a 936-cores HPC cluster processing 330 GB of scientific data. We show that it scales very well when running long-lasting scientific tasks, but its performance is lower for short-duration tasks.</jats:p>","journal":"Anais do XXXII Simpósio Brasileiro de Banco de Dados (SBBD 2017)","year":2017,"id":606310,"datarank":0.24141568686511508,"base_score":1.6094379124341003,"endowment":1.6094379124341003,"self_citation_contribution":0.24141568686511508,"citation_network_contribution":0.0,"self_endowment_contribution":0.24141568686511508,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":4,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1556515,"name":"Vítor Silva","orcid":null,"position":1,"is_corresponding":false},{"id":1556516,"name":"Pedro Miranda","orcid":null,"position":2,"is_corresponding":false},{"id":1556517,"name":"Alexandre A. B. Lima","orcid":null,"position":3,"is_corresponding":false},{"id":1556518,"name":"Patrick Valduriez","orcid":null,"position":4,"is_corresponding":false},{"id":1556519,"name":"Marta Mattoso","orcid":null,"position":5,"is_corresponding":false},{"id":1556514,"name":"Renan Souza","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Spark Scalability Analysis in a Scientific Workflow","abstract":"<jats:p>Spark is being successfully used for big data parallel processing in many business domains (social media, finance, retail). Spark’s scalability, usability, and large user community have motivated developers from scientific domains (bioinformatics, oil and gas, astronomy) to try it. However, scientific applications’ profile, e.g., black-box programs and intense file writes, differs from traditional business workflows, which may affect its scalability. We present a scalability analysis of Spark in a real case-study in Oil and Gas domain. We explore workloads on a 936-cores HPC cluster processing 330 GB of scientific data. We show that it scales very well when running long-lasting scientific tasks, but its performance is lower for short-duration tasks.</jats:p>","is_dataset_classified":null,"base_score":1.6094379124341003,"endowment":1.6094379124341003,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"21097893","pmcid":null,"openalex_id":"https://openalex.org/W2792181990","authors":[],"funders":[{"funder_name":"European Commission","grant_id":"689772","title":"HPC for Energy"}],"total_grants":1,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2018,"count":1},{"year":2019,"count":1},{"year":2021,"count":2}],"oa_status":"gold","license":"other-oa","oa_locations":[{"url":"https://sol.sbc.org.br/index.php/sbbd/article/download/24300/24123","host_type":""},{"url":"https://sol.sbc.org.br/index.php/sbbd/article/download/24300/24123","host_type":""},{"url":"http://dx.doi.org/10.5753/sbbd.2017.174092","host_type":""},{"url":"https://hal-lirmm.ccsd.cnrs.fr/lirmm-01620161","host_type":"repository"},{"url":"https://doi.org/10.5753/sbbd.2017.174092","host_type":""},{"url":"https://hal-lirmm.ccsd.cnrs.fr/lirmm-01620161v1","host_type":""}],"fields_of_study":["Scientific Computing and Data Management","Distributed and Parallel Computing Systems","Cloud Computing and Resource Management","0202 electrical engineering, electronic engineering, information engineering","02 engineering and technology"],"mesh_terms":[],"keywords":["Scalability","SPARK (programming language)","Workflow","Big data","Computer science","Usability","Data science","Domain (mathematical analysis)","Database","Data mining","Operating system","Programming language","[INFO.INFO-DC] Computer Science [cs]/Distributed, Parallel, and Cluster Computing [cs.DC]","[INFO.INFO-DB] Computer Science [cs]/Databases [cs.DB]"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Industry, innovation and infrastructure"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-30T04:17:43.325548Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}