{"doi":"10.1109/icmlc48188.2019.8949260","title":"Exploring and Evaluating the Scalability and Eficinecy of Apache Spark Using Educational Datasets","abstract":null,"journal":"2019 International Conference on Machine Learning and Cybernetics (ICMLC)","year":2019,"id":608055,"datarank":0.10397207708399181,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"self_citation_contribution":0.10397207708399181,"citation_network_contribution":0.0,"self_endowment_contribution":0.10397207708399181,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":1,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":829620,"name":"Zijiang Yang","orcid":"0000-0002-3736-9077","position":1,"is_corresponding":false},{"id":1561516,"name":"Younes Benslimane","orcid":null,"position":2,"is_corresponding":false},{"id":351484,"name":"Jian Zhang","orcid":"0000-0003-3838-7389","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Exploring and Evaluating the Scalability and Eficinecy of Apache Spark Using Educational Datasets","abstract":"The combination of data mining and machine learning technology with web-based education system is becoming an imperative research area to enhance the quality of education beyond the traditional concept. With the worldwide fast growth of the Information Communication Technology (ICT), data come with significant large volume, high velocity and extensive variety. In this paper, four popular data mining methods are applied on Apache Spark using large volume of datasets from Online Cognitive Learning Systems to explore the scalability and efficiency of Spark. Various volumes of datasets are tested on Spark MLlib with different running configurations and parameter tunings. The output of the paper convincingly presents useful strategies of computing resource allocation and tuning to make full advantage of the in-memory system of Apache Spark with the tasks of data mining and machine learning on educational datasets.","is_dataset_classified":null,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"23304386","pmcid":null,"openalex_id":"https://openalex.org/W2901787049","authors":[],"funders":[],"total_grants":0,"fwci":null,"citation_percentile":null,"influential_citations":0,"citation_trend":[{"year":2023,"count":1}],"oa_status":"closed","license":"https://doi.org/10.15223/policy-029","oa_locations":[{"url":"http://xplorestaging.ieee.org/ielx7/8942645/8949161/08949260.pdf?arnumber=8949260","host_type":"publisher"},{"url":"https://doi.org/10.1109/icmlc48188.2019.8949260","host_type":""}],"fields_of_study":["Online Learning and Analytics","Intelligent Tutoring Systems and Adaptive Learning","Distributed and Parallel Computing Systems"],"mesh_terms":[],"keywords":["SPARK (programming language)","Scalability","Computer science","Volume (thermodynamics)","Big data","Machine learning","Resource (disambiguation)","Variety (cybernetics)","Quality (philosophy)","Data mining","Data science","Artificial intelligence","Database"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-30T07:26:18.623695Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}