{"doi":"10.1145/3459930.3470855","title":"Search feasibility in distributed MS-proteomics big data","abstract":"Making large-scale Mass Spectrometry (MS) data FAIR (Findable, Accessible, Interoperable, Reusable) and democratizing access for the omics research community requires advance access and reuse mechanisms. In this work, we proposed a novel distributed data access infrastructure and developed a simulation test-bed to show the feasibility of this solution. In contrast to existing centralized approaches, participating nodes are relied upon to execute the search algorithm and search based on the comparison of raw spectra is supported as opposed to simple meta-data based searches. Simulation results using networking, stochastic modelling, and queuing theory, illustrated that search times were reduced by up-to 600 times for up-to a total of fifty billion spectra. Proteomics is vital because of the importance proteins to life and their role in state-of-the-art medicine such as custom drug delivery and cancer treatment. MS-based proteomics involves the fragmentation of proteins into peptide ions to generate raw MS spectra. Traditionally, scientists have relied on meta-data based searches of centralized repositories followed by complex database searches and protein sequencing. Though useful, this technique may result in missed datasets because of poor meta-data or sheer amount of effort and computational time needed. Recently, direct raw spectra search has been proposed with the development of centralized tools such as PeptideAtlas. However, PeptideAtlas hosts 13,000 spectra whereas systems supporting billions of spectra are needed. Let us assume users can submit one or more query spectra for search to a central controller. In the proposed novel distributed paradigm, the controller will forward the queries to several nodes hosting a total of multiple MS/MS datasets, where each of the nodes will run the search algorithm against against each spectrum in their local MS/MS dataset, and send the results as URLs/pointers and associated scores back to the controller. The controller will then collate the results and transmit them back to the users. To simulate the system performance, we focused on the distributed process between the controller and the participating nodes. We modeled the the nodes using computational devices present in typical research labs, communication links as the average achievable by combined fiber/Ethernet links, and data loads based on typical storage sizes of spectra and URLs. By running Monte Carlo simulations, we were able to obtain the response time to a single query for various scenarios and assuming an M/M/1 queue, we simulated the time degradation due to multiple requests by compounding over the number of requests with a load degradation factor. Testing results for fifty billion spectra indicated that using 500 distributed nodes can provide search results in 10s and 2000 nodes in 5s, a reduction by 100 and 200 times, respectively, compared to a centralized approach which requires 1000s. Considering typical capabilities of modern day servers and computers, a load factor of 0.001% was tested and indicated that the system provided constant time performance up-to 10k concurrent queries. Lastly, accounting for communication link degradation demonstrated that a trade-off can be achieved between performance and number of nodes. Therefore, it is worth investigating the implementation of a distributed big-data access infrastructure for proteomics.","journal":null,"year":2021,"id":221337,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":1,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9543,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":109046,"name":"Fahad Saeed","orcid":"0000-0002-3410-9552","position":1,"is_corresponding":false},{"id":821359,"name":"Umair Mohammad","orcid":"0000-0002-4499-548X","position":0,"is_corresponding":true}],"reference_count":0,"raw_metadata":null,"created_at":"2026-07-18T23:53:55.215284Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}