{"doi":"10.1007/s11227-025-07413-5","title":"Optimizing inference of segmentation on high-resolution images in MLExchange","abstract":"Abstract MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3 $$\\times$$ <mml:math xmlns:mml=\"http://www.w3.org/1998/Math/MathML\"> <mml:mo>×</mml:mo> </mml:math> speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40 $$\\times$$ <mml:math xmlns:mml=\"http://www.w3.org/1998/Math/MathML\"> <mml:mo>×</mml:mo> </mml:math> speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.","journal":"The Journal of Supercomputing","year":2025,"id":539632,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9441,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":577605,"name":"Tanny Chávez","orcid":"0000-0001-9317-2896","position":1,"is_corresponding":false},{"id":1426680,"name":"Wiebke Koepp","orcid":"0000-0002-3234-9368","position":2,"is_corresponding":false},{"id":1133136,"name":"Guanhua Hao","orcid":"0000-0003-3281-6816","position":3,"is_corresponding":false},{"id":542430,"name":"Petrus H. Zwart","orcid":"0000-0003-3315-4092","position":4,"is_corresponding":false},{"id":1133142,"name":"Alexander Hexemer","orcid":"0000-0002-5269-0125","position":5,"is_corresponding":false},{"id":1426679,"name":"Shizhao Lu","orcid":"0000-0003-3055-1877","position":0,"is_corresponding":true}],"reference_count":17,"raw_metadata":null,"created_at":"2026-07-19T02:52:30.048313Z","pmid":"40546286","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}