{"doi":"10.1002/mp.70206","title":"Uncertainty‐guided test‐time optimization for personalizing segmentation models in longitudinal medical imaging","abstract":"BACKGROUND: Accurate and consistent image segmentation across longitudinal scans is essential in many clinical applications, including surveillance, treatment monitoring, and adaptive interventions. While personalized model adaptation using patient-specific prior scans has shown promise, current approaches typically rely on fixed training durations and lack mechanisms to determine optimal stopping points on a per-patient basis, particularly in the absence of validation labels. PURPOSE: We propose an uncertainty-guided test-time optimization (TTO) framework that dynamically adjusts the personalization duration for each patient using a validation-free stopping criterion based on predictive uncertainty. METHODS: Our framework personalizes a generalized segmentation model using patient-specific prior imaging and selects the optimal checkpoint based on the minimum voxel-wise predictive uncertainty, estimated via Monte Carlo Dropout (TTO-MCD) or Deep Ensembling (TTO-DE). We evaluated the approach on three datasets: 214 pancreas (CT) scans, 243 liver (CT) scans, and 175 head-and-neck tumor (MRI) scans, each containing a subset of patients with paired longitudinal scans to enable patient-specific personalization. Each patient's follow-up scan was held out for testing. As a baseline, we implemented a fixed-epoch personalization strategy (Pre-TTO) using a fivefold cross-test design to emulate deployable model selection without test label leakage. RESULTS: TTO methods consistently outperformed the Pre-TTO and unpersonalized baseline across standard metrics, including the Dice Similarity Coefficient (DSC), 95th percentile Hausdorff Distance (HD95), Mean Surface Distance (MSD), and the proposed LogPenalty Score (LPS), which provides a bounded, interpretable scale that jointly reflects volumetric and boundary fidelity. Paired t-tests confirmed statistically significant improvements for pancreas and liver datasets (p < 0.05), while favorable trends were observed in the head-and-neck dataset despite greater anatomical variability. Both TTO-MCD and TTO-DE achieved near-optimal performance without requiring access to labels at test time. CONCLUSION: Uncertainty-guided TTO provides a robust, validation-free strategy for optimizing patient-specific segmentation models in longitudinal medical imaging. By tailoring personalization based on predictive uncertainty, our method improves segmentation quality across a range of imaging modalities and anatomical targets. This framework supports broad clinical deployment of personalized AI and motivates future extensions to contextual integration and multi-label segmentation. Code is publicly available at https://github.com/jchun-ai/uncertainty-tto.","journal":"Medical Physics","year":2025,"id":586615,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9519,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1196112,"name":"Austin H. Castelo","orcid":"0009-0009-3015-3570","position":1,"is_corresponding":false},{"id":1167810,"name":"McKell Woodland","orcid":"0000-0003-0995-3695","position":2,"is_corresponding":false},{"id":1465595,"name":"Caleb O'Connor","orcid":null,"position":3,"is_corresponding":false},{"id":1141615,"name":"Mais M Al Taie","orcid":null,"position":4,"is_corresponding":false},{"id":1196113,"name":"Mohamed Eltaher","orcid":"0009-0008-2078-4718","position":5,"is_corresponding":false},{"id":430415,"name":"Aashish C. Gupta","orcid":"0000-0002-0178-0921","position":6,"is_corresponding":false},{"id":1501618,"name":"Bilel Daoud","orcid":null,"position":7,"is_corresponding":false},{"id":1376203,"name":"Shanli Ding","orcid":null,"position":8,"is_corresponding":false},{"id":1501619,"name":"Jeddy Bennett","orcid":null,"position":9,"is_corresponding":false},{"id":68527,"name":"Anirban Maitra","orcid":"0000-0001-7923-9978","position":10,"is_corresponding":false},{"id":404907,"name":"Matthew A. Firpo","orcid":"0000-0002-7983-3982","position":11,"is_corresponding":false},{"id":336041,"name":"Kimberly S. Kirkwood","orcid":"0000-0002-6876-2752","position":12,"is_corresponding":false},{"id":227744,"name":"Eugene J. Koay","orcid":"0000-0001-7675-3461","position":13,"is_corresponding":false},{"id":310388,"name":"Kristy K. Brock","orcid":"0000-0001-9364-5040","position":14,"is_corresponding":false},{"id":881960,"name":"Jaehee Chun","orcid":"0000-0002-9695-6079","position":0,"is_corresponding":true}],"reference_count":22,"raw_metadata":null,"created_at":"2026-07-19T02:59:32.191237Z","pmid":"41423658","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}