{"doi":"10.1137/20m1333110","title":"Test Data Reuse for the Evaluation of Continuously Evolving Classification Algorithms Using the Area under the Receiver Operating Characteristic Curve","abstract":"Performance evaluation of continuously evolving machine learning algorithms presents new challenges, especially for high-risk application domains such as medicine. In principle, to obtain performance measures that generalize to a target population, a new independent test dataset randomly drawn from the target population should be used each time a new performance evaluation is required. However, test datasets of sufficient quality are often hard to acquire, and it is tempting to utilize a previously used test dataset for a new performance evaluation. With extensive experiments on simulated and real data we illustrate how such a \"naive\" approach to test data reuse can inadvertently result in overfitting the algorithm to the test data, resulting in a generalization loss and overly optimistic conclusions about the algorithm performance. We investigate the use of a modified version of the reusable holdout mechanism of Dwork et al. [Science, 349 (2015), pp. 636--638], which allows for repeated reuse of the same test dataset. We extend their approach to the use of AUC, the area under the receiver operating characteristic curve, as the reported performance metric. Theoretical guarantees for our method are proven to hold in extremely data-rich scenarios. However, our empirical results indicate promising performance of the proposed technique even on small data. With extensive simulation studies and experiments on real medical imaging data we show that our procedure indeed substantially reduces the problem of overfitting to the test data, even when the test dataset is small, at the cost of a mild additional uncertainty on the reported test performance.","journal":"SIAM Journal on Mathematics of Data Science","year":2021,"id":194820,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":9,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9316,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2021-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":764846,"name":"Aria Pezeshk","orcid":"0000-0002-3570-3051","position":1,"is_corresponding":false},{"id":764847,"name":"Yuping Wang","orcid":"0000-0001-6868-0004","position":2,"is_corresponding":false},{"id":764848,"name":"Berkman Sahiner","orcid":"0000-0003-2804-2264","position":3,"is_corresponding":false},{"id":764845,"name":"Alexej Gossmann","orcid":"0000-0001-9068-3877","position":0,"is_corresponding":true}],"reference_count":34,"raw_metadata":null,"created_at":"2026-07-18T23:49:59.476757Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}