{"doi":"10.1002/mp.70074","title":"Consistent performance between medical experts and non‐expert readers in forced‐choice lesion‐detection tasks with PET images","abstract":"BACKGROUND: Labeled data are used to train, validate, and test deep learning model observers (DLMOs) as well as linear model observers, such as channelized Hotelling observers (CHOs), for image quality assessment in many imaging modalities, including PET imaging. Ideally, these annotations would come from clinical readers, but these readers are often difficult to access, and it is unclear whether clinical experience is necessary in simple detection tasks involving a known signal profile at a specified location in an image using patient data for backgrounds. PURPOSE: We evaluate the potential for using non-medical observers in simple detection tasks to train DLMOs and CHOs. To this end, we compare nuclear medicine physicians to medically naïve readers using a two-alternative forced choice task with PET images and simulated low-contrast lesions. METHODS: Experimental conditions span 2 locations in the body (liver and lung), 3 lesion contrasts, 3 implementations of 2 different reconstruction algorithms based on ordered subsets expectation maximization and penalized likelihood. The image readers consist of four board-certified nuclear medicine physicians and four non-physician readers. We make direct comparisons of observer performance between the two reader groups and analyze the observer performance data using generalized linear mixed models. In addition, we train two DLMOs using the expert reader data and the non-expert reader data, respectively, and compare the performance of the DLMOs in predicting the performance of (held-out) expert readers. The DLMOs are based on a CNN and a combination of CNN and Transformer (SwinT) encoders. We also compare CHOs trained on the expert reader data and the non-expert reader data. RESULTS: Similarities among the expert readers and between expert and non-expert readers, measured by a concordance metric and Cohen's kappa, were not found to be significantly different (p > 0.11). Analysis of variance found that reader expertise was only marginally significant (p = 0.07) and all interactions involving expertise were non-significant (p > 0.2). Furthermore, the DLMOs trained on experts and on non-experts showed similar performance in predicting experts for both CNN and CNN-SwinT DLMOs (p > 0.26). The CHOs trained on experts and on non-experts showed similar performance (p > 0.31). CONCLUSIONS: We did not find substantive differences between the average performance of medical experts and non-experts. This suggests that in simple detection tasks like these, clinical experience is not necessary, and we can use non-medical observers instead of physicians when training models.","journal":"Medical Physics","year":2025,"id":578826,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.7925,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":760009,"name":"Sangtae Ahn","orcid":"0000-0001-7252-2607","position":1,"is_corresponding":false},{"id":1163622,"name":"Muhan Shao","orcid":"0000-0002-4686-4395","position":2,"is_corresponding":false},{"id":589109,"name":"Darrin Byrd","orcid":null,"position":3,"is_corresponding":false},{"id":916545,"name":"Kristen A. Wangerin","orcid":"0000-0002-9099-3223","position":4,"is_corresponding":false},{"id":974986,"name":"Scott D. Wollenweber","orcid":"0000-0001-6105-6831","position":5,"is_corresponding":false},{"id":657403,"name":"Fatemeh Behnia","orcid":"0000-0002-9110-909X","position":6,"is_corresponding":false},{"id":1284354,"name":"Jean H. Lee","orcid":null,"position":7,"is_corresponding":false},{"id":925823,"name":"Amir Iravani","orcid":"0000-0002-1273-5835","position":8,"is_corresponding":false},{"id":1284355,"name":"Murat Sadıç","orcid":null,"position":9,"is_corresponding":false},{"id":273859,"name":"Delphine L. Chen","orcid":"0000-0002-0150-5401","position":10,"is_corresponding":false},{"id":65783,"name":"Paul E. Kinahan","orcid":"0000-0001-6461-3306","position":11,"is_corresponding":false},{"id":414265,"name":"Craig K. Abbey","orcid":"0000-0003-4608-9402","position":0,"is_corresponding":true}],"reference_count":45,"raw_metadata":null,"created_at":"2026-07-19T02:58:24.957414Z","pmid":"41126367","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}