{"doi":"10.1093/jamiaopen/ooaf070","title":"Large-scale deep learning for metastasis detection in pathology reports","abstract":"Objectives: No existing algorithm can reliably identify metastasis from pathology reports across multiple cancer types and the entire US population. In this study, we develop a deep learning model that automatically detects patients with metastatic cancer by using pathology reports from many laboratories and of multiple cancer types. Materials and Methods: We use 60 471 unstructured pathology reports from 4 Surveillance, Epidemiology, and End Results (SEER) registries. The reports were coded into 1 of 3 labels: metastasis negative, metastases positive, or metastasis undetermined. We utilize a task-specific deep neural network trained from scratch and compare its performance with a widely used large language model (LLM). Results: Our deep learning architecture trained on task-specific data outperforms a general-purpose LLM, with a recall of 0.894 compared to 0.824. We quantified model uncertainty and used it to defer reports for human review. We found that retaining 72.9% of reports increased recall from 0.894 to 0.969. Discussion: A smaller deep learning architecture trained on task-specific data outperforms a general LLM. Equally critical to model performance is the incorporation of uncertainty quantification, achieved here through an abstention mechanism. Conclusions: This study's finding demonstrate the feasibility of developing algorithms to automatically identify metastatic cancer cases from unstructured pathology reports.","journal":"JAMIA Open","year":2025,"id":570094,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.7669,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":434090,"name":"Zachary Fox","orcid":"0000-0003-0138-2867","position":1,"is_corresponding":false},{"id":447318,"name":"Valentina I. Petkov","orcid":"0009-0009-8297-5381","position":2,"is_corresponding":false},{"id":1033598,"name":"Serban Negoita","orcid":"0000-0002-9327-9519","position":3,"is_corresponding":false},{"id":309,"name":"Jennifer Anne Doherty","orcid":"0000-0002-1454-8187","position":4,"is_corresponding":false},{"id":298442,"name":"Antoinette M. Stroup","orcid":"0000-0003-3341-4018","position":5,"is_corresponding":false},{"id":375900,"name":"Stephen M. Schwartz","orcid":"0000-0001-7499-8502","position":6,"is_corresponding":false},{"id":356471,"name":"Lynne Penberthy","orcid":"0000-0001-9372-9869","position":7,"is_corresponding":false},{"id":1268398,"name":"Elizabeth Hsu","orcid":"0000-0003-3762-9315","position":8,"is_corresponding":false},{"id":378744,"name":"John Gounley","orcid":"0000-0001-8424-4982","position":9,"is_corresponding":false},{"id":329182,"name":"Heidi A. Hanson","orcid":"0000-0003-0056-196X","position":10,"is_corresponding":false},{"id":810380,"name":"Patrycja Krawczuk","orcid":"0000-0003-4860-9619","position":0,"is_corresponding":true}],"reference_count":35,"raw_metadata":null,"created_at":"2026-07-19T02:57:07.857542Z","pmid":"40655537","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}