{"doi":"10.1093/jamiaopen/ooac049","title":"Automatic information extraction from childhood cancer pathology reports","abstract":"Objectives: The International Classification of Childhood Cancer (ICCC) facilitates the effective classification of a heterogeneous group of cancers in the important pediatric population. However, there has been no development of machine learning models for the ICCC classification. We developed deep learning-based information extraction models from cancer pathology reports based on the ICD-O-3 coding standard. In this article, we describe extending the models to perform ICCC classification. Materials and Methods: We developed 2 models, ICD-O-3 classification and ICCC recoding (Model 1) and direct ICCC classification (Model 2), and 4 scenarios subject to the training sample size. We evaluated these models with a corpus consisting of 29 206 reports with age at diagnosis between 0 and 19 from 6 state cancer registries. Results: Our findings suggest that the direct ICCC classification (Model 2) is substantially better than reusing the ICD-O-3 classification model (Model 1). Applying the uncertainty quantification mechanism to assess the confidence of the algorithm in assigning a code demonstrated that the model achieved a micro-F1 score of 0.987 while abstaining (not sufficiently confident to assign a code) on only 14.8% of ambiguous pathology reports. Conclusions: Our experimental results suggest that the machine learning-based automatic information extraction from childhood cancer pathology reports in the ICCC is a reliable means of supplementing human annotators at state cancer registries by reading and abstracting the majority of the childhood cancer pathology reports accurately and reliably.","journal":"JAMIA Open","year":2022,"id":281458,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":8,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9512,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2022-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":65507,"name":"Alina Peluso","orcid":"0000-0003-2895-0406","position":1,"is_corresponding":false},{"id":85407,"name":"Eric B. Durbin","orcid":"0000-0002-5600-7645","position":2,"is_corresponding":false},{"id":353078,"name":"Xiao‐Cheng Wu","orcid":"0000-0003-3663-5027","position":3,"is_corresponding":false},{"id":298442,"name":"Antoinette M. Stroup","orcid":"0000-0003-3341-4018","position":4,"is_corresponding":false},{"id":309,"name":"Jennifer Anne Doherty","orcid":"0000-0002-1454-8187","position":5,"is_corresponding":false},{"id":375900,"name":"Stephen M. Schwartz","orcid":"0000-0001-7499-8502","position":6,"is_corresponding":false},{"id":420945,"name":"Charles L. Wiggins","orcid":"0000-0002-0661-8261","position":7,"is_corresponding":false},{"id":234049,"name":"Linda Coyle","orcid":null,"position":8,"is_corresponding":false},{"id":356471,"name":"Lynne Penberthy","orcid":"0000-0001-9372-9869","position":9,"is_corresponding":false},{"id":356468,"name":"Hong‐Jun Yoon","orcid":"0000-0002-5450-5878","position":0,"is_corresponding":true}],"reference_count":13,"raw_metadata":null,"created_at":"2026-07-19T00:29:11.698554Z","pmid":"35721398","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}