{"doi":"10.1145/3745022","title":"Genomics-Enhanced Cancer Risk Prediction for Personalized LLM-Driven Healthcare Recommender Systems","abstract":"<jats:p>Cancer risk prediction is a cornerstone of personalized medicine that offers opportunities for early detection and preventive interventions. However, the current models are designed to predict cancer risk face several challenges. First, most rely on traditional statistical methods, which struggle to capture the complexity of genetic, family medical history, and lifestyle factors. Hence, the accuracy of these models is limited. Additionally, the models neglect to integrate multidimensional data sources, particularly genetic information like single nucleotide polymorphisms (SNPs), which could enhance prediction accuracy. Third, while the system might effectively predict risk, it cannot translate those predictions into actionable healthcare recommendations to reduce cancer risk.</jats:p>\n          <jats:p>\n            In this study, we address all three of these limitations. With a focus on six prevalent cancers—we extracted SNP data from the UK Biobank and designed a novel risk prediction model for cancer and personalized healthcare recommendations based upon the mixture of experts (MoE) paradigm and large language models (LLMs), respectively. Named MoE-HRS, experts based two router networks for separate processing by the Transformer and the convolutional neural network (CNN). Experiments on UK Biobank data show that our model outperforms state-of-the-art cancer risk prediction models. To bridge the gap between risk prediction and practical healthcare applications, we devised a healthcare recommender system powered by LLMs. This approach holds promise for enhancing early detection rates and promoting preventive healthcare management (relevant coding and data are available at\n            <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://github.com/bjtu-lucas-nlp/MoE-HRS)\">https://github.com/bjtu-lucas-nlp/MoE-HRS</jats:ext-link>\n            ).\n          </jats:p>","journal":"ACM Transactions on Information Systems","year":2025,"id":617138,"datarank":0.33195744977248254,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"self_citation_contribution":0.29188652235829704,"citation_network_contribution":0.0400709274141855,"self_endowment_contribution":0.29188652235829704,"citer_contribution":0.0400709274141855,"corpus_percentile":null,"corpus_rank":null,"citation_count":6,"citer_count":5,"citers_with_citation_signal":3,"citers_with_endowment":3,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":135810,"name":"Jie Lu","orcid":"0000-0003-0690-4732","position":1,"is_corresponding":false},{"id":350526,"name":"Hanshi Xu","orcid":"0000-0003-1129-5337","position":2,"is_corresponding":false},{"id":1416786,"name":"Kairui Guo","orcid":"0000-0003-1570-314X","position":3,"is_corresponding":false},{"id":484864,"name":"Qian Zhang","orcid":"0000-0001-7970-5684","position":4,"is_corresponding":false},{"id":343816,"name":"Hua Lin","orcid":"0000-0002-0840-6553","position":5,"is_corresponding":false},{"id":1416790,"name":"Mark Grosser","orcid":"0000-0002-4281-5369","position":6,"is_corresponding":false},{"id":537977,"name":"Yi Zhang","orcid":"0000-0003-1780-0963","position":7,"is_corresponding":false},{"id":1416792,"name":"Guangquan Zhang","orcid":"0000-0003-3960-0583","position":8,"is_corresponding":false},{"id":1591334,"name":"Kezhi Lu","orcid":"0000-0003-4979-5097","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Genomics-Enhanced Cancer Risk Prediction for Personalized LLM-Driven Healthcare Recommender Systems","abstract":"<jats:p>Cancer risk prediction is a cornerstone of personalized medicine that offers opportunities for early detection and preventive interventions. However, the current models are designed to predict cancer risk face several challenges. First, most rely on traditional statistical methods, which struggle to capture the complexity of genetic, family medical history, and lifestyle factors. Hence, the accuracy of these models is limited. Additionally, the models neglect to integrate multidimensional data sources, particularly genetic information like single nucleotide polymorphisms (SNPs), which could enhance prediction accuracy. Third, while the system might effectively predict risk, it cannot translate those predictions into actionable healthcare recommendations to reduce cancer risk.</jats:p>\n          <jats:p>\n            In this study, we address all three of these limitations. With a focus on six prevalent cancers—we extracted SNP data from the UK Biobank and designed a novel risk prediction model for cancer and personalized healthcare recommendations based upon the mixture of experts (MoE) paradigm and large language models (LLMs), respectively. Named MoE-HRS, experts based two router networks for separate processing by the Transformer and the convolutional neural network (CNN). Experiments on UK Biobank data show that our model outperforms state-of-the-art cancer risk prediction models. To bridge the gap between risk prediction and practical healthcare applications, we devised a healthcare recommender system powered by LLMs. This approach holds promise for enhancing early detection rates and promoting preventive healthcare management (relevant coding and data are available at\n            <jats:ext-link xmlns:xlink=\"http://www.w3.org/1999/xlink\" ext-link-type=\"uri\" xlink:href=\"https://github.com/bjtu-lucas-nlp/MoE-HRS)\">https://github.com/bjtu-lucas-nlp/MoE-HRS</jats:ext-link>\n            ).\n          </jats:p>","is_dataset_classified":null,"base_score":1.9459101490553132,"endowment":1.9459101490553132,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"19910364","pmcid":null,"openalex_id":"https://openalex.org/W4411489267","authors":[],"funders":[],"total_grants":0,"fwci":3.5556,"citation_percentile":0.93200409,"influential_citations":0,"citation_trend":[{"year":2025,"count":4},{"year":2026,"count":2}],"oa_status":"closed","license":null,"oa_locations":[{"url":"https://dl.acm.org/doi/pdf/10.1145/3745022","host_type":"publisher"},{"url":"https://doi.org/10.1145/3745022","host_type":"journal"}],"fields_of_study":["Cancer Genomics and Diagnostics","Radiomics and Machine Learning in Medical Imaging","AI in cancer detection"],"mesh_terms":[],"keywords":["Biobank","Computer science","Health care","Personalized medicine","Precision medicine","Predictive modelling","Risk assessment","Artificial intelligence","Risk analysis (engineering)","Data science","Data mining","Machine learning","Bioinformatics","Medicine","Computer security"],"sdg_mappings":[{"sdg_number":0,"sdg_label":"Good health and well-being"}],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-03T00:56:21.899805Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}