{"doi":"10.3389/fbinf.2025.1727953","title":"KG2ML: integrating knowledge graphs and positive unlabeled learning for identifying disease-associated genes","abstract":"<jats:sec>\n                    <jats:title>Background</jats:title>\n                    <jats:p>Biomedical knowledge graphs (KGs), such as the Data Distillery Knowledge Graph (DDKG), capture known relationships among entities (e.g., genes, diseases, proteins), providing valuable insights for research. However, these relationships are typically derived from prior studies, leaving potential unknown associations unexplored. Identifying such unknown associations, including previously unknown disease-associated genes, remains a critical challenge in bioinformatics and is crucial for advancing biomedical knowledge.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods</jats:title>\n                    <jats:p>Traditional methods, such as linkage analysis and genome-wide association studies (GWAS), can be time-consuming and resource-intensive. This highlights the need for efficient computational approaches to identify or predict new genes using known disease-gene associations. Recently, network-based methods and KGs, enhanced by advances in machine learning (ML) frameworks, have emerged as promising tools for inferring these unexplored associations. Given the technical limitations of the Neo4j Graph Data Science (GDS) machine learning pipeline, we developed a novel machine learning pipeline called KG2ML (Knowledge Graph to Machine Learning). This pipeline utilizes our Positive and Unlabeled (PU) learning algorithm, PULSCAR (Positive Unlabeled Learning Selected Completely At Random), and incorporates path-based feature extraction from ProteinGraphML.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Results</jats:title>\n                    <jats:p>KG2ML was applied to 12 diseases, including Bipolar Disorder, Coronary Artery Disease, and Parkinson’s Disease, to infer disease-associated genes not explicitly recorded in DDKG. For several of these diseases, 14 out of the 15 top-ranked genes lacked prior explicit associations in the DDKG but were supported by literature and TINX (Target Importance and Novelty Explorer) evidence. Incorporating PULSCAR-imputed genes as positives enhanced XGBoost classification, demonstrating the potential of PU learning in identifying hidden gene-disease relationships.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusion</jats:title>\n                    <jats:p>The observed improvement in classification performance after the inclusion of PULSCAR-imputed genes as positive examples, along with the subject matter experts’ (SME) evaluations of the top 15 imputed genes for 12 diseases, suggests that PU learning can effectively uncover disease-gene associations missing from existing knowledge graphs (KGs). By integrating KG data with ML-based inference, our KG2ML pipeline provides a scalable and interpretable framework to advance biomedical research while addressing the inherent limitations of current KGs.</jats:p>\n                  </jats:sec>","journal":"Frontiers in Bioinformatics","year":2026,"id":593740,"datarank":0.10397207708399181,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"self_citation_contribution":0.10397207708399181,"citation_network_contribution":0.0,"self_endowment_contribution":0.10397207708399181,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":1,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":846317,"name":"Vincent T. Metzger","orcid":"0000-0002-8041-0370","position":1,"is_corresponding":false},{"id":1457645,"name":"Swastika T. Purushotham","orcid":null,"position":2,"is_corresponding":false},{"id":1457646,"name":"Priyansh Kedia","orcid":null,"position":3,"is_corresponding":false},{"id":1519730,"name":"Cristian G. Bologa","orcid":null,"position":4,"is_corresponding":false},{"id":63962,"name":"Christophe G. Lambert","orcid":"0000-0003-1994-2893","position":5,"is_corresponding":false},{"id":68164,"name":"Jeremy J. Yang","orcid":"0000-0002-1476-6192","position":6,"is_corresponding":false},{"id":361660,"name":"Praveen Kumar","orcid":"0000-0002-7818-671X","position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"KG2ML: integrating knowledge graphs and positive unlabeled learning for identifying disease-associated genes","abstract":"<jats:sec>\n                    <jats:title>Background</jats:title>\n                    <jats:p>Biomedical knowledge graphs (KGs), such as the Data Distillery Knowledge Graph (DDKG), capture known relationships among entities (e.g., genes, diseases, proteins), providing valuable insights for research. However, these relationships are typically derived from prior studies, leaving potential unknown associations unexplored. Identifying such unknown associations, including previously unknown disease-associated genes, remains a critical challenge in bioinformatics and is crucial for advancing biomedical knowledge.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods</jats:title>\n                    <jats:p>Traditional methods, such as linkage analysis and genome-wide association studies (GWAS), can be time-consuming and resource-intensive. This highlights the need for efficient computational approaches to identify or predict new genes using known disease-gene associations. Recently, network-based methods and KGs, enhanced by advances in machine learning (ML) frameworks, have emerged as promising tools for inferring these unexplored associations. Given the technical limitations of the Neo4j Graph Data Science (GDS) machine learning pipeline, we developed a novel machine learning pipeline called KG2ML (Knowledge Graph to Machine Learning). This pipeline utilizes our Positive and Unlabeled (PU) learning algorithm, PULSCAR (Positive Unlabeled Learning Selected Completely At Random), and incorporates path-based feature extraction from ProteinGraphML.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Results</jats:title>\n                    <jats:p>KG2ML was applied to 12 diseases, including Bipolar Disorder, Coronary Artery Disease, and Parkinson’s Disease, to infer disease-associated genes not explicitly recorded in DDKG. For several of these diseases, 14 out of the 15 top-ranked genes lacked prior explicit associations in the DDKG but were supported by literature and TINX (Target Importance and Novelty Explorer) evidence. Incorporating PULSCAR-imputed genes as positives enhanced XGBoost classification, demonstrating the potential of PU learning in identifying hidden gene-disease relationships.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusion</jats:title>\n                    <jats:p>The observed improvement in classification performance after the inclusion of PULSCAR-imputed genes as positive examples, along with the subject matter experts’ (SME) evaluations of the top 15 imputed genes for 12 diseases, suggests that PU learning can effectively uncover disease-gene associations missing from existing knowledge graphs (KGs). By integrating KG data with ML-based inference, our KG2ML pipeline provides a scalable and interpretable framework to advance biomedical research while addressing the inherent limitations of current KGs.</jats:p>\n                  </jats:sec>","is_dataset_classified":null,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"41584517","pmcid":"PMC12823822","openalex_id":"https://openalex.org/W7118407843","authors":[],"funders":[{"funder_name":"NIH Office of the Director","grant_id":"3OT2OD030546-01S3","title":null},{"funder_name":"National Institute of Mental Health","grant_id":"R01MH129764","title":null},{"funder_name":"NIH HHS","grant_id":"OT2 OD030546","title":null}],"total_grants":3,"fwci":4.1252,"citation_percentile":0.86924283,"influential_citations":0,"citation_trend":[{"year":2026,"count":1}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://public-pages-files-2025.frontiersin.org/journals/bioinformatics/articles/10.3389/fbinf.2025.1727953/pdf","host_type":"journal"},{"url":"https://public-pages-files-2025.frontiersin.org/journals/bioinformatics/articles/10.3389/fbinf.2025.1727953/pdf","host_type":"publisher"},{"url":"https://www.frontiersin.org/articles/10.3389/fbinf.2025.1727953/full","host_type":"publisher"},{"url":"https://doi.org/10.3389/fbinf.2025.1727953","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/41584517","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12823822/","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC12823822","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC12823822?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Bioinformatics and Genomic Networks","Genetic Associations and Epidemiology","Advanced Graph Neural Networks"],"mesh_terms":[],"keywords":["Pipeline (software)","Scalability","Missing data","Subject (documents)","Inclusion (mineral)","Training set","Gene","Linkage analysis","Genome-wide Association Studies","Gwas","Disease-associated Genes","Network-based Methods","Biomedical Knowledge Graphs","Data Distillery Knowledge Graph","Ddkg"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-07-27T12:05:24.949149Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}