{"doi":"10.1016/j.aiopen.2025.11.003","title":"LLMKG＋: Systematically improving knowledge quality and coverage in KGs using LLMs – A case study in medical domain","abstract":"Knowledge graphs (KGs) encode structured information about real-world entities and their relations, supporting core NLP tasks such as question answering and retrieval. Existing LLM-based methods for knowledge extraction and fusion often struggle to balance quality and coverage when adapting to emerging knowledge. We propose LLMKG＋, a framework for KG expansion that integrates the generative strengths of LLMs with relevance verification. LLMKG＋features (1) a two-stage pipeline with retrieval-augmented generation followed by hierarchical expansion filtering, where the latter is the first to jointly assess semantic equivalence to eliminate triple-level redundancy while ensuring factual correctness, and (2) a novel KG Reconstruction Test that recognizes semantically equivalent triples to enable more accurate quality and coverage assessment. Evaluated on PubMed abstracts and the UMLS semantic network using eight state-of-the-art LLMs, LLMKG＋improves KG quality and coverage by 20.47%–73.71% over strong baselines. These results demonstrate that LLMKG＋offers an effective solution for KG expansion in domains requiring high quality, broad coverage, and continual knowledge growth. Code: https://github.com/xincanfeng/llmkg . • We propose LLMKG＋ , the first framework that systematically addresses triple-level semantic redundancy during knowledge graph (KG) expansion by combining retrieval-augmented generation with hierarchical verification. • LLMKG＋introduces a KG Reconstruction Test that accounts for semantically equivalent triples, enabling more accurate and comprehensive evaluation of knowledge quality and coverage than conventional exact-match metrics. • The framework integrates multi-stage correctness and relevance verification, combining BERT-based similarity filtering and LLM-based semantic reasoning to ensure both factual accuracy and non-redundant knowledge expansion. • Extensive experiments on PubMed abstracts and UMLS semantic network using eight state-of-the-art LLMs demonstrate that LLMKG＋improves KG quality and coverage by 20.47%–73.71% over strong baselines. • Human evaluation and ablation studies further confirm that LLMKG＋provides a robust, interpretable, and scalable solution for high-quality, high-coverage KG construction in domains that demand both precision and continual knowledge growth.","journal":"AI Open","year":2025,"id":582858,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9545,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":842654,"name":"Hejie Cui","orcid":"0000-0001-6388-2619","position":1,"is_corresponding":false},{"id":1494803,"name":"Kumiko Hayashi","orcid":"0009-0007-7327-5311","position":2,"is_corresponding":false},{"id":1494804,"name":"Huy Hien Vu","orcid":"0000-0002-5957-0345","position":3,"is_corresponding":false},{"id":1495193,"name":"Kenta T. Suzuki","orcid":null,"position":4,"is_corresponding":false},{"id":1494805,"name":"Noriki Nishida","orcid":"0000-0002-5851-7681","position":5,"is_corresponding":false},{"id":1494806,"name":"Hidetaka Kamigaito","orcid":"0000-0002-5249-5813","position":6,"is_corresponding":false},{"id":1230424,"name":"Yūji Matsumoto","orcid":"0000-0003-4946-9574","position":7,"is_corresponding":false},{"id":1494807,"name":"Taro Watanabe","orcid":"0000-0001-8349-3522","position":8,"is_corresponding":false},{"id":551800,"name":"Carl Yang","orcid":"0000-0001-9145-4531","position":9,"is_corresponding":false},{"id":1494802,"name":"Xincan Feng","orcid":"0000-0003-4647-7050","position":0,"is_corresponding":true}],"reference_count":3,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-07-19T02:58:59.653747Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}