{"doi":"10.51408/issi2025_078","title":"Leveraging Large Language Models for Post-Publication Peer Review: Potential and Limitations","abstract":"<jats:p>Peer review is the cornerstone of scientific evaluation, ensuring the quality, accuracy, and integrity\nof published research. However, challenges such as reviewer bias, time constraints, and the\nincreasing volume of submissions have strained traditional peer review systems, resulting in delays,\nlower-quality reviews, and reviewer fatigue. These limitations highlight the need for innovative\nsolutions. Large language models (LLMs) have emerged as promising tools to support or potentially\nreplace certain aspects of peer review. This study investigates the potential of LLMs to enhance\npost-publication peer review, offering quality assessments and recommendations for published\narticles. Specifically, we designed two tasks to evaluate the performance of LLMs in postpublication research evaluation: identifying high-quality articles (Task 1) and providing ratings on\nrecommended articles (Task 2). Six versions of three generative LLMs, including open-source\nmodels such as Qwen and Llama, the closed-source GPT-4o-mini model, and four BERT-based\nmodels, were assessed using in-context learning and fine-tuning approaches. The data for training\nand evaluation were sourced from H1 Connect (formerly Faculty Opinions), a platform for expert\nrecommendations in the biomedical domain. Results indicate that fine-tuning LLMs with labelled\ndata can significantly enhance their alignment with human expert evaluations. For Task 1, fine-tuned\nmodels performed well in identifying high-quality articles with an accuracy of 84%. However, for\nTask 2 - rating on recommended articles - LLMs struggled to match human judgement consistently\nwith an accuracy below 0.6, highlighting their current limitations in nuanced, context-dependent\ntasks.</jats:p>","journal":"International Conference on Scientometrics &amp; Informetrics","year":2025,"id":4849,"datarank":0.10397207708399181,"base_score":0.6931471805599453,"endowment":0.6931471805599453,"self_citation_contribution":0.10397207708399181,"citation_network_contribution":0.0,"self_endowment_contribution":0.10397207708399181,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":1,"citer_count":1,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.0463,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-07-10","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":50497,"name":"Mengjia Wu","orcid":"0000-0003-3956-7808","position":1,"is_corresponding":false},{"id":1647,"name":"Yi Zhang","orcid":"0000-0002-2789-0811","position":2,"is_corresponding":false},{"id":839,"name":"Robin Haunschild","orcid":"0000-0001-7025-7256","position":4,"is_corresponding":false},{"id":50498,"name":"Max Planck Institute for Solid State Research","orcid":null,"position":5,"is_corresponding":false},{"id":661,"name":"Lutz Bornmann","orcid":"0000-0003-0810-7091","position":6,"is_corresponding":false},{"id":50499,"name":"Max Planck Institute for Solid State Research; Science Policy and Strategy Department, Administrative Headquarters of the Max Planck Society","orcid":null,"position":7,"is_corresponding":false},{"id":50496,"name":"Australian Artificial Intelligence Institute, Faculty of Engineering and Information Technology, University of Technology Sydney","orcid":null,"position":0,"is_corresponding":true}],"reference_count":0,"raw_metadata":{"citation_network_status":"fetched"},"created_at":"2026-03-01T18:20:47.508186Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}