{"doi":"10.1093/jamia/ocae325","title":"Collaborative large language models for automated data extraction in living systematic reviews","abstract":"OBJECTIVE: Data extraction from the published literature is the most laborious step in conducting living systematic reviews (LSRs). We aim to build a generalizable, automated data extraction workflow leveraging large language models (LLMs) that mimics the real-world 2-reviewer process. MATERIALS AND METHODS: A dataset of 10 trials (22 publications) from a published LSR was used, focusing on 23 variables related to trial, population, and outcomes data. The dataset was split into prompt development (n = 5) and held-out test sets (n = 17). GPT-4-turbo and Claude-3-Opus were used for data extraction. Responses from the 2 LLMs were considered concordant if they were the same for a given variable. The discordant responses from each LLM were provided to the other LLM for cross-critique. Accuracy, ie, the total number of correct responses divided by the total number of responses, was computed to assess performance. RESULTS: In the prompt development set, 110 (96%) responses were concordant, achieving an accuracy of 0.99 against the gold standard. In the test set, 342 (87%) responses were concordant. The accuracy of the concordant responses was 0.94. The accuracy of the discordant responses was 0.41 for GPT-4-turbo and 0.50 for Claude-3-Opus. Of the 49 discordant responses, 25 (51%) became concordant after cross-critique, increasing accuracy to 0.76. DISCUSSION: Concordant responses by the LLMs are likely to be accurate. In instances of discordant responses, cross-critique can further increase the accuracy. CONCLUSION: Large language models, when simulated in a collaborative, 2-reviewer workflow, can extract data with reasonable performance, enabling truly \"living\" systematic reviews.","journal":"Journal of the American Medical Informatics Association","year":2025,"id":509346,"datarank":0.8921630326672245,"base_score":3.6109179126442243,"endowment":3.6109179126442243,"self_citation_contribution":0.5416376868966337,"citation_network_contribution":0.35052534577059075,"self_endowment_contribution":0.5416376868966337,"citer_contribution":0.35052534577059075,"corpus_percentile":77.40388334493696,"corpus_rank":2922,"citation_count":36,"citer_count":32,"citers_with_citation_signal":10,"citers_with_endowment":10,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.7101,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1326901,"name":"Umair Ayub","orcid":"0000-0001-8961-3978","position":1,"is_corresponding":false},{"id":246665,"name":"Syed Arsalan Ahmed Naqvi","orcid":"0000-0002-3131-2088","position":2,"is_corresponding":false},{"id":891181,"name":"Kaneez Zahra Rubab Khakwani","orcid":null,"position":3,"is_corresponding":false},{"id":1327372,"name":"Zaryab bin Riaz Sipra","orcid":null,"position":4,"is_corresponding":false},{"id":1327373,"name":"Ammad Raina","orcid":null,"position":5,"is_corresponding":false},{"id":1363286,"name":"Sihan Zhou","orcid":"0009-0004-4338-8575","position":6,"is_corresponding":false},{"id":251073,"name":"Huan He","orcid":"0000-0003-1312-4195","position":7,"is_corresponding":false},{"id":1364137,"name":"Amir Saeidi","orcid":null,"position":8,"is_corresponding":false},{"id":857725,"name":"Bashar Hasan","orcid":"0000-0001-9531-4990","position":9,"is_corresponding":false},{"id":591963,"name":"R. Bryan Rumble","orcid":"0000-0002-2786-2925","position":10,"is_corresponding":false},{"id":229099,"name":"Danielle S. Bitterman","orcid":"0000-0003-0345-2232","position":11,"is_corresponding":false},{"id":105190,"name":"Jeremy L. Warner","orcid":"0000-0002-2851-7242","position":12,"is_corresponding":false},{"id":1326902,"name":"Jia Zou","orcid":"0000-0001-9849-0711","position":13,"is_corresponding":false},{"id":280773,"name":"Amyé Tevaarwerk","orcid":"0000-0002-8087-5119","position":14,"is_corresponding":false},{"id":412008,"name":"Konstantinos Leventakos","orcid":"0000-0001-9784-8603","position":15,"is_corresponding":false},{"id":492213,"name":"Kenneth L. Kehl","orcid":"0000-0001-5339-9797","position":16,"is_corresponding":false},{"id":604682,"name":"Jeanne Palmer","orcid":"0000-0002-2185-1504","position":17,"is_corresponding":false},{"id":56883,"name":"Mohammad Hassan Murad","orcid":"0000-0001-5502-5975","position":18,"is_corresponding":false},{"id":1326903,"name":"Chitta Baral","orcid":"0000-0002-7549-723X","position":19,"is_corresponding":false},{"id":246664,"name":"Irbaz Bin Riaz","orcid":"0000-0003-4249-0311","position":20,"is_corresponding":false},{"id":1326900,"name":"Muhammad Ali Khan","orcid":"0000-0001-8235-1733","position":0,"is_corresponding":true}],"reference_count":25,"raw_metadata":null,"created_at":"2026-07-19T02:47:24.513904Z","pmid":"39836495","pmcid":"PMC12005628","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}