{"doi":"10.1002/cesm.70033","title":"Data Extractions Using a Large Language Model (Elicit) and Human Reviewers in Randomized Controlled Trials: A Systematic Comparison","abstract":"<jats:title>ABSTRACT</jats:title>\n                  <jats:sec>\n                    <jats:title>Aim</jats:title>\n                    <jats:p>We aimed at comparing data extractions from randomized controlled trials by using Elicit and human reviewers.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Background</jats:title>\n                    <jats:p>Elicit is an artificial intelligence tool which may automate specific steps in conducting systematic reviews. However, the tool's performance and accuracy have not been independently assessed.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods</jats:title>\n                    <jats:p>For comparison, we sampled 20 randomized controlled trials of which data were extracted manually from a human reviewer. We assessed the variables study objectives, sample characteristics and size, study design, interventions, outcome measured, and intervention effects and classified the results into “more,” “equal to,” “partially equal,” and “deviating” extractions. STROBE checklist was used to report the study.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Results</jats:title>\n                    <jats:p>We analysed 20 randomized controlled trials from 11 countries. The studies covered diverse healthcare topics. Across all seven variables, Elicit extracted “more” data in 29.3% of cases, “equal” in 20.7%, “partially equal” in 45.7%, and “deviating” in 4.3%. Elicit provided “more” information for the variable study design (100%) and sample characteristics (45%). In contrast, for more nuanced variables, such as “intervention effects,” Elicit's extractions were less detailed, with 95% rated as “partially equal.”</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions</jats:title>\n                    <jats:p>Elicit was capable of extracting data partly correct for our predefined variables. Variables like “intervention effect” or “intervention” may require a human reviewer to complete the data extraction. Our results suggest that verification by human reviewers is necessary to ensure that all relevant information is captured completely and correctly by Elicit.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Implications</jats:title>\n                    <jats:p>Systematic reviews are labor‐intensive. Data extraction process may be facilitated by artificial intelligence tools. Use of Elicit may require a human reviewer to double‐check the extracted data.</jats:p>\n                  </jats:sec>","journal":"Cochrane Evidence Synthesis and Methods","year":2025,"id":627380,"datarank":0.41588830833596724,"base_score":2.772588722239781,"endowment":2.772588722239781,"self_citation_contribution":0.41588830833596724,"citation_network_contribution":0.0,"self_endowment_contribution":0.41588830833596724,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":15,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":null,"is_data_producer":false,"deposit_databanks":null,"is_oa":false,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":null,"fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":27467,"name":"Julian Hirt","orcid":"0000-0001-6589-3936","position":1,"is_corresponding":false},{"id":1623771,"name":"Magdalena Vogt","orcid":"0009-0002-1281-8304","position":2,"is_corresponding":false},{"id":851205,"name":"Janine Vetsch","orcid":"0000-0002-5564-7093","position":3,"is_corresponding":false},{"id":1623770,"name":"Joleen Bianchi","orcid":null,"position":0,"is_corresponding":false}],"reference_count":0,"raw_metadata":{"has_enrichment":true,"resolved":true,"title":"Data Extractions Using a Large Language Model (Elicit) and Human Reviewers in Randomized Controlled Trials: A Systematic Comparison","abstract":"<jats:title>ABSTRACT</jats:title>\n                  <jats:sec>\n                    <jats:title>Aim</jats:title>\n                    <jats:p>We aimed at comparing data extractions from randomized controlled trials by using Elicit and human reviewers.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Background</jats:title>\n                    <jats:p>Elicit is an artificial intelligence tool which may automate specific steps in conducting systematic reviews. However, the tool's performance and accuracy have not been independently assessed.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods</jats:title>\n                    <jats:p>For comparison, we sampled 20 randomized controlled trials of which data were extracted manually from a human reviewer. We assessed the variables study objectives, sample characteristics and size, study design, interventions, outcome measured, and intervention effects and classified the results into “more,” “equal to,” “partially equal,” and “deviating” extractions. STROBE checklist was used to report the study.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Results</jats:title>\n                    <jats:p>We analysed 20 randomized controlled trials from 11 countries. The studies covered diverse healthcare topics. Across all seven variables, Elicit extracted “more” data in 29.3% of cases, “equal” in 20.7%, “partially equal” in 45.7%, and “deviating” in 4.3%. Elicit provided “more” information for the variable study design (100%) and sample characteristics (45%). In contrast, for more nuanced variables, such as “intervention effects,” Elicit's extractions were less detailed, with 95% rated as “partially equal.”</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions</jats:title>\n                    <jats:p>Elicit was capable of extracting data partly correct for our predefined variables. Variables like “intervention effect” or “intervention” may require a human reviewer to complete the data extraction. Our results suggest that verification by human reviewers is necessary to ensure that all relevant information is captured completely and correctly by Elicit.</jats:p>\n                  </jats:sec>\n                  <jats:sec>\n                    <jats:title>Implications</jats:title>\n                    <jats:p>Systematic reviews are labor‐intensive. Data extraction process may be facilitated by artificial intelligence tools. Use of Elicit may require a human reviewer to double‐check the extracted data.</jats:p>\n                  </jats:sec>","is_dataset_classified":null,"base_score":2.772588722239781,"endowment":2.772588722239781,"datacite_reuse_total":0,"file_count":0,"downloads":0,"views":0,"has_version_chain":false,"is_dataset":false,"is_oa":false,"pmid":"41019842","pmcid":"PMC12462964","openalex_id":"https://openalex.org/W4411160616","authors":[],"funders":[],"total_grants":0,"fwci":10.2607,"citation_percentile":0.98684729,"influential_citations":0,"citation_trend":[{"year":2025,"count":2},{"year":2026,"count":13}],"oa_status":"gold","license":"cc-by","oa_locations":[{"url":"https://onlinelibrary.wiley.com/doi/pdfdirect/10.1002/cesm.70033","host_type":"journal"},{"url":"https://onlinelibrary.wiley.com/doi/pdfdirect/10.1002/cesm.70033","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/pdf/10.1002/cesm.70033","host_type":"publisher"},{"url":"https://onlinelibrary.wiley.com/doi/full-xml/10.1002/cesm.70033","host_type":"publisher"},{"url":"https://doi.org/10.1002/cesm.70033","host_type":"journal"},{"url":"https://pubmed.ncbi.nlm.nih.gov/41019842","host_type":"repository"},{"url":"https://doaj.org/article/d3ebb3324435458fafa85feff3baabf5","host_type":"repository"},{"url":"https://www.ncbi.nlm.nih.gov/pmc/articles/12462964","host_type":"repository"},{"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC12462964/","host_type":"repository"},{"url":"https://europepmc.org/articles/PMC12462964","host_type":"Europe_PMC"},{"url":"https://europepmc.org/articles/PMC12462964?pdf=render","host_type":"Europe_PMC"}],"fields_of_study":["Meta-analysis and systematic reviews","Artificial Intelligence in Healthcare and Education","Academic Writing and Publishing"],"mesh_terms":[],"keywords":["Randomized controlled trial","Computer science","Natural language processing","Medical physics","Information retrieval","Medicine","Internal medicine","Artificial intelligence","Systematic review","Data Extraction","Human Reviewer"],"sdg_mappings":[],"linked_datasets":[],"clinical_trials":[],"software_tools":[],"database_accessions":[],"source":"live","citation_network_status":"fetched"},"created_at":"2026-08-04T17:18:27.247463Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}