{"doi":"10.5281/zenodo.17602575","title":"React-to-Me Expert Evaluation Dataset","abstract":"This dataset provides a complete expert evaluation framework for benchmarking grounded versus ungrounded large language models on domain-specific biomedical question answering. It accompanies \"React-to-Me: A Conversational AI System for Interactive Exploration of the Reactome Pathway Knowledgebase\" and contains all materials needed to replicate a blinded comparative assessment. The evaluation set comprises 109 molecular biology questions derived from Voet & Voet, Biochemistry (2nd ed.), stratified by cognitive complexity: 76 query-like questions requiring factual recall, and 33 reasoning questions requiring synthesis and multi-step inference. For each question, the dataset includes paired responses from React-to-Me (retrieval-augmented generation grounded in Reactome pathways) and GPT-4o-mini (ungrounded baseline), plus optional textbook hint excerpt citations. Responses were evaluated using a standardized four-point ordinal rubric across three independent dimensions: factual accuracy, level of granularity (biological specificity), and relational depth (mechanistic integration). Cumulative-link mixed-effects modeling revealed that grounded responses were twice as likely to receive higher expert ratings (OR=2.01, 95% CI: 1.50-2.69, p<0.01), with performance gains generalizing across both factual and reasoning tasks. This dataset enables replication of the reported evaluation and provides a reusable benchmark for assessing domain-grounded AI systems in molecular biology. The three-metric rubric, validated through blinded expert assessment, offers a standardized framework for evaluating factual reliability, biological precision, and systems-level reasoning in biomedical question-answering systems.","journal":"Zenodo (CERN European Organization for Nuclear Research)","year":2025,"id":588070,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":0.0,"corpus_rank":10062,"citation_count":0,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.9334,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":900591,"name":"Fatemeh Almodaresi","orcid":"0000-0003-0993-3600","position":1,"is_corresponding":false},{"id":1504405,"name":"Gregory F. J. Hogue","orcid":null,"position":2,"is_corresponding":false},{"id":49664,"name":"Adam Wright","orcid":"0000-0002-5719-4024","position":3,"is_corresponding":false},{"id":614230,"name":"M Orlic-Milacic","orcid":"0000-0002-3218-5631","position":4,"is_corresponding":false},{"id":1161509,"name":"Nancy T. Li","orcid":"0000-0002-2663-5245","position":5,"is_corresponding":false},{"id":1498488,"name":"Amin Mawani","orcid":"0000-0003-2054-7487","position":6,"is_corresponding":false},{"id":14310,"name":"Lincoln Stein","orcid":"0000-0002-1983-4588","position":7,"is_corresponding":false},{"id":1363649,"name":"Helia Mohammadi","orcid":"0009-0001-9112-5559","position":0,"is_corresponding":true}],"reference_count":0,"raw_metadata":null,"created_at":"2026-07-19T02:59:43.096742Z","pmid":null,"pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}