{"doi":"10.1016/j.xops.2025.101034","title":"Performance of GPT-5 Frontier Models in Ophthalmology Question Answering","abstract":"Purpose: Novel large language models (LLMs) such as Generative Pretrained Transformer-5 (GPT-5) integrate advanced reasoning capabilities that may enhance performance on complex medical question-answering tasks. For this latest generation of reasoning models, the configurations that maximize both accuracy and cost-efficiency have yet to be established. Our objective was to evaluate the performance and cost-accuracy trade-offs of OpenAI's GPT-5 compared with previous generation LLMs on ophthalmic question answering. Design: Evaluation of diagnostic test or technology. Participants: Generative Pretrained Transformer-5 is a publicly available LLM. Methods: In August 2025, 12 configurations of OpenAI's GPT-5 series (3 model tiers across 4 reasoning effort settings) were evaluated alongside o1-high, o3-high, and GPT-4o, using 260 closed-access multiple-choice questions from the American Academy of Ophthalmology Basic Clinical Science Course data set. The study did not include human participants. Main Outcome Measures: The primary outcome was accuracy on the 260-item ophthalmology multiple-choice question set for each model configuration. The secondary outcomes included head-to-head ranking of configurations using a Bradley-Terry model applied to paired win/loss comparisons of answer accuracy, and evaluation of generated natural language rationales using a reference-anchored, pairwise LLM-as-a-judge framework. Additional analyses assessed the accuracy-cost trade-off by calculating mean per-question cost from token usage and identifying Pareto-efficient configurations. Results: < 0.001), but not o3-high (0.958; 95% CI, 0.931-0.981). The configuration GPT-5-high ranked first in accuracy (1.66x stronger than o3-high) and rationale quality (1.11x stronger than o3-high), as judged by a reference-anchored LLM-as-a-judge autograder. Cost-accuracy analysis identified multiple GPT-5 configurations on the Pareto frontier, with GPT-5-mini-low providing the most optimal low-cost, high-performance configuration. Conclusions: This study benchmarks the GPT-5 series on a high-quality ophthalmology question-answering data set, demonstrating that GPT-5 with high reasoning effort achieved near-perfect accuracy and outperformed prior reasoning LLMs. This study also introduces an autograder framework for scalable, automated evaluation of LLM-generated answers against reference standards in ophthalmology. Financial Disclosures: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.","journal":"Ophthalmology Science","year":2025,"id":527750,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":3,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9514,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2025-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1405123,"name":"David Mikhail","orcid":"0009-0009-0831-1915","position":1,"is_corresponding":false},{"id":1405124,"name":"Daniel Milad","orcid":"0000-0002-0693-3421","position":2,"is_corresponding":false},{"id":1026445,"name":"Danny A. Mammo","orcid":"0000-0002-7496-5118","position":3,"is_corresponding":false},{"id":323239,"name":"Sumit Sharma","orcid":"0000-0001-5769-0717","position":4,"is_corresponding":false},{"id":341327,"name":"Sunil K. Srivastava","orcid":"0000-0002-0398-8806","position":5,"is_corresponding":false},{"id":1405125,"name":"Bing Yu Chen","orcid":"0000-0003-4049-7528","position":6,"is_corresponding":false},{"id":1405126,"name":"Samir Touma","orcid":"0000-0002-6365-0946","position":7,"is_corresponding":false},{"id":1405638,"name":"Mertcan Sevgi","orcid":null,"position":8,"is_corresponding":false},{"id":1405127,"name":"Jonathan El‐Khoury","orcid":"0000-0003-3186-2351","position":9,"is_corresponding":false},{"id":107563,"name":"Pearse A. Keane","orcid":"0000-0002-9239-745X","position":10,"is_corresponding":false},{"id":236881,"name":"Qingyu Chen","orcid":"0000-0002-6036-1516","position":11,"is_corresponding":false},{"id":291483,"name":"Yih Chung Tham","orcid":"0000-0002-6752-797X","position":12,"is_corresponding":false},{"id":1405128,"name":"Renaud Duval","orcid":"0000-0002-3845-3318","position":13,"is_corresponding":false},{"id":1209631,"name":"Fares Antaki","orcid":"0000-0001-6679-7276","position":0,"is_corresponding":true}],"reference_count":10,"raw_metadata":null,"created_at":"2026-07-19T02:50:44.062153Z","pmid":"41552655","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}