{"doi":"10.1016/j.xops.2024.100667","title":"Glaucoma Detection and Feature Identification via GPT-4V Fundus Image Analysis","abstract":"Purpose: The aim is to assess GPT-4V's (OpenAI) diagnostic accuracy and its capability to identify glaucoma-related features compared to expert evaluations. Design: Evaluation of multimodal large language models for reviewing fundus images in glaucoma. Subjects: A total of 300 fundus images from 3 public datasets (ACRIMA, ORIGA, and RIM-One v3) that included 139 glaucomatous and 161 nonglaucomatous cases were analyzed. Methods: Preprocessing ensured each image was centered on the optic disc. GPT-4's vision-preview model (GPT-4V) assessed each image for various glaucoma-related criteria: image quality, image gradability, cup-to-disc ratio, peripapillary atrophy, disc hemorrhages, rim thinning (by quadrant and clock hour), glaucoma status, and estimated probability of glaucoma. Each image was analyzed twice by GPT-4V to evaluate consistency in its predictions. Two expert graders independently evaluated the same images using identical criteria. Comparisons between GPT-4V's assessments, expert evaluations, and dataset labels were made to determine accuracy, sensitivity, specificity, and Cohen kappa. Main Outcome Measures: The main parameters measured were the accuracy, sensitivity, specificity, and Cohen kappa of GPT-4V in detecting glaucoma compared with expert evaluations. Results: GPT-4V successfully provided glaucoma assessments for all 300 fundus images across the datasets, although approximately 35% required multiple prompt submissions. GPT-4V's overall accuracy in glaucoma detection was slightly lower (0.68, 0.70, and 0.81, respectively) than that of expert graders (0.78, 0.80, and 0.88, for expert grader 1 and 0.72, 0.78, and 0.87, for expert grader 2, respectively), across the ACRIMA, ORIGA, and RIM-ONE datasets. In Glaucoma detection, GPT-4V showed variable agreement by dataset and expert graders, with Cohen kappa values ranging from 0.08 to 0.72. In terms of feature detection, GPT-4V demonstrated high consistency (repeatability) in image gradability, with an agreement accuracy of ≥89% and substantial agreement in rim thinning and cup-to-disc ratio assessments, although kappas were generally lower than expert-to-expert agreement. Conclusions: GPT-4V shows promise as a tool in glaucoma screening and detection through fundus image analysis, demonstrating generally high agreement with expert evaluations of key diagnostic features, although agreement did vary substantially across datasets. Financial Disclosures: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.","journal":"Ophthalmology Science","year":2024,"id":427168,"datarank":0.6373109268138797,"base_score":2.833213344056216,"endowment":2.833213344056216,"self_citation_contribution":0.42498200160843247,"citation_network_contribution":0.21232892520544716,"self_endowment_contribution":0.42498200160843247,"citer_contribution":0.21232892520544716,"corpus_percentile":68.90229751682524,"corpus_rank":4021,"citation_count":16,"citer_count":15,"citers_with_citation_signal":7,"citers_with_endowment":7,"datacite_reuse_total":0,"is_dataset":true,"is_dataset_confidence":0.7812,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2024-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":1113328,"name":"Anuwat Jiravarnsirikul","orcid":"0000-0001-6054-8850","position":1,"is_corresponding":false},{"id":277210,"name":"Christopher Bowd","orcid":"0000-0001-9330-9987","position":2,"is_corresponding":false},{"id":817266,"name":"Benton Chuter","orcid":"0000-0002-0354-0569","position":3,"is_corresponding":false},{"id":457861,"name":"Akram Belghith","orcid":null,"position":4,"is_corresponding":false},{"id":281319,"name":"Michael H. Goldbaum","orcid":"0000-0002-7721-2736","position":5,"is_corresponding":false},{"id":349921,"name":"Sally L. Baxter","orcid":"0000-0002-5271-7690","position":6,"is_corresponding":false},{"id":6890,"name":"Robert N. Weinreb","orcid":"0000-0001-9553-3202","position":7,"is_corresponding":false},{"id":326179,"name":"Linda M. Zangwill","orcid":"0000-0002-1143-5224","position":8,"is_corresponding":false},{"id":456446,"name":"Mark Christopher","orcid":"0000-0001-6165-6175","position":9,"is_corresponding":false},{"id":1227267,"name":"Jalil Jalili","orcid":"0000-0003-4920-0815","position":0,"is_corresponding":true}],"reference_count":39,"raw_metadata":null,"created_at":"2026-07-19T01:58:47.805440Z","pmid":"39877464","pmcid":"PMC11773068","fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}