{"doi":"10.1038/s41598-024-55568-7","title":"The model student: GPT-4 performance on graduate biomedical science exams","abstract":"The GPT-4 large language model (LLM) and ChatGPT chatbot have emerged as accessible and capable tools for generating English-language text in a variety of formats. GPT-4 has previously performed well when applied to questions from multiple standardized examinations. However, further evaluation of trustworthiness and accuracy of GPT-4 responses across various knowledge domains is essential before its use as a reference resource. Here, we assess GPT-4 performance on nine graduate-level examinations in the biomedical sciences (seven blinded), finding that GPT-4 scores exceed the student average in seven of nine cases and exceed all student scores for four exams. GPT-4 performed very well on fill-in-the-blank, short-answer, and essay questions, and correctly answered several questions on figures sourced from published manuscripts. Conversely, GPT-4 performed poorly on questions with figures containing simulated data and those requiring a hand-drawn answer. Two GPT-4 answer-sets were flagged as plagiarism based on answer similarity and some model responses included detailed hallucinations. In addition to assessing GPT-4 performance, we discuss patterns and limitations in GPT-4 capabilities with the goal of informing design of future academic examinations in the chatbot era.","journal":"Scientific Reports","year":2024,"id":418323,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":52,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9246,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2024-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":361150,"name":"Yuxing Xia","orcid":"0000-0002-8415-8256","position":1,"is_corresponding":false},{"id":1206684,"name":"Maha K Amer","orcid":null,"position":2,"is_corresponding":false},{"id":5658,"name":"Kiley Graim","orcid":"0000-0002-4569-8444","position":3,"is_corresponding":false},{"id":1058515,"name":"Connie J. Mulligan","orcid":"0000-0002-4360-2402","position":4,"is_corresponding":false},{"id":527364,"name":"Rolf Renne","orcid":"0000-0001-7391-8806","position":5,"is_corresponding":false},{"id":720555,"name":"Daniel Stribling","orcid":"0000-0002-0649-9506","position":0,"is_corresponding":true}],"reference_count":45,"raw_metadata":null,"created_at":"2026-07-19T01:57:02.483590Z","pmid":"38453979","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}