ChatGPT performance on pharmacology examination and board review questions: Implications for medical education and knowledge assessment

Pharmacy Practice

  • Rima A. Hijazeen1B.Sc. Pharmacy, M.Sc. Clinical Pharmacy, PhD Clinical Pharmacy Practice; Associate Professor in Clinical Pharmacy Practice; The University of Jordan, Faculty of Pharmacy, Department of Biopharmaceutics and Clinical Pharmacy, Amman 11942, Jordan.
  • Al-Motassem Yousef2B.Sc. Pharmacy, PhD in Pharmacology and Therapeutics; Professor of Pharmacology and Therapeutics; The University of Jordan, Faculty of Pharmacy, Department of Biopharmaceutics and Clinical Pharmacy, Amman 11942, Jordan.
  • Ahmed Almousa3RPh, MSc, PhD; Assistant Professor of Clinical Pharmacy; Department of Biopharmaceutics and Clinical Pharmacy, Faculty of Pharmacy, University of Jordan, Amman, Jordan.
  • Aya N. Alzoghair4Undergraduate Pharmacy student; The University of Jordan, Faculty of Pharmacy, Amman 11942, Jordan.
  • Jude K. Dwairi4Undergraduate Pharmacy student; The University of Jordan, Faculty of Pharmacy, Amman 11942, Jordan.
  • Majd I. Sawaqed4Undergraduate Pharmacy student; The University of Jordan, Faculty of Pharmacy, Amman 11942, Jordan.
  • Ghaith F. Al-Ryahneh4Undergraduate Pharmacy student; The University of Jordan, Faculty of Pharmacy, Amman 11942, Jordan.
  • Marwan H Ali4Undergraduate Pharmacy student; The University of Jordan, Faculty of Pharmacy, Amman 11942, Jordan.

Volume 24 Issue 2 Pages 1-10

DOI: 10.18549/PharmPract.2026.2.3488

Abstract

Objectives: This study aimed to evaluate ChatGPT’s performance on pharmacology exam questions by assessing its accuracy in basic and clinical pharmacology, reasoning processes, and response consistency over time. Methods: A dataset of 583 multiple-choice questions from the Pharmacology Examination and Board Review (13th edition) was used. ChatGPT’s responses were evaluated for logical justification, use of internal question stem information, and integration of external knowledge. Statistical analyses, including chi-square and McNemar tests, assessed associations and changes in response accuracy over a four-week interval. Results: ChatGPT achieved 76.2% accuracy (444/583 questions), demonstrating logical reasoning in 97% of responses. Internal information was used in 99.7% of cases, while external information was incorporated in 98% of correct and 93.5% of incorrect responses (p = 0.008). Information errors were the most common reason for incorrect answers. A statistically significant improvement in accuracy upon re-evaluation (χ² = 37.3, p < 0.0001) was observed, suggesting potential temporal variation in performance. Conclusion: ChatGPT meets or exceeds typical passing standards in many educational settings, with evidence of improved response accuracy over time. These findings highlight its capabilities in processing pharmacological content, with potential implications for future research into AI-assisted educational tools.

Keywords

  • Artificial intelligence
  • ChatGPT
  • Medical education
  • Pharmacology
  • Reasoning
  • Multiple-choice questions
  • Large language model
Pharmacy Practice

Loading…