Evaluation of large language models in explaining black-box warnings: A comparative analysis of fidelity and clarity across models

Pharmacy Practice

  • Kannan Sridharan1Professor, Department of Pharmacology & Therapeutics, College of Medicine & Health Sciences, Arabian Gulf University, Manama, Kingdom of Bahrain.
  • Gowri Sivaramakrishnan2Scientific Researcher, Bahrain Defence Force Royal Medical Services, Riffa, Kingdom of Bahrain.

Volume 24 Issue 3 Pages 1-138

DOI: 10.18549/PharmPract.2026.3.3652

Abstract

Background: Black box warnings (BBWs) issued by the United States Food and Drug Administration (USFDA) serve as the most stringent alerts for serious or life-threatening drug-related adverse events. Despite their clinical importance, multiple studies indicate gaps in healthcare professionals’ awareness and patients’ understanding of BBWs. Large language models (LLMs) offer promising potential to simplify, clarify, and disseminate regulatory safety information. This study aimed to evaluate the performance of three advanced LLM, ChatGPT-4, DeepSeek-V3, and Google Gemini 2.5 Flash, in accurately and effectively conveying BBW content to healthcare personnel and patients. Methods: Thirteen FDA-issued BBWs were selected across a range of drugs involving pediatric, adult, psychiatric, reproductive health, and rare disease contexts. Each LLM was prompted separately to generate responses tailored to two audiences: healthcare professionals and patients. Responses were independently evaluated using structured rubrics across six domains, including fidelity to FDA language, completeness, clarity, actionability, understandability (for patient outputs), and absence of hallucinations. Output was scored and categorized as low-, moderate-, or high-quality based on the in-house developed rubric. Results: All three LLMs successfully generated responses for every scenario. Across both clinician- and patient-directed prompts, outputs from ChatGPT, DeepSeek, and Gemini were consistently rated as high quality. ChatGPT excelled in clarity and structured communication, DeepSeek provided detailed clinical guidance and regulatory references, and Gemini delivered comprehensive and empathetic narratives with enhanced personalization. Differences across models emerged in tone, output length, format, and inclusion of supplemental content. Importantly, no model exhibited critical hallucinations, although isolated overextensions of regulatory scope were noted in some DeepSeek responses. Conclusion: LLMs demonstrate strong potential in conveying complex BBW content to both clinical and lay audiences. While each model has distinct advantages, all three effectively translated regulatory language into clear and actionable guidance. Integration of LLMs into safety communication strategies may enhance risk comprehension and support safer prescribing practices.

Keywords

  • Serious adverse events
  • Pharmacovigilance
  • Black box warnings
  • BBW
  • AI
  • Artificial intelligence
Pharmacy Practice

Loading…