Close
MIDAS Network Member

Paper Information

Title

Evaluating large language models on multilingual vaccine knowledge: a benchmark study.

Abstract

Large language models (LLMs) are increasingly used by clinicians and the public for vaccine information, yet their factual accuracy across languages and vaccine domains remains insufficiently characterized. We evaluated 13 LLMs using VaxEval, a multilingual vaccine-knowledge benchmark of 1886 vaccine-related multiple-choice questions spanning 14 vaccines in English (71%), Spanish (13%), and Chinese (16%). All items underwent quality control, with reference answers verified against authoritative guidance and peer-reviewed sources. Model performance was evaluated under zero-shot, few-shot, and chain-of-thought (CoT) prompting, with exact-match accuracy defined as selecting the pre-specified reference option. We used mixed-effect logistic regression to estimate associations between model group (newer flagship models vs earlier models), prompting strategy, language, and vaccine type, and answer correctness. Mean accuracy across models was 86.0% in English, 83.7% in Spanish, and 80.0% in Chinese. Flagship models had higher odds of correctness than earlier versions (OR 1.57; 95% CI 1.50-1.65; P < .001). Few-shot prompting was associated with higher correctness (OR 1.17; P < .001), whereas CoT prompting was associated with lower correctness (OR 0.79; P < .001). Performance varied by vaccine type and question category, underscoring the need for rigorous evaluation, structured guardrails, and targeted refinement before using LLMs for vaccine communication.

Journal

NPJ vaccines

Citation

MIDAS Authors