Categories
Nevin Manimala Statistics

Safety, accuracy, empathy, reliability, and readability of large language model chatbot responses to public-facing vegetarian and vegan nutrition advice questions: a cross-sectional comparative study

Front Public Health. 2026 Aug 3;14:1921819. doi: 10.3389/fpubh.2026.1921819. eCollection 2026.

ABSTRACT

BACKGROUND: Publicly accessible large language model (LLM) chatbots are increasingly used to seek nutrition advice. Although vegetarian and vegan nutrition advice is often perceived as low risk, it may involve supplementation, vulnerable life stages, chronic disease, and symptoms that require clinical assessment.

OBJECTIVE: To compare the safety, accuracy, empathy, reliability, information quality, and readability of five publicly accessible LLM chatbot products in responding to public-facing vegetarian and vegan nutrition advice questions.

METHODS: In this cross-sectional comparative study, 58 predefined public-facing vegetarian and vegan nutrition advice questions were submitted once in English to ChatGPT, Gemini, Microsoft Copilot, DeepSeek, and Doubao through their official web interfaces. The final dataset included 290 chatbot responses. Five blinded raters with clinical nutrition training evaluated safety, accuracy, and empathy using predefined reference-answer anchors, and assessed reliability and information quality using four established instruments: the DISCERN instrument, Ensuring Quality Information for Patients (EQIP), Journal of the American Medical Association (JAMA) benchmark criteria, and Global Quality Score (GQS). Readability was assessed using six formula-based indices: the Automated Readability Index (ARI), Coleman-Liau Index (CL), Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), Simple Measure of Gobbledygook (SMOG), and Flesch Reading Ease Score (FRES). Model comparisons were paired by question.

RESULTS: Inter-rater agreement was high for all human-rated outcomes. Potentially harmful responses occurred in all models, although adjusted pairwise safety comparisons were not statistically significant. Recurring harm patterns involved vitamin B12 source reliability, uncontrolled iodine or selenium intake, vulnerable life stages, chronic disease, symptom triage, and poorly traceable or overconfident statements. Accuracy, empathy, reliability, information quality, and readability differed significantly across models. ChatGPT, Copilot, and DeepSeek showed higher accuracy; ChatGPT showed higher empathy; Copilot performed best on DISCERN and EQIP; and Gemini produced the most readable responses.

CONCLUSION: Publicly accessible LLM chatbots can provide useful general information on vegetarian and vegan nutrition, but their performance is uneven and safety limitations remain. Chatbot-generated advice should be interpreted cautiously, especially for supplementation, vulnerable groups, chronic disease, and symptoms requiring clinical assessment. Future systems should improve safety guidance, referral cues, source traceability, and plain-language communication.

PMID:42609511 | PMC:PMC13478225 | DOI:10.3389/fpubh.2026.1921819

By Nevin Manimala

Portfolio Website for Nevin Manimala