JMIR Form Res. 2026 Aug 26;10:e95883. doi: 10.2196/95883.
ABSTRACT
BACKGROUND: Large language models (LLMs) are increasingly used to support digital health communication, yet their reliability in patient-facing cardiovascular imaging education remains uncertain. Cardiovascular imaging involves complex terminology and procedural details that many patients struggle to understand, creating a need for accurate, clear, and reassuring explanations. While prior evaluations of conversational AI have focused primarily on diagnostic reasoning or clinician-oriented tasks, few studies have systematically compared contemporary LLMs in their ability to communicate effectively with patients.
OBJECTIVE: This study aimed to compare the accuracy, clarity, completeness, and patient-centered communication quality of responses generated by 3 state-of-the-art conversational agents (DeepSeek, GPT-o1, and GPT-4o) when addressing real-world patient questions about cardiovascular imaging.
METHODS: A prospective methodological evaluation was conducted using 84 unique patient-centered questions curated from authoritative cardiovascular information sources and online patient forums. Each question was independently submitted to DeepSeek, GPT-o1, and GPT-4o in isolated sessions to avoid contextual contamination. Two cardiovascular radiologists scored each response across 4 domains (accuracy, clarity and appropriateness, completeness, and user engagement and reassurance) using a standardized 3-point rubric (total score range 4-12). Discrepancies were resolved through predefined adjudication procedures. Because the scores were ordinal, median domain and composite scores with IQRs were summarized and compared across the 3 models using the Kruskal-Wallis test, with ε2 as an effect size. Statistical significance was defined as an α value of .05.
RESULTS: Across the 84 patient questions, all 3 models produced largely accurate, clear, and complete responses, with comparably high scores across the accuracy, clarity, and completeness domains (median 3 of 3, IQR 3-3 in each). The only meaningful difference appeared in user engagement and reassurance. A “good” engagement rating was assigned to 96.4% (81/84) of DeepSeek responses and 98.8% (83/84) of GPT-o1 responses but only 53.6% (45/84) of GPT-4o responses (Kruskal-Wallis P<.001). Composite scores were correspondingly lower for GPT-4o (median 11, IQR 10-12) than for DeepSeek and GPT-o1 (both median 12, IQR 11-12; P<.001). No significant differences were observed across models for accuracy (P=.91), clarity (P=.06), or completeness (P=.65), and no unsafe statements were identified in any model.
CONCLUSIONS: DeepSeek and GPT-o1 consistently delivered accurate, clear, and patient-centered explanations of cardiovascular imaging questions, whereas GPT-4o, despite comparable technical accuracy, provided less engaging and reassuring communication. These findings suggest that affective qualities rather than factual correctness represent the main differentiator among current LLMs in patient education tasks. As conversational agents become integrated into cardiovascular imaging workflows, attention to communication tone, emotional support, and health literacy alignment will be essential to ensure safe and effective patient use.
PMID:42647860 | DOI:10.2196/95883