Categories
Nevin Manimala Statistics

Performance of Large Language Models for Oncology Nursing Decision Support: Cross-Sectional Study

J Med Internet Res. 2026 Jul 24;28:e97802. doi: 10.2196/97802.

ABSTRACT

BACKGROUND: Large language models (LLMs) are increasingly used in health care, with emerging applications in clinical decision support and nursing education. However, evidence on their performance in nursing contexts, particularly in oncology nursing, remains limited. Given the complexity and high-risk nature of oncology care, it is important to evaluate the performance and clinical relevance of LLM-generated responses in oncology nursing contexts.

OBJECTIVE: This study aimed to compare the performance of LLMs in oncology nursing decision support tasks using standardized examination questions and case-based clinical scenarios and explore LLMs’ potential applicability and current limitations in oncology nursing practice.

METHODS: A total of 33 case-based questions derived from 10 oncology nursing clinical scenarios in a nationally used training manual, along with standardized examination-oriented questions from a commercially published preparation book for the Chinese Nursing (Intermediate) Qualification Examination, were used to evaluate the performance of 5 LLMs (DeepSeek, Qwen, Spark-Desk, WiseDiag, and ChatGPT). All models generated responses using a standardized prompt. Two oncology nurses with more than 5 years of clinical experience independently rated the case-based responses using 3 evaluation dimensions: correctness, clarity, and conciseness. Interrater reliability was assessed using the quadratic weighted Cohen κ, intraclass correlation coefficient, and Spearman rank correlation coefficient. Differences among models were analyzed using the Kruskal-Wallis test with the Dunn post hoc test. In addition, examination performance was evaluated based on total score, accuracy rate, and completion efficiency.

RESULTS: Interrater reliability analyses indicated moderate agreement between evaluators. The median correctness, clarity, and conciseness scores were as follows: 11.50 (IQR 10.50-12.00) for DeepSeek, 11.00 (IQR 10.50-12.00) for Qwen, 10.50 (IQR 9.50-11.50) for Spark-Desk, 10.00 (IQR 9.50-11.50) for WiseDiag, and 10.00 (IQR 9.00-11.50) for ChatGPT. The Kruskal-Wallis test indicated statistically significant differences among models (H=11.416; P<.05), with post hoc analysis showing a significant difference only between DeepSeek and ChatGPT (P<.05). In examination-based tasks, all models achieved passing performance, with accuracy rates ranging from 77% (77/100) to 93% (93/100). In terms of response completion, DeepSeek and ChatGPT completed all tasks in a single interaction, whereas other models required multiple interactions due to output interruptions.

CONCLUSIONS: LLMs showed relatively strong performance on structured knowledge and examination-based tasks but remained limited in complex oncology nursing scenarios requiring individualized assessment and dynamic clinical judgment. Their potential use may be most relevant to information retrieval, knowledge organization, and patient education. Because the correctness, clarity, and conciseness rubric showed only moderate interrater reliability, the case-based comparisons should be interpreted as preliminary signals rather than definitive evidence of between-model differences. LLM outputs should therefore be used as supportive information and interpreted alongside professional clinical judgment.

PMID:42497362 | DOI:10.2196/97802

By Nevin Manimala

Portfolio Website for Nevin Manimala