J Med Internet Res. 2026 Jul 23;28:e99136. doi: 10.2196/99136.
ABSTRACT
BACKGROUND: Molecular Tumor Boards (MTBs) generate highly technical recommendations. The language used in their protocols is rarely accessible to patients. Lay-language patient protocols could support patient-clinician communication, yet manual production is difficult to sustain in high-volume oncology settings. Large language models (LLMs) may offer scalable drafting assistance, yet clinical usability remains largely uninvestigated under real-world deployment constraints. Existing evaluations rely predominantly on synthetic data or closed-source models that are incompatible with strict data protection requirements.
OBJECTIVE: This study evaluated whether open-weight LLMs can provide clinically usable drafting support for German MTB patient protocols under real-world deployment constraints and developed a transferable evaluation framework for patient-facing text generation.
METHODS: Eight open-weight LLMs were evaluated under zero-shot (A1) and one-shot (A2) prompting with constrained decoding, which ensures section-schema compliance. Automatic evaluation used ROUGE-1 (Recall-Oriented Understudy for Gisting Evaluation), BERTScore-F1 (Bidirectional Encoder Representations From Transformers Score), Wiener Sachtextformel version 4, and DistilBERT (Distilled Version of Bidirectional Encoder Representations From Transformers)-based complexity using a corpus of 316 MTB protocols and 47 expert-written patient protocols. For expert evaluation, 7 medical oncologists evaluated 50 protocols from the best-performing model across 3 International Organization for Standardization 9241-11 usability dimensions using fine-grained error annotation, perceived postediting effort (PPEE), and net promoter score. Critical errors were defined as bearing the risk of patient harm.
RESULTS: Llama-3.3-70B-Instruct achieved the strongest automatic performance. Across models, A2 significantly improved most automatic metrics compared to A1. However, expert usability evaluation of Llama-3.3-70B-Instruct showed the opposite picture: the proportion of protocols containing at least 1 critical error doubled under A2 (10/25, 40% vs 5/25, 20%) compared with A1, and the dominant error type shifted from language (40/108, 37%) errors to factual errors (69/145, 48%). Overall, 16% (230/1420) of the annotated paragraphs contained errors. Median PPEE was 2 (IQR 2.0-3.0; low), and median net promoter score was 7 (IQR 5.0-9.0). Detractors (46/100, 46%) outweighed promoters (29/100, 29%), which suggests hesitation toward routine adoption. These differences in expert evaluation between A2 and A1 were directionally consistent but did not reach individual statistical significance for the paired samples (n=25).
CONCLUSIONS: Prompting strategies that improve automatic metrics can simultaneously increase the number of critical errors. Surface-level metric gains were, therefore, insufficient proxies for clinical safety. This was observed as a consistent directional pattern for a single model, but generalization to other models remains to be investigated. Nonetheless, the low paragraph-level error rate and favorable PPEE suggest that structured open-weight LLM generation may be a useful drafting support in a clinician-supervised setting. The proposed evaluation framework provides a text-quality-focused basis for future assessment of patient-facing LLM applications in real-world clinical settings.
PMID:42495812 | DOI:10.2196/99136