JMIR AI. 2026 Aug 4;5:e78485. doi: 10.2196/78485.
ABSTRACT
BACKGROUND: Patient-reported outcome measures (PROMs) are central to multinational clinical research, but high-quality translation and linguistic validation remain resource-intensive. AI-powered translation may accelerate this process, but its performance relative to validated human PROM translations requires systematic evaluation.
OBJECTIVE: This benchmarking study evaluated the quality and comparability of 4 AI-powered translation services for the EuroQol 5-dimension 5-level (EQ-5D-5L) across 5 target languages, using official, linguistically validated human translations as the reference standard (gold standard).
METHODS: The 43 text segments of the EQ-5D-5L were translated from English into Danish, Dutch, French, German, and Spanish using Google Translate, GPT-4.1, Amazon Translate, and DeepL. GPT-4.1 was evaluated with a structured medical-translator prompt, whereas Google Translate, Amazon Translate, and DeepL were evaluated using standard unprompted application programming interfaces without domain-specific glossary constraints. Outputs were benchmarked against official, validated human translations using 4 automated metrics: BLEU (bilingual evaluation understudy), METEOR (metric for evaluation of translation with explicit ordering), COMET (cross-lingual optimized metric for evaluation of translation), and BLEURT (bilingual evaluation understudy with representations from transformers). Friedman tests were used to assess overall between-service differences within each metric-language combination. When the Friedman test was significant, paired Wilcoxon signed-rank post hoc tests with Holm-Bonferroni correction were conducted. Descriptive summaries, score distributions, and sentence-level hotspot analyses were used to evaluate semantic similarity patterns and identify localized low-scoring deviations.
RESULTS: Friedman tests assessed whether the AI services differed in performance, whereas descriptive summaries and visualizations were used to determine whether scores clustered in ranges consistent with strong semantic similarity to the gold standard. Friedman tests identified statistically significant between-service differences in 11 of the 20 (55%; P<.05) metric-language combinations. Subsequent paired Wilcoxon signed-rank post hoc tests with Holm-Bonferroni correction identified 11 significant pairwise differences, with adjusted P values ranging from <.001 to .049. Most of these differences were detected by surface-overlap metrics (10/11 for BLEU or METEOR), whereas only 1 of 11 was detected by a semantic metric (BLEURT), suggesting that many between-service differences were stylistic rather than meaning-altering. Descriptive and visual analyses further showed that semantic similarity was generally high across services, while low-scoring deviations clustered in specific linguistic hotspots, particularly domain headers, abstract health concepts, and short, context-dependent interface strings.
CONCLUSIONS: Among the evaluated high-resource European languages, AI translation services showed high semantic similarity to the validated human translations, although localized conceptual deviations persisted. These findings suggest that AI can support the generation of translations for PROM workflows in these languages; however, expert human review may still be required to confirm conceptual equivalence. The practical relevance of isolated header differences could not be assessed in the present study, whereas abstract health concepts and other clinically sensitive phrasing should be evaluated in further research.
PMID:42550950 | DOI:10.2196/78485