Evid Based Dent. 2026 Aug 8. doi: 10.1038/s41432-026-01238-8. Online ahead of print.
ABSTRACT
OBJECTIVE: To evaluate the accuracy and reliability of three AI platforms ChatGPT, Perplexity, and Google Gemini in assessing the methodological quality of systematic reviews using the AMSTAR 2 checklist, compared with expert manual evaluation in dental research.
METHODS: A cross-sectional comparative study was conducted to assess the performance of three AI platforms ChatGPT, Perplexity, and Google Gemini in evaluating the methodological quality of 35 systematic reviews using the AMSTAR 2 checklist. Manual assessments by a domain expert served as the reference standard. Each AI system was prompted with a standardized AMSTAR 2 query, and item-level outputs were collected for direct comparison. Key metrics included percentage agreement, error proportions, and inter-rater reliability measured by Cohen’s kappa. Error proportions represent the proportion of discordant assessments out of total valid pairwise comparisons across 16 AMSTAR-2 items. Differences between LLM-generated and reference AMSTAR-2 ratings were summarized using effect estimates with corresponding 95% confidence intervals. Comparative performance across platforms was assessed based on confidence-interval overlap rather than hypothesis testing. Results are presented as effect estimates with corresponding 95% confidence intervals, without hypothesis testing or statistical dichotomization. This approach provided a robust and reproducible framework to benchmark AI-assisted quality appraisal in dental evidence synthesis.
RESULTS: Among 35 systematic reviews assessed, Perplexity demonstrated the highest agreement with expert AMSTAR-2 ratings (error proportion: 19.0%; weighted κ_w: 0.78, 95% CI 0.71-0.85), followed by ChatGPT (error proportion: 22.9%; weighted κ_w: 0.62, 95% CI 0.54-0.70) and Google Gemini (error proportion: 43.9%; weighted κ_w: 0.41, 95% CI 0.33-0.49). Perplexity also achieved the best sensitivity (81.3%, 95% CI 76.5-85.4%) and specificity (82.7%, 95% CI 78.1-86.5%) for correctly identifying high-quality reviews. Non-overlapping 95% confidence intervals suggest meaningful differences in performances among platforms, with Perplexity showing superior agreement across all metrics. Across all platforms, agreement was generally higher for non-critical AMSTAR-2 domains involving clear and structured reporting, whereas performance was weaker for critical domains requiring interpretation of complex methodological details, risk-of-bias considerations, and evidence synthesis procedures.
CONCLUSIONS: Perplexity demonstrated the highest accuracy and agreement with expert assessments of the methodological quality of systematic reviews, suggesting its potential as a supportive AI tool for AMSTAR-2-based appraisal in dental evidence synthesis. In contrast, systematic biases observed in ChatGPT and Google Gemini underscore the continued need for human oversight to ensure the validity of methodological assessments. Differences in agreement and error proportions were observed across all models when compared with expert AMSTAR-2 evaluations, indicating meaningful variability in methodological appraisal performance, reinforcing that AI-assisted appraisal of systematic review methodology should complement rather than replace expert human judgment in dental research.
PMID:42570960 | DOI:10.1038/s41432-026-01238-8