Categories
Nevin Manimala Statistics

Automating Motivational Interviewing Coding in Adolescent Substance Use Prevention: Human-AI Agreement Study

JMIR AI. 2026 Aug 25;5:e95964. doi: 10.2196/95964.

ABSTRACT

BACKGROUND: Motivational interviewing (MI) is widely used in preventive interventions, yet coding MI techniques and monitoring intervention adherence remain resource-intensive due to the reliance on manual transcription and expert review. Large language models (LLMs) offer a promising approach to automate these tasks, but their agreement with human coders in the context of prevention interventions has not been established.

OBJECTIVE: This study evaluated the agreement between an AI-based coder (OpenAI’s GPT 4.1) and trained human coders on two tasks: (1) identification of MI techniques (eg, open questions, affirmations, giving information) at the facilitator-message level and (2) completing a 21-item checklist of implementation adherence for a brief MI-based preventive intervention for adolescent substance use.

METHODS: Two certified MI facilitators independently coded 72 facilitator messages from 2 standardized Spanish-language Brief Intervention Based on Motivational Interviewing program (Intervención Breve Basada en Entrevista Motivacional [IBEM]) practice sessions with chatbot-simulated adolescent responses. The facilitators classified MI techniques using the OARS (open questions, affirmations, reflections, and summaries) framework and completed a 21-item implementation-adherence checklist. An AI-based coder (OpenAI’s GPT-4.1, accessed through the API) classified the same facilitator messages and checklist items using a structured prompt derived from the MI coding manual. MI techniques were compared at the facilitator-message level and implementation adherence at the session level. Intercoder agreement in use of MI techniques and implementation adherence was assessed using Cohen κ, Fleiss κ, and Cochran Q tests.

RESULTS: For use of MI techniques, the AI coder demonstrated moderate-to-substantial agreement with human coders across most techniques, including open questions (κ=0.66-0.69), affirmations (κ=0.66-0.77), and giving information (κ=0.91). No statistically significant differences in percentages of MI technique use were observed among the 3 coders, although agreement was the weakest for higher-inference categories such as complex reflections (κ=0.00). For MI implementation adherence, overall agreement was moderate (Fleiss κ=0.487), and pairwise agreement between the AI coder and 1 human coder was substantial (κ=0.67), exceeding the agreement observed between the 2 human coders (κ=0.53).

CONCLUSIONS: These findings provide support for the feasibility of using LLMs to recognize MI techniques and assess implementation adherence. The results support a human-AI collaborative model in which the AI coder “precodes” facilitator messages and flags sessions for expert review, while human coders retain responsibility for higher-inference judgments and shift their effort from routine coding toward contextual review and coaching feedback. Because the analyses are based on only 2 sessions, these results should be interpreted as early-stage, proof-of-concept evidence rather than a basis for large-scale deployment. Future research should compare different LLMs and evaluate whether AI-assisted coding improves the scalability of routine implementation monitoring.

PMID:42640265 | DOI:10.2196/95964

By Nevin Manimala

Portfolio Website for Nevin Manimala