Categories
Nevin Manimala Statistics

Comparing a Large Language Model to Human-Generated Retention Messages for a Family Healthy Weight Program

Am J Health Promot. 2026 Jul 29:8901171261472636. doi: 10.1177/08901171261472636. Online ahead of print.

ABSTRACT

BackgroundText messaging can improve attendance and retention in community health programs; however, message development can be resource intensive.PurposeTo compare human and large language model (LLM)-generated retention messages in terms of creation time, clarity, appropriateness, and alignment with behavior change principles.DesignMixed methods using expert message ratings and qualitative feedback from message developers.SettingBuilding Healthy Families, a family healthy weight program.SampleExperts in behavioral science or related fields (n = 21) and message developers (n = 3).MeasuresMatched message pairs were rated on a 5-point Likert scale for clarity, appropriateness, and alignment with Social Cognitive Theory principles.AnalysisWilcoxon signed-rank tests compared message ratings. Equivalence testing using a ±10% equivalence interval and 90% confidence intervals assessed similarity between message types.ResultsLLM-generated messages yielded significantly higher ratings than human-generated messages for 4 of 5 message pairs on clarity and 3 of 5 rated on appropriateness (all P‘s < 0.05). Messages were statistically equivalent for 3 of 5 message pairs rated for behavior change theory alignment (all P‘s < 0.05). Time to develop the human- and LLM-generated messages was approximately the same with similar averaged Flesch-Kincaid grade level scores. However, the LLM-generated messages were shorter and more concise.ConclusionThis study suggests the potential for LLMs to create messages more efficiently and reduce, but not eliminate, workload for health promotion professionals.

PMID:42522506 | DOI:10.1177/08901171261472636

By Nevin Manimala

Portfolio Website for Nevin Manimala