Categories
Nevin Manimala Statistics

From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology

Annu Rev Psychol. 2026 Aug 10. doi: 10.1146/annurev-psych-100925-034807. Online ahead of print.

ABSTRACT

Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human instruments to LLMs without establishing that the outputs are reliable or interpretable, raising the risk of measurement phantoms-statistical regularities mistaken for genuine psychological phenomena. This review argues that robust AI psychological research requires integrating two methodological traditions: psychometric validation of what a score means and causal inference standards for what the results warrant. It develops a dual-validity framework in which evidentiary demands scale with scientific ambition: from tool use through behavioral characterization and human simulation to cognitive modeling. Classifying text may require only accuracy and reliability; claiming that an LLM simulates anxiety or illuminates cognitive mechanisms requires additional evidence, including construct validity evidence and experimental controls. Progress depends on developing computational analogs of psychological constructs rather than assuming human measures automatically apply to language models.

PMID:42574761 | DOI:10.1146/annurev-psych-100925-034807

Categories
Nevin Manimala Statistics

Nurse-Led Ambient AI Scribe for Patient Safety Incident Investigation Reports (Project NARRATE): Retrospective Pre-Post Comparative Document-Quality Study

JMIR Nurs. 2026 Aug 10;9:e100775. doi: 10.2196/100775.

ABSTRACT

BACKGROUND: Patient safety investigation reports support organizational learning only when they are complete, usable, and sufficiently detailed. Conventional free-text reports are often inconsistent and may omit information needed for review and learning. Project NARRATE (Nursing AI-Refined for Accurate Transcription of Events) is a nursing-led ambient artificial intelligence workflow that uses prompts aligned with the World Health Organization Minimal Information Model for Patient Safety Incident Reporting and Learning Systems, Situation-Background-Assessment-Recommendation output, and visible missing-information cues to support structured supervisor reporting.

OBJECTIVE: This study aimed to compare the completeness and narrative quality of conventional and NARRATE-period supervisor investigation reports for falls and medication administration-related incidents.

METHODS: We conducted a retrospective pre-post document-quality study at a tertiary academic medical center in Singapore. We reviewed 150 deidentified completed supervisor investigation reports: 75 conventional reports from June to August 2025 and 75 confirmed NARRATE reports from January to March 2026. NARRATE use was voluntary, and recorded use represented approximately 40% of eligible postimplementation reports. Two blinded reviewers rated reports using a World Health Organization (WHO)-aligned completeness checklist and an adapted 8-domain Physician Documentation Quality Instrument (PDQI). Report-level comparisons were adjusted for repeated reports by the same supervisor using random-intercept linear mixed-effects models. A stratified 60-report plain-paragraph rerating examined whether visible structure influenced ratings.

RESULTS: All 150 reports were analyzed. Unadjusted mean WHO total completeness was 11.81 (SD 3.39) for conventional reports and 13.61 (SD 2.54) for NARRATE reports; the unadjusted difference was 1.80 points, and the cluster-adjusted mean difference was 1.95 (95% CI 0.91-3.00; P<.001). The adapted PDQI mean was 3.61 (SD 0.52) and 4.13 (SD 0.34), respectively; the unadjusted difference was 0.52 points, and the cluster-adjusted mean difference was 0.53 (95% CI 0.37-0.69; P<.001). In the plain-paragraph sensitivity analysis, the completeness advantage remained (adjusted mean difference 1.70, 95% CI 0.27-3.14; P=.02), as did the adapted PDQI mean advantage (adjusted mean difference 0.25, 95% CI 0.06-0.43; P=.009). Explanation, organization, and comprehensibility remained significantly higher after deformatting; actions were borderline (P=.05), and synthesis, internal consistency, and fairness/balance were not statistically significant.

CONCLUSIONS: Among voluntary early adopters, NARRATE use was associated with more complete reports and higher adapted PDQI mean scores after accounting for supervisor clustering. Because recorded use represented approximately 40% of eligible postimplementation reports and users self-selected, findings may reflect adopter and supervisor characteristics. Results support the structured workflow as a whole, not any single AI component, and do not demonstrate downstream patient-safety effects. Confirmatory evaluation under broader adoption with a concurrent, reliably classified comparison group is needed.

PMID:42574744 | DOI:10.2196/100775

Categories
Nevin Manimala Statistics

Attrition in Digital Self-Management Interventions for Patients With Metabolic Dysfunction Associated Steatotic Liver Disease (MASLD): Mixed Methods Systematic Review

J Med Internet Res. 2026 Aug 10;28:e89124. doi: 10.2196/89124.

ABSTRACT

BACKGROUND: Lifestyle modification delivered through digital self-management is central to metabolic dysfunction-associated steatotic liver disease (MASLD) care, yet long-term engagement remains the threshold beyond which clinical benefit is realized. Understanding attrition requires examining both retention (dropout) and adherence (usage quality), which are often evaluated in isolation. Existing systematic reviews of digital interventions for MASLD have focused predominantly on clinical effectiveness, leaving less attention on attrition.

OBJECTIVE: This study aimed to integrate quantitative retention metrics with qualitative adherence insights and characterize the determinants of attrition in digital MASLD self-management interventions.

METHODS: Following PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta-Analyses literature search extension) and PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines, a comprehensive search of five databases (PubMed, Web of Science, Embase, Cochrane Library, and CINAHL) was conducted. The initial search was conducted in June 2025, and a subsequent update was made on April 17, 2026. Eligible studies enrolled adults with MASLD or nonalcoholic fatty liver disease in structured digital self-management interventions reporting retention or adherence data. Methodological quality was assessed using the Mixed Methods Appraisal Tool. A convergent segregated design was adopted. Retention proportions were pooled using random-effects meta-analysis with logit transformation, restricted maximum likelihood estimation, and Hartung-Knapp-Sidik-Jonkman adjustment. Adherence data were synthesized through inductive framework synthesis. Findings were subsequently integrated narratively.

RESULTS: In total, 21 studies met the eligibility criteria, of which 15 (n=1,032) contributed to the quantitative synthesis. The pooled retention proportion was 80% (95% CI 72%-87%) with substantial between-study heterogeneity (I²=73.6%). App-based platforms showed the highest point estimate and the lowest within-group heterogeneity, although no subgroup difference reached statistical significance. Adherence varied widely and was not amenable to meta-analytic pooling. Thematic synthesis identified 4 interacting domains shaping adherence, namely platform and design, human support and professional integration, motivational and behavioral strategies, and patient-level characteristics. Access friction at entry, gated coaching architecture, the absence of proximal biological feedback, and psychological comorbidity recurred as attenuators of long-term engagement.

CONCLUSIONS: This review innovatively integrates retention and adherence to provide a comprehensive framework of attrition dynamics specific to MASLD. While retention compared favorably with adjacent fields, long-term adherence depended less on platform type than on accessible human support, alignment of feedback with the disease’s silent course, and psychological screening. Although evidence certainty was rated very low under Grading of Recommendations Assessment, Development and Evaluation, reflecting blinding constraints intrinsic to digital interventions and a predominance of pilot or feasibility designs, these findings carry clear real-world implications. Future interventions would benefit from establishing standardized, component-level reporting that distinguishes retention from adherence, to reliably evaluate the true therapeutic potential of digital MASLD interventions.

PMID:42574740 | DOI:10.2196/89124

Categories
Nevin Manimala Statistics

Automated Features, Algorithms, and Technologies of Electronic Early Warning/Track-and-Trigger Systems: Systematic Review

J Med Internet Res. 2026 Aug 10;28:e58233. doi: 10.2196/58233.

ABSTRACT

BACKGROUND: Electronic early warning/track-and-trigger systems (EW/TTS) are crucial for patient monitoring, detecting clinical deterioration (CD), and activating rapid response teams. Understanding the current level of automation in EW/TTS is essential.

OBJECTIVE: This study aimed to provide a comprehensive overview and critical assessment of electronic EW/TTS, including automated features, algorithms, and technologies, following a published registered study protocol.

METHODS: Based on the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, we included studies from PubMed, Web of Science, and Scopus published between January 2010 and December 2025 describing EW/TTS applied in real-world settings, and electronic systems for CD detection. We excluded studies outside the clinical context or those that used manual scoring charts. We applied a descriptive narrative approach and a methodological quality assessment according to the Joanna Briggs Institute Critical Appraisal Checklist.

RESULTS: After removing outliers and duplicates, the query returned 1181 studies. The selected studies (n=43) reported CD as the primary objective in 24 of 44 (54.5%) reported primary objectives, with ICU transfer in 16 of 68 (23.5%) reported secondary objectives, and mortality prediction in 10 of 68 (14.7%) reported secondary objectives. EW/TTS primarily relied on vital signs and assessment scores, accounting for 42 of 67 (62.7%) reported clinical indexes to detect and predict CD effectively. Among the included systems, 18 of 43 (41.9%) had a measured automation level, 11 of 43 (25.6%) had a managed automation level, and 7 of 43 (16.3%) had a defined automation level. The studies focused on several technological domains, with a strong emphasis on data analytics (24/43, 55.8%) and hardware technologies (7/43, 16.3%). Predictive algorithms, including statistical and machine learning approaches, were used in 11 of 43 (25.6%) systems. Interoperable connectivity was reported in 30 of 43 (69.8%) systems, including connectivity with electronic health records, wearable devices, and communication platforms such as Ascom Unite, as well as integrations using standards such as Health Level Seven Fast Healthcare Interoperability Resource and Health Level Seven. Evaluations of the systems showed earlier warning (14/70, 20%), higher accuracy (12/70, 17.1%), and lower specificity (9/70, 12.9%) as the main reported outcomes. Electronic EW/TTS were most prevalent in the United States (15/43, 34.9%), the United Kingdom (6/43, 14%), and the Netherlands (6/43, 14%).

CONCLUSIONS: Current EW/TTS systems implemented a measured level of automation and primarily focused on patient monitoring in hospital surgery wards. More than half of EW/TTS featured data exchange capabilities and connectivity with other systems. Reported outcomes of EW/TTS included early warning, high accuracy, and lower specificity. However, the included evidence was limited by heterogeneous prediction targets, inconsistent performance metrics and time horizons, and poor reporting of development history and system failure. Using clinically validated wearable devices and establishing a standardized data collection framework may further improve system accuracy and reliability.

PMID:42574739 | DOI:10.2196/58233

Categories
Nevin Manimala Statistics

Psychedelic Use Covaries with Psychological Flexibility Through Mystical Experiences: Results of a Retrospective Web Survey

J Psychoactive Drugs. 2026 Aug 10:1-9. doi: 10.1080/02791072.2026.2710057. Online ahead of print.

ABSTRACT

The therapeutic effects of psychedelic-assisted treatments covary with acute mystical experiences as well as enhancements in psychological flexibility. Psychological flexibility, a key construct in acceptance and commitment therapy (ACT), has broad benefits for adaptive functioning and well-being, making it a vital focus of psychedelic-assisted therapy. More than 200 participants (54.9% female, 70% Caucasian, 72% with a college degree) completed the Mystical Experiences Questionnaire (MEQ-30) addressing their most profound mystical experience as well as the Multidimensional Psychological Flexibility Index (MPFI) and their psychedelic use. Analyses revealed that psychedelic use had an indirect effect on psychological flexibility through mystical experiences, underscoring the potential role of mystical experiences in fostering psychological growth. Participants who had used a psychedelic reported significantly higher scores on the MEQ-30 and MPFI compared to those with only non-drug mystical experiences. The flipped model did not achieve statistical significance. These findings suggest that mystical experiences might act as a catalyst for increases in psychological flexibility, and they underscore the need for further research on the mechanisms linking mystical experiences and psychological flexibility, particularly in clinical contexts. Psychedelic-assisted therapy, informed by frameworks like ACT, shows promise for enhancing psychological flexibility and fostering long-term mental health improvements. Preparatory and integration sessions that emphasize flexibility may improve treatment outcomes.

PMID:42574736 | DOI:10.1080/02791072.2026.2710057

Categories
Nevin Manimala Statistics

Depressive and anxiety symptoms, their predictors, and pregnancy outcomes among Omani pregnant women: A prospective cohort study

Womens Health (Lond). 2026 Jan-Dec;22:17455057261476574. doi: 10.1177/17455057261476574. Epub 2026 Aug 10.

ABSTRACT

BackgroundMaternal depression during pregnancy is a major global health concern that affects both mothers and infants. However, limited studies have examined maternal depression across pregnancy stages in Arabic-speaking populations, where fertility rates are high.ObjectivesThis study aims to evaluate the relationship between antenatal depression and anxiety symptoms during the early (8-12 weeks) and later stages (24-28 weeks) of pregnancy and their effects on maternal and neonatal outcomes among Omani women.DesignProspective cohort study.MethodsA prospective cohort design was used, involving 302 pregnant Omani women receiving antenatal care at Al Buraimi Hospital. Eligible participants were aged 18-45 years, between 8-12 weeks of gestation and expected to continue care at the same clinic. Depression and anxiety were measured at both stages using the Arabic version of the Edinburgh Postnatal Depression Scale (EPDS) and the EPDS-3A subscale. Statistical analyses, including chi-square tests and logistic regression, were used to examine associations between anxiety, depression, and pregnancy outcomes. Multiple regression analyses controlled for maternal age, marital status, parity, pre-pregnancy BMI, and household income.ResultsHigh levels of depressive and anxiety symptoms were identified, particularly as pregnancy advanced. Among 302 pregnant Omani women, the prevalence of depressive symptoms was 29.8% and anxiety symptoms was 24.8% in early pregnancy. Women with elevated EPDS scores had higher risks of caesarean delivery, low birth weight, and preterm birth. Elevated anxiety was associated with greater maternal distress and poorer neonatal outcomes.ConclusionFindings highlight the importance of integrating mental health screening into routine antenatal care in Oman. Early identification and management of depression and anxiety may reduce adverse outcomes for mothers and infants. Further research should investigate barriers to mental health services for pregnant women and the long-term developmental effects of antenatal depression on children.

PMID:42574733 | DOI:10.1177/17455057261476574

Categories
Nevin Manimala Statistics

Clear Masks Do Not Prevent Gains During Dynamic Temporal and Tactile Cueing Treatment: An Explanatory Sequential Mixed Methods Pilot Study

Am J Speech Lang Pathol. 2024 Sep 18;33(5):2438-2460. doi: 10.1044/2024_AJSLP-23-00473. Epub 2024 Aug 6.

ABSTRACT

PURPOSE: This study aimed to determine the outcomes and impact of Dynamic Temporal and Tactile Cueing (DTTC) treatment when clear vinyl masks were worn. DTTC is one of the few evidence-based treatments for children with childhood apraxia of speech (CAS). Given that DTTC relies on visual, auditory, and tactile cues, it was unknown if treatment gains would be demonstrated when masks were worn and how masking would impact the therapy experience for clinicians and caregivers.

METHOD: A sequential mixed methods design was used to study the efficacy of DTTC treatment in children with CAS when clear masks were worn. The quantitative phase used a multiple-baseline across-participants design. Four children (each 4 years of age) participated in the treatment protocol in which 24 sessions of DTTC were provided over 8 weeks while clear vinyl face masks were worn by participants and clinicians. Whole word accuracy on treated items and generalization to easy and hard untreated items were assessed during baseline, treatment, and follow-up. Semistructured interviews were conducted with clinicians and caregivers following treatment to explore the experience of masks being worn during treatment. Qualitative data were analyzed using descriptive thematic analysis.

RESULTS: Three children completed the treatment protocol. Visual and statistical analyses revealed that two participants demonstrated significant treatment effects, with one also demonstrating generalization. The remaining participant demonstrated marginal treatment gains. Qualitative findings revealed two main themes: “mask wearing was inconvenient but did not prevent therapy gains” and “in-person therapy with face masks was preferable to teletherapy.”

CONCLUSIONS: Masks did not prohibit treatment gains during DTTC therapy, with treatment effects of varying degrees shown for the three participants who completed the protocol. Together, quantitative and qualitative results indicate that mask wearing was, for most, a minor inconvenience that did not substantially interfere with the efficacy of DTTC treatment.

SUPPLEMENTAL MATERIAL: https://doi.org/10.23641/asha.26408854.

PMID:42574723 | DOI:10.1044/2024_AJSLP-23-00473

Categories
Nevin Manimala Statistics

Human-Edited Generative AI-Assisted Multiple-Choice Questions in Postgraduate Family Medicine: Blinded Cross-Sectional Comparative Psychometric Study

JMIR Med Educ. 2026 Aug 10;12:e100179. doi: 10.2196/100179.

ABSTRACT

BACKGROUND: Generative artificial intelligence (GenAI) is increasingly used to draft multiple-choice questions (MCQs) for health professions education, but much evidence concerns raw model outputs, expert ratings, or item difficulty alone. Educators edit GenAI drafts before use, and whether such items are psychometrically ready for postgraduate assessment remains unclear.

OBJECTIVE: This study aimed to compare human-edited GenAI-assisted and educator-crafted MCQs for postgraduate Family Medicine Applied Knowledge Test-level assessment, examining difficulty, discrimination, reliability, distractor functioning, and participant perceptions.

METHODS: We conducted a blinded cross-sectional, within-participant comparative psychometric evaluation in Singapore. Sixty best-of-five single-best-answer MCQs were evaluated, 30 human-edited GenAI-assisted items and 30 educator-crafted items, topic-matched across postgraduate FM domains and randomized across 2 assessment sets. Eligible participants were postgraduate doctors enrolled in FM residency or postgraduate family medicine programs, preparing for the Applied Knowledge Test, and blinded to item origin; incomplete paired responses were excluded. Outcomes included paired total scores, score correlation and agreement, Kuder-Richardson Formula 20 reliability, item difficulty index, corrected point-biserial discrimination, distractor functioning, and perceived difficulty, clarity, and relevance. Analyses used paired-sample tests, Pearson correlation, Fisher exact tests, and item-level psychometric statistics, with α=.05 and Bonferroni correction within comparison families.

RESULTS: Of 74 participants, 73 completed both item sets and were included in the analysis. The final sample comprised 36 graduate diploma in FM trainees, 5 MMed FM trainees, and 32 FM residents. Paired-sample testing showed lower scores on GenAI-assisted than educator-crafted items (mean 19.12, SD 2.83 vs mean 21.10, SD 3.42 out of 30; mean difference -1.97, 95% CI -2.72 to -1.23; P<.001; Cohen d=0.62), indicating that GenAI-assisted items were not easier. Scores were positively correlated (r=0.49, 95% CI 0.30-0.64; P<.001), but Bland-Altman analysis indicated limited agreement. Kuder-Richardson Formula 20 reliability was lower for GenAI-assisted items (0.38 vs 0.60). Mean difficulty index did not differ significantly (0.64 vs 0.70; mean difference -0.07, 95% CI -0.19 to 0.06; P=.29), and more GenAI-assisted items fell within the acceptable difficulty range (18/30, 60.0% vs 13/30, 43.3%). However, mean corrected point-biserial discrimination was lower for GenAI-assisted items (0.09 vs 0.18; mean difference -0.08, 95% CI -0.16 to -0.01; P=.04), and negative discrimination was more common (6/30, 20% vs 3/30, 10%). GenAI-assisted items also had more nonfunctioning and negatively discriminating distractors, although these differences were not statistically significant. Participant ratings of perceived difficulty, clarity, and practice relevance did not differ by origin.

CONCLUSIONS: Human-edited GenAI-assisted MCQs can achieve plausible difficulty, but difficulty and surface acceptability did not ensure assessment readiness. Using trainee response data, this study extends work on raw outputs or expert opinion. GenAI should be used as a drafting adjunct within educator-led workflows prioritizing key verification, distractor engineering, pilot testing, empirical item analysis, and repair before item-bank or summative use.

PMID:42574719 | DOI:10.2196/100179

Categories
Nevin Manimala Statistics

Assessing Overall and Mental/Emotional Health Among People Living With HIV/AIDS

AIDS Educ Prev. 2026 Aug;38(4):307-323. doi: 10.1521/aeap.2026.38.4.307.

ABSTRACT

To characterize the lived experience of persons living with HIV/AIDS (PLWHA) during the COVID-19 pandemic, we utilized an innovative survey to assess the experiences of 94 adult respondents receiving medical care and/or case management services from two Ryan White funded sites in New Brunswick, NJ, from May 2020 to November 2021. The Local Inventory of Needs and Knowledge-HIV (LINK-HIV) survey includes five indices and assesses overall and mental/emotional health of PLWHA. We demonstrate internal validity of the indices, describe relationships between indices and health outcomes, and characterize the health of this population. We identified specific unmet needs across the constructs of social determinants of health (SDOH) Needs, Stress due to unmet SDOH, Access to Healthcare, Stress due to COVID-19, and Person-Centered Primary Care. Respondents’ self-reported health was suboptimal (47.8% reported better overall health, 40.4% reported better mental/emotional health), and associations were seen between responses to the five indices and self-reported health outcomes.

PMID:42574704 | DOI:10.1521/aeap.2026.38.4.307

Categories
Nevin Manimala Statistics

Understanding the Relationship Between Intimate Partner Violence and HIV Status Disclosure Across Health Care Settings in Eastern and Southern Africa: A Scoping Review

AIDS Educ Prev. 2026 Aug;38(4):284-306. doi: 10.1521/aeap.2026.38.4.284.

ABSTRACT

HIV status disclosure is an ethical obligation and, in some jurisdictions, a legal requirement that supports treatment adherence. In Eastern and Southern Africa (ESA), where intimate partner violence (IPV) is prevalent, disclosure may have harmful consequences. The objective of this scoping review was to map the relationship between IPV and HIV disclosure across health care settings in ESA. Following PRISMA-ScR guidelines, peer-reviewed English-language studies (2012-2024) were identified through EBSCOhost, PubMed, and Google Scholar using a SPICE-informed search strategy. Thirty-six quantitative, qualitative, and mixed-methods studies met the inclusion criteria and were analyzed thematically. Three themes emerged: factors influencing disclosure, including relationship dynamics and fear of violence; the positive and negative consequences of disclosure, particularly IPV; and the role of health care workers, whose limited IPV training tended to increase risk. Integrating IPV screening, safety planning, and gender-sensitive training into HIV counseling is essential to support safe, client-led disclosure.

PMID:42574702 | DOI:10.1521/aeap.2026.38.4.284