Categories
Nevin Manimala Statistics

Large language models for early childhood caries risk triage: A validation study using standardized pediatric dental vignettes

J Indian Soc Pedod Prev Dent. 2026 Jul 1;44(4):392-398. doi: 10.4103/jisppd.jisppd_203_26. Epub 2026 Aug 25.

ABSTRACT

PURPOSE: To compare the diagnostic accuracy and public health utility of Gemini 2.5 Flash and GPT-4 for early childhood caries (ECC) risk triage in classifying presentations as high risk (warranting immediate referral) or low risk (amenable to preventive counseling).

MATERIALS AND METHODS: A standardized dataset of 50 clinical vignettes (25 high risk and 25 low risk) was evaluated against both large language models via official application programming interfaces using a blinded, prospective comparative design (CONSORT-AI; TRIPOD+AI). This constitutes a vignette-based internal validation study and does not represent a clinical diagnostic accuracy study. Ground truth was established by two consultant pediatric dentists (intraclass correlation coefficient = 0.92). Performance was quantified using sensitivity, specificity, positive predictive value, negative predictive value, and false negative rate (FNR). Readability, actionability, and multilingual competence in Hindi and Tamil were additionally assessed.

RESULTS: Both models exceeded the 80% accuracy threshold. GPT-4 achieved significantly higher sensitivity (92.0% vs. 80.0%; P = 0.032) and a lower FNR (8.0% vs. 20.0%), with superior readability (Flesch-Kincaid Grade Level 5.9 vs. 8.1), actionability (4.8 vs. 4.2), and cross-lingual consistency. Both models achieved 100% safety disclaimer compliance in English.

CONCLUSIONS: GPT-4 demonstrates a statistically superior safety profile for ECC risk triage. As a Level 1 adjunct, it can assist pediatric dentists by prescreening caregiver-reported presentations, reducing diagnosis lag, and directing specialist capacity toward confirmed high-risk pathology without replacing clinical judgment.

CLINICAL RELEVANCE: This study establishes that GPT-4 is a safer and more accessible AI triage tool than Gemini 2.5 Flash for pediatric dental screening, particularly in linguistically diverse, resource-constrained settings where specialist access is limited.

PMID:42641091 | DOI:10.4103/jisppd.jisppd_203_26

Categories
Nevin Manimala Statistics

Impact of a Visualization Training Platform on PICU Nurses’ Pain Assessment Skill in Critically Ill Children: A Mixed-Methods Randomized Controlled Trial

Nurs Open. 2026 Aug;13(8):e70741. doi: 10.1002/nop2.70741.

ABSTRACT

AIMS: Pain is common and challenging to assess in the paediatric intensive care unit (PICU), highlighting the need for targeted training. We previously developed PainScan, a visualization-based training platform and aimed to investigate the efficacy of PainScan in improving PICU nurses’ pain assessment competencies and to describe their learning experience with PainScan.

DESIGN: This study employed a mixed-methods design, integrating a superiority, randomized parallel controlled trial with embedded qualitative research. The RCT component of this study was prospectively registered at ClinicalTrials.gov (NCT06431802) prior to participant enrollment.

METHODS: Eligible PICU nurses from a tertiary children’s hospital in Shanghai, China, took part in the trial. 70 PICU nurses were recruited and randomly assigned to either the PainScan training group or the on-site training group using stratified randomization, with 35 individuals in each group, over 4 weeks. Primary outcomes were PICU nurses’ knowledge and skill levels in pain assessment. Outcome assessors were blinded to group allocation. Following the training, a descriptive qualitative study was conducted using semi-structured interviews with 11 PICU nurses, who had participated in PainScan training, to gain insights into their learning experience.

RESULTS: A total of 69 participants (35 in the PainScan group, 34 in the on-site training group) were included in the primary outcome analysis, as one participant (1.4%) in the on-site training group withdrew due to a rotation change. Both PainScan training and on-site training significantly improved PICU nurses’ knowledge (PainScan: t = 10.249, on-site: t = 10.592, both p < 0.001) and skill levels in pain assessment (PainScan: Z = -5.138, on-site: Z = -4.847, both p < 0.001). Compared to traditional on-site training, the PainScan showed superior skill development (Z = -3.644, p < 0.001). Qualitative findings revealed that the platform: (1) competency enhancement in pain assessment through PainScan training; (2) advantages of the digital learning platform; and (3) limitations of the PainScan training.

CONCLUSION: PainScan effectively enhanced PICU nurses’ pain assessment competencies. It also enriched their learning experience through a realistic, immersive and flexible approach. However, improvements are needed in supervision, engagement and scalability.

REPORTING METHODS: The randomized controlled trial was reported according to CONSORT 2025, and the qualitative component was reported following COREQ.

PATIENT OR PUBLIC CONTRIBUTION: This study shows that PainScan is a practical and cost-effective training tool for PICU nurses. It can enhance pain assessment skills with flexible learning. Addressing feedback on supervision and engagement could further optimize its use in clinical practice to improve patient care.

PMID:42641083 | DOI:10.1002/nop2.70741

Categories
Nevin Manimala Statistics

Estimated Prevalence of Developmental Language Disorder Risk in West Virginia Schools

Lang Speech Hear Serv Sch. 2026 Aug 25:1-18. doi: 10.1044/2026_LSHSS-25-00248. Online ahead of print.

ABSTRACT

PURPOSE: Developmental language disorder (DLD) is linked to long-term academic and social difficulties. Despite these lasting impacts, state-level prevalence estimates are limited. This study estimated prevalence among early-elementary students in West Virginia who may be at risk for DLD and examined variation by grade, locale, and assessment instrument.

METHOD: A cross-sectional screening battery was administered to 252 students in Grades 1-2 across six Title I schools. Children with Kaufman Brief Intelligence Test-Second Edition Matrices standard scores ≤ 70 were excluded prior to analysis. A dual-criterion definition of DLD risk was used: ≤ 80 on the Clinical Evaluation of Language Fundamentals-Fifth Edition (CELF-5) core composite and ≤ 92 on the Test of Narrative Language-Second Edition (TNL-2). Descriptive statistics and mixed-effects logistic regression (random intercept for school) tested effects of grade, gender, and locale.

RESULTS: Using the dual-criterion definition, 72 of 252 students (28.6%) met criteria for being at risk for DLD. Adjusted models showed higher odds of DLD risk in town versus urban schools (odds ratio = 3.06, 95% confidence interval [1.31, 7.15], p = .010), whereas grade and gender were not significant predictors after accounting for school-level clustering. Instrument-specific classification rates diverged substantially, with the TNL-2 identifying a higher proportion of students relative to the CELF-5.

CONCLUSIONS: Although based on a relatively small yet representative sample of West Virginia students, prevalence estimates substantially exceed commonly cited large-scale population estimates. Results underscore the influence of instrument choice on suspected prevalence and the need for locally calibrated screening protocols and targeted service planning.

PMID:42641081 | DOI:10.1044/2026_LSHSS-25-00248

Categories
Nevin Manimala Statistics

Automating Motivational Interviewing Coding in Adolescent Substance Use Prevention: Human-AI Agreement Study

JMIR AI. 2026 Aug 25;5:e95964. doi: 10.2196/95964.

ABSTRACT

BACKGROUND: Motivational interviewing (MI) is widely used in preventive interventions, yet coding MI techniques and monitoring intervention adherence remain resource-intensive due to the reliance on manual transcription and expert review. Large language models (LLMs) offer a promising approach to automate these tasks, but their agreement with human coders in the context of prevention interventions has not been established.

OBJECTIVE: This study evaluated the agreement between an AI-based coder (OpenAI’s GPT 4.1) and trained human coders on two tasks: (1) identification of MI techniques (eg, open questions, affirmations, giving information) at the facilitator-message level and (2) completing a 21-item checklist of implementation adherence for a brief MI-based preventive intervention for adolescent substance use.

METHODS: Two certified MI facilitators independently coded 72 facilitator messages from 2 standardized Spanish-language Brief Intervention Based on Motivational Interviewing program (Intervención Breve Basada en Entrevista Motivacional [IBEM]) practice sessions with chatbot-simulated adolescent responses. The facilitators classified MI techniques using the OARS (open questions, affirmations, reflections, and summaries) framework and completed a 21-item implementation-adherence checklist. An AI-based coder (OpenAI’s GPT-4.1, accessed through the API) classified the same facilitator messages and checklist items using a structured prompt derived from the MI coding manual. MI techniques were compared at the facilitator-message level and implementation adherence at the session level. Intercoder agreement in use of MI techniques and implementation adherence was assessed using Cohen κ, Fleiss κ, and Cochran Q tests.

RESULTS: For use of MI techniques, the AI coder demonstrated moderate-to-substantial agreement with human coders across most techniques, including open questions (κ=0.66-0.69), affirmations (κ=0.66-0.77), and giving information (κ=0.91). No statistically significant differences in percentages of MI technique use were observed among the 3 coders, although agreement was the weakest for higher-inference categories such as complex reflections (κ=0.00). For MI implementation adherence, overall agreement was moderate (Fleiss κ=0.487), and pairwise agreement between the AI coder and 1 human coder was substantial (κ=0.67), exceeding the agreement observed between the 2 human coders (κ=0.53).

CONCLUSIONS: These findings provide support for the feasibility of using LLMs to recognize MI techniques and assess implementation adherence. The results support a human-AI collaborative model in which the AI coder “precodes” facilitator messages and flags sessions for expert review, while human coders retain responsibility for higher-inference judgments and shift their effort from routine coding toward contextual review and coaching feedback. Because the analyses are based on only 2 sessions, these results should be interpreted as early-stage, proof-of-concept evidence rather than a basis for large-scale deployment. Future research should compare different LLMs and evaluate whether AI-assisted coding improves the scalability of routine implementation monitoring.

PMID:42640265 | DOI:10.2196/95964

Categories
Nevin Manimala Statistics

Geographical Access to Preferred Pharmacies in Medicare Part D

Am J Manag Care. 2026 Jul 1;32(7):e242-e248. doi: 10.37765/ajmc.2026.89989.

ABSTRACT

OBJECTIVES: Objective: To examine geographic access to preferred pharmacies in Medicare Part D prescription drug plans (PDPs) and Medicare Advantage PDPs (MAPDs).

STUDY DESIGN: We conducted a nationwide, retrospective analysis using Part D preferred pharmacies’ data, Zip Code Tabulation Areas (ZCTAs) from the US census, and socioeconomic characteristics from the American Community Survey and the Area Health Resources Files from 2010 to 2024.

METHODS: We computed mean distance from the centroid of each ZCTA to the nearest preferred pharmacy in stand-alone PDPs and in MAPDs. We compared mean distance as well as additional distance to a preferred vs a nonpreferred pharmacy, across rural and urban ZCTAs, and across geographic areas with relatively high densities of different racial/ethnic groups. We employed regression models to evaluate the association between ZCTA-level characteristics and additional distance to a preferred pharmacy.

RESULTS: Rural areas in the West North Central, Mountain, and East South Central divisions of the US tended to lack preferred pharmacies in PDPs and MAPDs, as did some urban areas. Among ZCTAs with preferred pharmacies, a rural ZCTA was associated with a 0.734-mile longer additional distance ( P < .01) among PDPs, which was more than double the mean additional distance. For MAPD plans, the additional distance for rural ZCTAs was 0.320 miles longer ( P < .01), an 80% increase relative to the mean.

CONCLUSIONS: Many rural and urban areas in the US lack preferred pharmacies. In areas with preferred pharmacies, distances were generally minimal relative to the cost savings, but rural-urban disparities persist.

PMID:42640262 | DOI:10.37765/ajmc.2026.89989

Categories
Nevin Manimala Statistics

State Medicaid Budgetary Implications of New Cancers

Am J Manag Care. 2026 Jul 1;32(7):e234-e241. doi: 10.37765/ajmc.2026.89988.

ABSTRACT

OBJECTIVES: The burden of a new cancer diagnosis on state Medicaid programs is not well understood. Our study aimed to (1) quantify Medicaid program spending, use of health care services, and enrollment patterns among Medicaid beneficiaries newly diagnosed with cancer; and (2) assess how spending and health care use differ between a new metastatic and a new nonmetastatic cancer diagnosis.

STUDY DESIGN: Cross-sectional study.

METHODS: This study used Transformed Medicaid Statistical Information System Analytic Files (2016-2019) to identify Medicaid beneficiaries 50 years and older who were newly diagnosed with cancer in 2017 and 2018. We provide descriptive evidence about their use of health care resources and used linked Medicare-Medicaid claims to follow enrollment patterns among dual-eligible beneficiaries. We also used generalized estimating equations to compare Medicaid program spending among those diagnosed with a new metastatic vs nonmetastatic cancer. Study measures were differences in total Medicaid spending, hospitalizations and emergency department visits per year, and Medicaid enrollment status at the end of the study (ie, continuously enrolled, disenrolled, or died) by metastatic status.

RESULTS: We identified 291,014 new cancer diagnoses among Medicaid beneficiaries 50 years and older (18.8% metastatic vs 81.2% nonmetastatic). A metastatic cancer diagnosis was associated with higher hospitalizations, emergency department use, and spending compared with a nonmetastatic cancer diagnosis.

CONCLUSIONS: Our results quantify the cost and health care resource utilization burden borne by state Medicaid programs following a new cancer diagnosis. These findings can inform policy makers about the budgetary considerations associated with investing in reducing late-stage cancer diagnoses, such as through early detection and increased access to care.

PMID:42640261 | DOI:10.37765/ajmc.2026.89988

Categories
Nevin Manimala Statistics

Blood Pressure Control Among Adherent vs Nonadherent Medicare Patients With Hypertension

Am J Manag Care. 2026 Jul 1;32(7):e228-e233. doi: 10.37765/ajmc.2026.89987.

ABSTRACT

OBJECTIVES: Medication adherence measures for hypertension drive financial incentives and patient care strategies for many value-based health care programs. The association between medication adherence and hypertension control as a value-based quality measure is not well described. The objectives of this study were to compare blood pressure (BP) control among Medicare and Medicare Advantage patients who meet vs do not meet the medication adherence for hypertension (MAH) measure and compare patient and prescription factors associated with adherence vs disease control.

STUDY DESIGN: Retrospective cohort study using linked payer and electronic health record data from a single health system.

METHODS: Medicare patients eligible for MAH in 2023 were categorized as adherent or nonadherent according to measure specifications, and hypertension control was defined as a primary care office BP less than 140/90 mm Hg or less than 130/80 mm Hg. Multivariable logistic regression was used to identify factors independently associated with adherence and disease control.

RESULTS: Of the 5206 patients evaluated, 77.9% were classified as adherent to the MAH measure. Of adherent patients, 77.2% achieved a BP less than 140/90 mm Hg vs 73.5% of nonadherent patients ( P = .01), and 45.9% and 43.4%, respectively, achieved a BP of less than 130/80 mm Hg ( P = .14). White race and younger age were positively associated with hypertension control, whereas extended days’ supply and non-White race were associated with meeting the MAH measure.

CONCLUSIONS: Patients classified as adherent were more likely to achieve a BP less than 140/90 mm Hg, but achievement of BP less than 130/80 mm Hg did not differ between groups. A sizeable proportion of patients who did not meet the adherence measure (73.5%) achieved a BP less than 140/90 mm Hg. Factors associated with adherence vs disease control differed. Emphasis on disease control over medication adherence measures may more meaningfully reflect high-value clinical care.

PMID:42640260 | DOI:10.37765/ajmc.2026.89987

Categories
Nevin Manimala Statistics

Invasive vs Non-invasive Management of Elderly Patients with Non-ST-Elevation Myocardial Infarction (NSTEMI)

Ir Med J. 2026 Aug 20;119(7):131.

ABSTRACT

AIM: Management of non-ST-elevation myocardial infarction (NSTEMI) in older patients remains challenging, particularly following the SENIOR-RITA trial, which did not demonstrate a mortality benefit with invasive management. This audit aims to compare outcomes between invasive and non-invasive management strategies in patients aged ≥75 years presenting with NSTEMI at Connolly Hospital, Blanchardstown, and to relate these findings to European Society of Cardiology (ESC) guidelines and SENIOR-RITA.

METHODS: A retrospective audit of 50 NSTEMI patients aged ≥75 admitted from 2020 to 2024 examined mortality, length of stay, complications, and readmission rates. Survival was analysed with Kaplan-Meier.

RESULTS: Of 50 patients, 29 (58%) underwent invasive management. Baseline features were similar between the invasive and conservative groups, though frailty was infrequently documented with formal assessment tools. Invasive patients had shorter stays (p=0.041) and lower 6-month mortality (14.3% vs 85.7%, p=0.03254). Kaplan-Meier analysis showed higher survival in the invasive group. Complication and readmission rates were similar.

DISCUSSION: Invasive management was associated with a shorter hospital stay and lower mortality. However, the findings should be interpreted cautiously, given the retrospective design and possible selection bias. These findings do not contradict the SENIOR-RITA trial, and local practice aligned with ESC guidelines.

PMID:42640250

Categories
Nevin Manimala Statistics

Safety Risk Analysis and Systematic Improvement Strategies for Intrahospital Transport of Patients Following General Anesthesia: A Retrospective Study Based on 255 Cases

J Perianesth Nurs. 2026 Aug 25:S1089-9472(26)00453-3. doi: 10.1016/j.jopan.2026.08.006. Online ahead of print.

ABSTRACT

PURPOSE: To systematically analyze the current safety status of intrahospital transport for patients recovering from general anesthesia, identify primary risk factors, and propose data-driven, systematic improvement strategies.

DESIGN: A retrospective analysis was conducted on transport monitoring records of 255 patients who underwent general anesthesia in a tertiary hospital.

METHODS: Transport duration, data integrity rate, and changes in oxygen saturation (SpO₂) and pulse rate (PR) were statistically analyzed.

FINDINGS: The mean transport duration was (9.68 ± 4.25) minutes, with a median time of 9.40 minutes. The overall monitoring data integrity rate was 75.35%. The incidence of hypoxemia (SpO₂ less than 90%) reached 50.59% (129/255). The mean minimum SpO₂ was (88.96 ± 4.40)%, and the mean maximum PR was (89.22 ± 16.46) bpm, indicating significant physiological fluctuations.

CONCLUSION: Even after meeting discharge criteria, patients undergoing general anesthesia exhibit physiological instability during intrahospital transport, characterized by a high incidence of hypoxemia and significant gaps in monitoring data, which may precipitate anesthesia-related safety incidents. Therefore, constructing a safety system encompassing standardized procedures, enhanced team safety awareness, and technological empowerment is a crucial measure to mitigate transport risks and ensure patient safety following general anesthesia.

PMID:42640243 | DOI:10.1016/j.jopan.2026.08.006

Categories
Nevin Manimala Statistics

Reverse Total Wrist Arthroplasty: New Design and Biomechanical Advantages

J Hand Surg Am. 2026 Aug 25:S0363-5023(26)00513-7. doi: 10.1016/j.jhsa.2026.06.005. Online ahead of print.

ABSTRACT

PURPOSE: Although fourth generation total wrist arthroplasties (TWAs) have improved implant longevity compared to earlier generations, high failure rates persist. This study aimed to assess how a shift in the center of rotation will affect kinematic outcomes of the joint. This has been accomplished through a TWA redesign, known as the reverse total wrist arthroplasty (RTWA).

METHODS: Eight cadaveric specimens (75.1 ± 10.9 years, 8 men) were implanted with both a custom load sensing TWA and a novel redesigned RTWA. For both reconstructed states, the specimens were mounted and actuated through flexion-extension (FE) and radial-ulnar deviation (RUD) using an active motion simulator. Muscle forces and articular loading patterns were recorded in both states for comparison of the implant designs. Statistical analysis completed showed RUD motion was underpowered and are therefore presented as exploratory and descriptive because of limited power.

RESULTS: In general, during both FE and RUD the RTWA required a lower magnitude of force from the muscles compared to the TWA state. Throughout FE, four of the five muscles examined required less force after RTWA implantation. Through RUD, three of five required less force. In addition, through both ranges of motion the carpal component joint loading was decreased in all wrist positions (41.6 ± 8.6% and 38.1 ± 7.1% for extension and flexion motion, respectively) after implantation of the RTWA compared to the TWA.

CONCLUSIONS: The redesigned RTWA performed well compared to a standard TWA implant and there was a noted decrease in both the muscle forces and joint loading following RTWA implantation.

CLINICAL RELEVANCE: The new RTWA design may have the potential to reduce the magnitude of force required by muscles to actuate the joint and reduce the high loading patterns within the distal carpal component that are noted after implantation of TWA implants.

PMID:42640233 | DOI:10.1016/j.jhsa.2026.06.005