Categories
Nevin Manimala Statistics

Intranational variation in prosthetic user populations: A comparative analysis of prosthetic and orthotic centers in central and northern Sri Lanka

Prosthet Orthot Int. 2026 Aug 5. doi: 10.1097/PXR.0000000000000574. Online ahead of print.

ABSTRACT

BACKGROUND: Population-level data on people with limb loss (PwLL) are needed for effective prosthetic provision and rehabilitation. These data are rare in low- and middle-income countries, and low- and middle-income countries are routinely dealt with as homogenous entities.

OBJECTIVES: Here, we investigate intranational variation in PwLL populations in Sri Lanka using long-term routinely collected data from prosthetic and orthotic (P&O) clinics.

STUDY DESIGN: Retrospective study.

METHODS: Until 2009, Sri Lanka experienced a 26-year civil war affecting primarily northern/eastern districts. Published data from 2 P&O centers in northern (Jaffna Jaipur Centre for Disability Rehabilitation, JJCDR; 1987-2018) and central (Centre for Handicapped, CFH; 1987-2024) districts were compared to investigate intranational variation in PwLL demographics and temporal trends. Summative statistics, χ2, and Mann-Whitney U tests (α < 0.05) were used to quantify differences.

RESULTS: Major differences included a larger population of males at the CFH (88.8%) compared with the JJCDR (77.4%) and-although war-related amputations were most common at both clinics-there were relatively more war-related amputees at the JJCDR (CFH: 44.8%; JJCDR: 62.7%). In addition, the JJCDR had a higher proportion of patients registered and injured in the same district (CFH: 62.22%; JJCDR: 66.11%).

CONCLUSIONS: Intranational variation in PwLL populations exists in Sri Lanka as evidenced by differences in populations between northern and central P&O clinics. This points toward differences in provisional needs and temporal health/societal trends underlying limb loss, which must be considered when provisioning health services.

PMID:42551032 | DOI:10.1097/PXR.0000000000000574

Categories
Nevin Manimala Statistics

Patient acceptance and satisfaction with Telehealth clinical consultation after an orthotic prescription: A cross-sectional questionnaire study

Prosthet Orthot Int. 2026 Aug 4. doi: 10.1097/PXR.0000000000000570. Online ahead of print.

ABSTRACT

BACKGROUND: Telehealth was more widely utilized in various medical field over the world since the COVID-19 pandemic. It could be an option to meet the growing health service demand. Objectives: This study aims to identify patient’s expectation and satisfaction in Telehealth as an alternative follow-up format.

STUDY DESIGN: Fifty patients presenting Hospital Department of Prosthetics and Orthotics were recruited to complete a questionnaire. Methods: The questionnaire evaluated subject’s expectation and satisfaction towards Telehealth in 5-point scale with result analysed with Fisher Exact test.

RESULTS: Statistical analysis revealed that: (1) shorter travel time and waiting time greatly encouraged subject to attend Telehealth; (2) body part exposure to camera, risk of diagnosis error, network and technological issue; (3) No significant satisfaction difference between Telehealth and traditional face-to-face follow-up subjects.

CONCLUSIONS: This is the first study in Hong Kong to evaluate patient’s feedback to Telehealth service in orthotic treatment. The findings could speed up the implementation of Telehealth into P&O treatment and improve the overall service quality.

PMID:42551026 | DOI:10.1097/PXR.0000000000000570

Categories
Nevin Manimala Statistics

Observational study of people with complete hand amputation using a multigrip myoelectric prosthesis after using a standard myoelectric prosthesis

Prosthet Orthot Int. 2026 Aug 4. doi: 10.1097/PXR.0000000000000571. Online ahead of print.

ABSTRACT

BACKGROUND: Clinical evidence substantiating the advantages multi-grip myoelectric hand prosthesis over conventional myoelectric hand prosthesis in daily life remains inconclusive.

OBJECTIVES: The study aimed to assess the efficacy of using a multi-grip myoelectric hand prosthesis after utilizing a conventional myoelectric hand prosthesis.

STUDY DESIGN: A before-and-after, multi-center, observational study included adults with complete hand amputation using a standard myoelectric hand prosthesis for more than 4 weeks.

METHODS: Patients were assessed at T0 with their standard myoelectric hand prosthesis, after a 4 to 6 weeks training and trial period with a multi-grip myoelectric hand prosthesis (T1) and then after an 8 to 10 weeks period of wearing (T2). The primary outcome measure was the change of upper limb function via the Orthotics and Prosthetics Users’ Survey (OPUS). Results: twenty-two participants were prescribed a definitive multi-grip hand myoelectric prosthesis. On average, the duration since amputation was 6.5 ± 11.4 years. Traumatic causes accounted for 77.3% (17/22) of the cases. The upper limb function component score of the OPUS exhibited a mean increase of 6.2 ± 10.8 points (95% CI 1.5 to 11.0) from T0 to T2. Statistically significant improvements were also noted in the quality-of-life component of OPUS, the Quick DASH score, and overall satisfaction with prosthesis.

CONCLUSIONS: In this study, the transition from using daily living standard myoelectric hand prosthesis to using a multi-grip myoelectric hand prosthesis yielded improvements in functionality, quality of life, and overall satisfaction (self-drafted) among individuals with complete hand amputation.

PMID:42551018 | DOI:10.1097/PXR.0000000000000571

Categories
Nevin Manimala Statistics

Sensor-Based Mobile App for Assessment and Tracking (SMAAT): Cross-Sectional Usability Study

JMIR Form Res. 2026 Aug 4;10:e98984. doi: 10.2196/98984.

ABSTRACT

BACKGROUND: Smartphone-based ecological momentary assessment (EMA) is increasingly used in digital health research to capture behaviors, symptoms, and psychological states in real time and in natural environments, offering advantages over traditional retrospective measures in terms of ecological validity and reduced recall bias. Despite the potential of EMA for advancing mobile health and digital phenotyping apps, accessible technical solutions that enable researchers without software engineering expertise to design and deploy EMA studies remain limited.

OBJECTIVE: This study aimed to (1) describe the development and core features of a sensor-based mobile app for assessment and tracking (SMAAT), a research platform for designing and deploying mobile surveys and EMA studies, and (2) report initial usability findings from a cross-sectional study using a novel swiping response format.

METHODS: SMAAT consists of a web-based dashboard for researchers and companion iOS (Apple Inc) and Android (Google) apps for participants, providing a visual survey builder, multiple notification schedules (including fixed, random, interval, event-based, and geofenced prompts), gamification mechanics, and tools to monitor participant compliance. To evaluate the platform in use, we conducted a between-participants cross-sectional study in which 97 university students and 132 online panel participants completed blocks of binary questions on smartphones using either a swiping or tapping response format, followed by usability and user-experience questionnaires.

RESULTS: In the cross-sectional study, SMAAT supported successful study setup, enrollment, and survey completion across both iOS and Android devices without major technical problems, and participants in both samples completed the full protocol in a single session. Performance on the binary tasks was generally high, with swiping and tapping showing broadly comparable response-time and accuracy patterns, and no clear disadvantages for swiping or consistent effects of response orientation. Usability and pragmatic user-experience ratings were high across conditions, with no meaningful differences between swiping and tapping, while Prolific participants reported higher usability and pragmatic quality than students. Hedonic ratings descriptively favored swiping, although this difference did not reach conventional statistical significance.

CONCLUSIONS: SMAAT is a flexible smartphone-based research platform for configuring and deploying mobile surveys and EMA studies with diverse item types and notification logics. Initial findings show that SMAAT can support reliable cross-sectional data collection across heterogeneous devices and samples, and that a swipe-based response format can be implemented without compromising task performance, usability, or pragmatic user experience relative to tapping.

PMID:42551009 | DOI:10.2196/98984

Categories
Nevin Manimala Statistics

Effects of Remotely Delivered and Web-Based Interventions on Depression Severity During the COVID-19 Pandemic: 3-Arm Randomized Controlled Trial

JMIR Ment Health. 2026 Aug 4;13:e88388. doi: 10.2196/88388.

ABSTRACT

BACKGROUND: The COVID-19 pandemic highlighted a critical need for effective population mental health approaches to target the most prevalent disorders (eg, depression) during periods of elevated community distress. The effectiveness of remotely delivered and web-based interventions should be investigated to identify and innovate high-quality models for population mental health service delivery.

OBJECTIVE: The primary objective investigated the effectiveness of adding Mindfulness-Based Cognitive Therapy for Resilience (MBCT-R)-a live, online, synchronous, remotely delivered, group-based intervention-to Cambridge Health Alliance MindWell (CHA-MW), a web-based population health screening and stratified support program, compared with CHA-MW alone, on depression symptom severity. The secondary objective evaluated adding internet Cognitive Behavioral Therapy (iCBT)-an asynchronous, web-based, individual, digital intervention-to CHA-MW, compared with CHA-MW alone.

METHODS: Participants (N=97) were randomized in a 2:2:1 ratio to receive MBCT-R+CHA-MW (n=37), iCBT+CHA-MW (n=41), or CHA-MW alone (n=19) in a 3-arm randomized clinical trial, from May 2021 to September 2022 in an urban public safety net hospital outpatient setting. CHA-MW served as a low-intensity control condition. For the MBCT-R+CHA-MW arm, MBCT-R was an 8-session program mildly adapted from MBCT to address COVID-19-related risks for depression. For the iCBT+CHA-MW arm, iCBT was a 6-session curriculum added to CHA-MW. All study procedures, including regular mental health symptom screenings, were conducted remotely or via a web-based platform. The primary outcome was change in depression symptom severity during the 24-week study period using an intention-to-treat approach that used generalized linear mixed-effects models to evaluate the comparative effectiveness of MBCT-R+CHA-MW vs CHA-MW over time. A secondary analysis compared iCBT+CHA-MW vs CHA-MW on depression severity. Completer analyses were conducted (per-protocol 6+ sessions). The secondary outcome was mental health visit utilization frequency during the study period.

RESULTS: Both MBCT-R+CHA-MW (mean difference -14.1, 95% CI -21.0 to -7.2) and CHA-MW (mean difference -15.2, 95% CI -21.8 to -8.6) had significant reductions in depression symptom severity, with no statistically significant between-group differences. iCBT+CHA-MW (mean difference -12.7, 95% CI -17.4 to -8.1) also reduced depression symptoms but without between-group differences when compared with CHA-MW. Intervention completion rates were low (MBCT-R: 30% and iCBT: 24%), and completers demonstrated significantly greater reductions in depression severity than noncompleters (mean difference -8.5, 95% CI -16.2 to -0.8). Overall mental health clinician visits by group had no statistically significant differences. CHA-MW had the largest increase in participants with new psychopharmacology treatment visits during the 24-week study (CHA-MW +21%, MBCT-R +10%, and iCBT -5%).

CONCLUSIONS: MBCT-R+CHA-MW, iCBT+CHA-MW, and CHA-MW were each effective in treating depression, without any intervention demonstrating superiority in intention-to-treat analyses. CHA-MW was as efficacious during the COVID-19 pandemic as more resource-intensive interventions that demanded greater time and effort from participants. Low completion rates for MBCT-R and iCBT during the COVID-19 pandemic may have contributed to these results.

PMID:42551008 | DOI:10.2196/88388

Categories
Nevin Manimala Statistics

Association of Enlarged Perivascular Spaces and Total Small Vessel Disease Burden With Kidney Function

Neurology. 2026 Aug 25;107(4):e218406. doi: 10.1212/WNL.0000000000218406. Epub 2026 Aug 4.

ABSTRACT

BACKGROUND AND OBJECTIVES: Enlarged perivascular spaces (EPVSs) in the basal ganglia (BG-EPVS) are an important marker of cerebral small vessel disease (cSVD), and EPVS in the centrum semiovale (CSO-EPVS) are part of the diagnostic criteria for cerebral amyloid angiopathy. We aimed to investigate associations of EPVS with reduced estimated glomerular filtration rate (eGFR) and glomerular hyperfiltration (higher than normal eGFR), which have scarcely been studied previously.

METHODS: In this cross-sectional study, we used pooled individual patient data from the Microbleeds International Collaborative Network which includes patients with ischemic stroke or transient ischemic attack. We investigated associations of impaired kidney function, defined as an eGFR of 30-60 or <30 mL/minute/1.73 m2, and glomerular hyperfiltration, defined as eGFR above the age-adjusted and sex-adjusted 95th centile, with BG-EPVS and CSO-EPVS severity. EPVS were rated according to a validated 5-point ordinal scale, and combined cSVD burden was rated using a validated 5-point ordinal scale with 1 point assigned for the presence of each of the following: severe white matter hyperintensities, ≥1 cerebral microbleed, ≥1 lacune, and BG-EPVS ≥11. Normal glomerular filtration was defined as eGFR ≥60 without hyperfiltration. We used multivariable ordinal logistic regression models to estimate risk of increased EPVS and cSVD burden severity adjusted for age, sex, and comorbidities.

RESULTS: Seven thousand two hundred fifty-four patients (mean age 71 ± 13 years, 43% female) were included in the analysis, 357 with glomerular hyperfiltration, 1,692 with eGFR 30-60, and 256 with eGFR <30. Compared with normal glomerular filtration, hyperfiltration was independently associated with BG-EPVS (adjusted odds ratio [aOR] 1.38, 95% CI 1.11-1.70, p < 0.001) and CSO-EPVS (aOR 1.34, 95% CI 1.08-1.64, p = 0.011). Associations of eGFR 30-60 and eGFR <30 with EPVS were not statistically significant. Compared with normal glomerular filtration, eGFR <30 (aOR 1.27, 95% CI 1.03-1.57) was independently associated with increased cSVD burden, but eGFR 30-60 (aOR 1.06, 95% CI 0.95-1.20) and hyperfiltration (aOR 1.15, 95% CI 0.98-1.34) were not.

DISCUSSION: Glomerular hyperfiltration was independently associated with EPVS severity, in both the basal ganglia and centrum semiovale. eGFR <30 was independently associated with total cSVD burden. A key limitation was a lack of repeated eGFR measurements.

PMID:42551001 | DOI:10.1212/WNL.0000000000218406

Categories
Nevin Manimala Statistics

Correctness, Harmfulness, and Diversity of Large Language Models for Colonoscopy Preparation Assistance: Comparative Evaluation Study

JMIR AI. 2026 Aug 4;5:e88581. doi: 10.2196/88581.

ABSTRACT

BACKGROUND: Colorectal cancer is a leading cause of cancer-related deaths in the United States, and colonoscopy remains the gold standard for early detection and prevention. However, many procedures are postponed due to inadequate bowel preparation, a preventable failure often caused by patients’ difficulty in understanding and following written prep instructions. Prior interventions such as reminder apps and instructional videos have improved adherence only modestly, largely because they cannot answer patient-specific questions. Recent advances in large language models (LLMs) raise the possibility of developing conversational assistants that can provide interactive support to patients in procedure preparation.

OBJECTIVE: This study evaluated the correctness, harmfulness, and diversity of synthetic dialogues generated by leading LLMs acting as both simulated AI Coaches and patients for colonoscopy preparation.

METHODS: Five leading LLMs-OpenAI’s o3, GPT-4.1, and GPT-5.1; Meta’s Llama 3.3 70B; and Mistral’s Large-2411-were used to generate 250 patient-AI Coach dialogues per model. Dialogues consisted of 3 to 7 question-answer pairs concerning diet, medications, and other prep-related topics. A multiprompt, multiquestion approach was designed to elicit diverse patient questions, and an error taxonomy was established to assess model capabilities in responding to questions. Human raters, including 3 medical experts, evaluated the generated questions for difficulty and the responses for correctness, error type, and potential harmfulness. Automatic evaluation using an LLM-as-a-judge approach complemented human evaluation. Question diversity was assessed using lexical diversity metrics (Distinct-1 and Distinct-2) and entropy. In addition, we evaluated a safety filtering mechanism in which responses judged incorrect by an automated evaluator were replaced with a deferral message instructing patients to contact their health care provider. Differences in response correctness across models were evaluated using permutation tests conducted at the dialogue level. Interrater agreement among human evaluators was assessed using the Gwet AC1 statistic. The study was conducted between May and September 2025.

RESULTS: Automatic evaluation results closely aligned with human judgments: leading models approached but did not achieve adequate performance. Closed-weight models (GPT-5.1, GPT-4.1, and o3) outperformed open-weight models (Llama and Mistral) on correctness, with the reasoning models (GPT-5.1 and o3) performing best. This turn-level ranking was preserved under the supplementary single-prompt baseline, although dialogue-level rankings differed. All models produced harmful errors, primarily due to omissions or misinterpretations of prep instructions. The multiprompt generation strategy substantially increased the diversity of patient questions compared with a single-prompt baseline. Applying an automated safety filter reduced overall error rates but failed to eliminate harmful responses.

CONCLUSIONS: Although LLMs demonstrate strong potential to support colonoscopy preparation, none are yet reliable enough for unsupervised deployment in patient-facing contexts. Persistent harmful errors and the limited effectiveness of simple filtering mechanisms highlight the need for improved instruction adherence, stronger safety mechanisms, and validation using real patient queries.

PMID:42550998 | DOI:10.2196/88581

Categories
Nevin Manimala Statistics

Correlates of Engagement and Associations With Outcomes in a Cannabis Harm-Reduction Mobile App for Youth With First-Episode Psychosis: Exploratory Analysis of the CHAMPS Pilot Randomized Controlled Trial

JMIR Form Res. 2026 Aug 4;10:e84836. doi: 10.2196/84836.

ABSTRACT

BACKGROUND: Continued cannabis use among young people with first-episode psychosis (FEP) has been linked to poorer clinical and functional outcomes (eg, increased symptom severity and higher relapse rates). Digital harm-reduction interventions may represent a promising, person-centered approach to reduce at-risk cannabis use behaviors in this population. However, evidence remains limited regarding which subgroups are more likely to engage with these interventions and whether specific levels of engagement are required to achieve more favorable cannabis-related outcomes.

OBJECTIVE: This exploratory, hypothesis-generating analysis of the CHAMPS (Cannabis Harm-Reducing App to Manage Practices Safely) pilot randomized controlled trial (RCT) evaluated a cannabis harm-reduction mobile app for youth with FEP in early intervention services (EIS). The objectives of this study were to assess engagement by examining associations between selected sociodemographic factors and module completion and to determine whether achieving specific completion thresholds was associated with improvements in cannabis-related outcomes.

METHODS: Cannabis-related outcomes (Marijuana Problems Scale [MPS], Protective Behavioral Strategies for Marijuana [PBSM] scores, and days of cannabis use) were self-assessed at baseline and at week 6 (primary end point), week 12, and week 18 (postrandomization). Participants were categorized into low to moderate (0-4 modules) and high (5-6 modules) completion groups, and selected sociodemographic factors were compared between groups using bivariate analyses. Mixed-effects models, adjusted for baseline values and the covariates sex and cannabis use disorder (CUD) status, were fitted to evaluate whether module thresholds were associated with improvements in cannabis-related outcomes.

RESULTS: Data from 96 participants, including 46 in the CHAMPS+EIS arm and 50 in the EIS-only arm (1:1 ratio), were analyzed under a modified intention-to-treat principle. Individuals in the high-engagement group reported higher baseline social support and higher educational level (P=.005 and P=.01). No statistically significant associations were observed between specific module completion thresholds and cannabis-related outcomes in adjusted mixed-effects models.

CONCLUSIONS: Participants with higher social support and higher education were more likely to engage with CHAMPS. Although descriptive analyses suggested a potential gradient of improvement with increasing module completion, no specific completion threshold was robustly associated with improvements in outcomes. Given the multifactorial nature of engagement, supporting subgroups at risk of lower app use and examining metrics beyond module completion may enhance the impact of CHAMPS. These findings are hypothesis-generating, given their exploratory nature, and require replication in a future efficacy trial.

PMID:42550986 | DOI:10.2196/84836

Categories
Nevin Manimala Statistics

Prediction of Blood Transfusion Need and Dose in Patients With Upper Gastrointestinal Bleeding: Retrospective Multicenter Prediction Model Study

JMIR Med Inform. 2026 Aug 4;14:e83889. doi: 10.2196/83889.

ABSTRACT

BACKGROUND: Transfusion thresholds in upper gastrointestinal bleeding are debated; hemoglobin cutoffs of 70-80 g/L are widely cited yet inconsistently applied. Common risk scores offer limited individualized guidance and rarely provide calibrated, interpretable predictions for transfusion decisions.

OBJECTIVE: This study aimed to develop and validate a two-stage, clinically constrained gradient-boosting framework (Medically Constrained Gradient Boosting [MCGB]) that predicts transfusion need and estimates transfusion dose with quantified uncertainty and to implement a prototype recommendation system for clinical use.

METHODS: We analyzed a retrospective multicenter cohort of 849 adults with endoscopically confirmed upper gastrointestinal bleeding admitted to 3 hospitals in Chongqing, China (January 2019 to August 2025). Predictors available before the transfusion decision included demographics, first recorded vital signs, initial laboratory indices, and clinician-adjudicated etiology. Stage 1 used a calibrated classifier with prespecified monotonic constraints and stability-screened, clinically justified interactions. Stage 2 modeled transfusion dose via quantile predictions with conformal adjustment to generate 95% prediction intervals. Performance was assessed using a cross-site hold-out design. Overall, 2 hospitals were used as the development cohort, within which stratified 5-fold cross-validation was performed for model development, hyperparameter tuning, interaction screening, and calibration. The remaining hospital was held out as an independent test cohort for final evaluation. Hospital-wise alternating external testing was further conducted as a supplementary robustness analysis to assess performance stability across institutions. Classification performance was evaluated using discrimination metrics (area under the receiver operating characteristic curve and area under the precision-recall curve), calibration metrics, and decision-curve analysis; regression performance was evaluated using R², mean absolute error, and prediction-interval coverage. A graphical user interface was implemented to enable clinicians to input patient data and obtain calibrated predictions of transfusion probability and corresponding dose recommendations.

RESULTS: MCGB achieved strong discrimination and good calibration across subgroups (area under the receiver operating characteristic curve=0.97 and area under the precision-recall curve=0.91). At a reference probability threshold of .50, sensitivity, specificity, and F1-scores were 0.99, 0.87, and 0.85, respectively, providing a representative operating point for comparison. For dose prediction among transfused patients, MCGB achieved R² of 0.95 and mean absolute error 0.04; 95% prediction-interval coverage was 0.94, indicating accurate point estimates with reliable uncertainty quantification. The software prototype further demonstrated feasibility of real-time decision support at the bedside.

CONCLUSIONS: MCGB provides calibrated, interpretable predictions of transfusion need and individualized dose in upper gastrointestinal bleeding and may support bedside decision-making and blood-bank planning, with a prototype interface demonstrating potential for clinical deployment. External validation in additional settings is warranted to confirm generalizability.

PMID:42550984 | DOI:10.2196/83889

Categories
Nevin Manimala Statistics

Large Language Model-Based Clinical Decision Support for Antibiotic Selection and Dose Recommendation in Hospitalized Patients With Pneumonia: Multicenter Retrospective Study

JMIR Med Inform. 2026 Aug 4;14:e98207. doi: 10.2196/98207.

ABSTRACT

BACKGROUND: Pneumonia is a common infectious disease, and antibiotic treatment in hospitalized patients must balance efficacy, safety, and resistance risk. However, antibiotic selection and dose adjustment still rely heavily on clinician experience. Although large language models (LLMs) are promising for clinical reasoning, their direct use for antibiotic selection and dose recommendation is limited by hallucinations and weak adherence to clinical constraints.

OBJECTIVE: This study aimed to develop and externally validate a constrained LLM-based clinical decision support pipeline for antibiotic selection and dose recommendation in hospitalized patients with pneumonia.

METHODS: We conducted a multicenter retrospective study using electronic health record narratives, antibiotic orders, and laboratory indicators of hepatic and renal function from 331 hospitalized patients with pneumonia from 2 hospitals in China. The development cohort included 233 patients, and the external validation cohort included 98 patients. The pipeline integrated dual-branch retrieval (similar-case vector retrieval plus guideline-based knowledge graph retrieval), clinician-defined rule constraints, and hybrid-context reasoning. DeepSeek-V3, GLM-4.6, and GPT-4o were evaluated using F1-score and Jaccard accuracy.

RESULTS: On the internal test set, the full pipeline using DeepSeek-V3 achieved the best performance, with an F1-score of 0.8110 (95% CI 0.7371-0.8762) and Jaccard accuracy of 0.7624 (95% CI 0.6810-0.8386) for antibiotic selection and an F1-score of 0.7538 (95% CI 0.6671-0.8329) and Jaccard accuracy of 0.7076 (95% CI 0.6145-0.7938) for joint antibiotic selection plus dosing recommendation. On the external validation set, performance remained high, with an F1-score of 0.8605 (95% CI 0.7891-0.9252) and Jaccard accuracy of 0.8571 (95% CI 0.7857-0.9184) for antibiotic selection, and an F1-score of 0.8503 (95% CI 0.7789-0.9150) and Jaccard accuracy of 0.8469 (95% CI 0.7755-0.9133) for antibiotic selection plus dosing recommendation. The system also provided traceable evidence and rule trigger information to support clinician review.

CONCLUSIONS: A constrained, retrieval-augmented LLM pipeline improved the consistency and interpretability of antibiotic selection and dose recommendation for hospitalized patients with pneumonia and provided preliminary evidence of cross-site generalizability.

PMID:42550965 | DOI:10.2196/98207