Categories
Nevin Manimala Statistics

Hidden burden and economic vulnerability of Narcolepsy Patients – Accessibility as the major challenge for a new therapeutic era

Sleep Med. 2026 Jul 17;147:109147. doi: 10.1016/j.sleep.2026.109147. Online ahead of print.

NO ABSTRACT

PMID:42492132 | DOI:10.1016/j.sleep.2026.109147

Categories
Nevin Manimala Statistics

Evaluating the Google CT Foundation model for central pulmonary embolism detection on computed tomography pulmonary angiograms

Eur J Radiol. 2026 Jul 20;204:113092. doi: 10.1016/j.ejrad.2026.113092. Online ahead of print.

ABSTRACT

Early detection of pulmonary embolism (PE) is critical for clinical outcomes, yet existing deep-learning models often fail to generalize across institutions. The recently released Google CT Foundation model, pre-trained on a large, diverse CT corpus, produces compact volumetric embeddings that may transfer to downstream tasks without fine-tuning. We evaluate a classification pipeline that pairs these frozen embeddings with three lightweight heads – Multi-Layer Perceptron (MLP), Random Forest (RF), and a stacking ensemble – for central PE detection, training on the public RSNA CTPA dataset and externally validating on the Stanford INSPECT cohort. The pipeline first reproduces the published data-size scaling behavior of the foundation model on other medical pathologies. The highest point-estimate test AUC of 0.79 was obtained by the RF trained on a balanced training set of 408 CT studies and evaluated on the held-out RSNA test set of 144 studies; paired DeLong tests showed that this advantage over the MLP and stacking heads was not statistically significant. Direct comparison to specialized state-of-the-art PE pipelines is task-mismatched – those target any-PE rather than central PE – so the result anchors a non-fine-tuned baseline rather than a deficit. In a low-data regime, hard vascular segmentation did not improve performance. On the external INSPECT cohort, AUC dropped by 0.17 and specificity at the transferred operating point collapsed, so simple global threshold recalibration does not restore deployability-frozen generalist embeddings alone do not guarantee cross-institutional reliability.

PMID:42492117 | DOI:10.1016/j.ejrad.2026.113092

Categories
Nevin Manimala Statistics

A Digital Acceptance and Commitment Therapy and Education Intervention for Caregivers of Very Preterm Infants in the Neonatal Intensive Care Unit: Randomized Controlled Trial

JMIR Ment Health. 2026 Jul 23;13:e92021. doi: 10.2196/92021.

ABSTRACT

BACKGROUND: Parents of very preterm infants admitted to the neonatal intensive care unit (NICU) experience high levels of psychological distress, yet access to timely, evidence-based mental health support is limited by staffing and resource constraints. Digital mental health interventions offer a scalable approach to addressing this gap; however, their effectiveness has not been well established in NICU caregiver populations, particularly during periods of acute stress.

OBJECTIVE: This study aims to evaluate the effectiveness of a self-guided digital acceptance and commitment therapy (ACT)-based intervention combined with NICU-specific education (NICU parent acceptance and commitment therapy [NPACT]). The study explored the intervention’s effects on stress among parents and primary caregivers of very preterm infants, compared to a digital education-only intervention, and active control.

METHODS: We conducted a 3-arm, single-center, randomized controlled cluster trial in a tertiary NICU. Parents and primary caregivers of very preterm infants (<32 wk’ gestational age,<1 wk old) were randomized by family cluster to (1) NPACT (ACT+ education), (2) a digital education-only intervention, or (3) active control. Digital interventions were delivered via a web-based platform over 2 weeks. The primary outcome was NICU-related stress on the Parent Stressor Scale: Neonatal Intensive Care Unit (PSS:NICU) at 2 weeks postrandomization. Secondary outcomes included caregiver anxiety, depression, perceived stress, and selected neonatal outcomes. Engagement and perceived helpfulness were assessed for digital interventions.

RESULTS: A total of 102 caregivers from 68 family clusters (79 infants; mean gestational age 28.1, SD 2.2 wk) were enrolled. There were no statistically significant between-group differences in the mean PSS:NICU scores at 2 weeks (NPACT 3.0, SD 0.9; education-only 2.5, SD 1; active control 2.6, SD 0.9; adjusted mean difference for NPACT vs active control 0.04, 95% CI -0.39 to 0.47). No between-group differences were observed for secondary psychological outcomes at any time point. However, caregivers in both digital intervention groups had higher odds of full breastfeeding at discharge compared with active control. Engagement with the digital interventions was high, with 97% (28/29) of NPACT participants and 76% (19/25) of education-only participants completing at least 5 of 7 modules, and both interventions were rated as very helpful.

CONCLUSIONS: In this trial, an unguided digital mental health intervention delivered during NICU admission did not reduce NICU-specific parental stress or other psychological outcomes relative to active control. However, the intervention was highly used by caregivers. These findings suggest that while a brief digital mental health intervention can be successfully implemented in a high-stress clinical setting with caregivers, its capacity to reduce acute psychological distress may be limited. Secondary findings indicate potential benefits of the digital intervention on breastfeeding, generating hypotheses for future research. Digital mental health interventions in neonatal settings may be most effective when integrated within hybrid models of care and/or delivered beyond the acute admission phase.

PMID:42492071 | DOI:10.2196/92021

Categories
Nevin Manimala Statistics

Investigation of Vaccine-Induced Immune Response After Vaccination Against Respiratory Viruses in Immunocompromised Patients With or Without Hemato-Oncological Diseases (RESPONSE): Protocol for a Prospective Cohort Study

JMIR Res Protoc. 2026 Jul 23;15:e88520. doi: 10.2196/88520.

ABSTRACT

BACKGROUND: Infections with respiratory viruses such as SARS-CoV-2 and influenza are significant international public health concerns. While patients with cancer remain the most vulnerable group, they show poor vaccine response in general. Immunological data in this population are limited and mainly focus on serological parameters. However, in these patients, cellular, and especially T-cell, responses often seem to be induced more reliably than humoral responses.

OBJECTIVE: To gain further insights into vaccine-induced immunity, the RESPONSE study will analyze the effect of early and late booster vaccination on humoral and cellular responses, with special focus on T cell-induced immune responses. In addition, we aim to investigate factors influencing humoral and cellular vaccine-induced immunity in patients with hematological and oncological malignancies, including state of disease, treatment, and demographic factors.

METHODS: Humoral immune responses will be assessed by measuring binding and neutralizing antibodies using standardized assays. Cellular immunity will be evaluated using functional assays such as flow cytometry and FluoroSpot, as well as in-depth analyses using additional exploratory assays as appropriate. Immune responses will be correlated with clinical parameters, including disease status, treatment, and demographic factors.

RESULTS: This study was initiated following ethics approval and is currently recruiting participants. Enrollment commenced on March 25, 2025, and is ongoing, whereas biosample collection and follow-up visits are nearing completion for most participants. Final data cleaning, dataset integration, and statistical analyses of adaptive immune responses are planned from the third quarter of 2026 onward.

CONCLUSIONS: This study intends to lay a foundation for a structured translational research platform on vaccination to aim for best protection from infection by different respiratory pathogens. Long-term objectives are reaching best possible protection from vaccine-preventable disease with a first focus on influenza infection. In addition, we plan to investigate vaccine-induced immune responses to the recently approved respiratory syncytial virus vaccine using this platform and possibly extend this to further vaccines in the future. Urgent questions, such as the influence of different targeted therapies on vaccine immune response, will be part of these projects.

TRIAL REGISTRATION: ClinicalTrials.gov NCT06612515; https://clinicaltrials.gov/study/NCT06612515.

INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/88520.

PMID:42492070 | DOI:10.2196/88520

Categories
Nevin Manimala Statistics

Exploring the Design Space of Glanceable Smartwatch Feedback Displays: Experimental Study

JMIR Mhealth Uhealth. 2026 Jul 23;14:e81972. doi: 10.2196/81972.

ABSTRACT

BACKGROUND: Self-monitoring technologies are commonly used to promote health behavior change, with glanceable displays offering continuous feedback throughout the day. Yet, it is still unclear how various aspects of these glanceable representations affect their interpretability and usability.

OBJECTIVE: This study aimed to investigate the effects of 3 design factors-stylization, granularity, and salience-on users’ ability to understand glanceable smartwatch-based feedback on daily step goals.

METHODS: We conducted an online simulation study to examine how 3 design dimensions-stylization, salience, and granularity-influence the effectiveness of glanceable feedback displays. Stylization and salience were crossed in a 2×2 factorial design, while granularity varied from 1% to 20% progress increments. A total of 202 Amazon Mechanical Turk participants were randomly assigned to 1 of 16 smartwatch display conditions. In each condition, participants viewed feedback on daily step progress and estimated the level of progress shown. We measured estimation error and questionnaire-assessed perceived usability and acceptability. The collected data were analyzed using generalized estimating equations and linear regression.

RESULTS: High stylization reduced accuracy (+4.52 error points; P<.001) and negatively affected perceptions across 6 dimensions, including comprehension (P=.003), complexity (P<.001), and usability (P=.001). Granularity had a nonlinear effect: error was lowest around 5%-10%, with sharp increases at 20%. The 10% level also received the most favorable ratings, for example, comprehension (+0.656; P=.003). Salience had no effect. Previous smartwatch users were less accurate than never-users (+7.46 points) but rated displays as more useful (P=.002) and easier to focus on (P<.001). Current users gave similarly positive ratings on attention and usefulness.

CONCLUSIONS: These findings could help researchers design effective glanceable smartwatch feedback displays and expand the design space for glanceable feedback.

PMID:42492069 | DOI:10.2196/81972

Categories
Nevin Manimala Statistics

National Estimates of Human Flourishing and Associated Acts of Charity

JMIR Form Res. 2026 Jul 23;10:e90951. doi: 10.2196/90951.

ABSTRACT

Higher levels of human flourishing were moderately associated with helping strangers and volunteering time but showed little association with financial donations.

PMID:42492058 | DOI:10.2196/90951

Categories
Nevin Manimala Statistics

Performance of GPT-4o and Claude in Medical Ethics Scenarios: Comparative Study

JMIR Med Educ. 2026 Jul 23;12:e70199. doi: 10.2196/70199.

ABSTRACT

BACKGROUND: The emergence of AI technology has sparked curiosity regarding the capabilities of large language models (LLMs) in the field of medicine. Minimal research exists regarding the proficiency of various AI models in ethics scenarios, specifically in specialty-based scenarios.

OBJECTIVE: This study aimed to compare the performance of GPT-4o and Claude Sonnet 4 on ethics questions with that of medical students and orthopedic residents.

METHODS: A total of 200 ethical or legal scenario questions were randomly selected from question banks targeted for third- and fourth-year medical students (UWorld, AMBOSS) and orthopedic residents (OrthoBullets). Questions at the medical student level were exclusively text-based, while resident-level questions included text-based questions accompanied by images. Each question was entered identically into each AI model 3 separate times. If answers varied between trials, the answer provided most frequently by the model was used as the selected answer.

RESULTS: GPT-4o correctly answered 140 (70%) of 200 questions, which was similar to the average human test taker score of 71% (~142/200 questions). Claude correctly answered 180 (89%) questions, a score greater than that of human test takers and significantly better than GPT-4o (P<.001). Claude scored significantly higher than GPT-4o in almost all question categories. GPT-4o provided different responses to identically worded trials for 27 (21%) of 130 general questions and 3 (4%) of 70 orthopedic questions (P=.002), while Claude did not have a significant difference in variability between these 2 groups (general: 16/130, 12% vs orthopedic: 3/70, 4%; P=.06). GPT-4o selected the incorrect response for 60 (30%) total questions and chose the incorrect response most commonly selected by humans significantly more frequently on UWorld interpersonal-specific questions (30/40, 75%) than on UWorld all social sciences (27/40, 68%; P=.03). Claude showed no significant difference in the rate of most common incorrect response selection between question categories.

CONCLUSIONS: These results suggest that GPT-4o can potentially answer both general and specialty-specific ethical questions with similar proficiency to sample groups of both medical students and orthopedic residents, while Claude AI performs significantly better than both humans and GPT-4o. Variables such as AI model framework and training data may drive the observed difference in performance, but the exact cause cannot be definitively isolated without intentional testing. Therefore, further research is needed to ensure safety by minimizing output variability before integrating AI as a patient-facing resource.

PMID:42492056 | DOI:10.2196/70199

Categories
Nevin Manimala Statistics

Strategies for Deploying Large Language Models for Ascertaining Clinical Outcomes and Sites of Metastases From Radiology Impressions in Patients With Cancer

JCO Clin Cancer Inform. 2026 Jul-Sep;10(3):e2500164. doi: 10.1200/CCI-25-00164. Epub 2026 Jul 23.

ABSTRACT

PURPOSE: To evaluate open-source large language models (LLMs) for extracting cancer-specific phenotypic data, benchmark their performance against GPT4 models, and assess the impact of fine-tuning with training data sizes.

METHODS: Open-source LLMs (Mistral, LLaMa, MAMBA, BioMistral) were evaluated in zero-/one-shot and fine-tuned setups against GPT4-turbo/GPT4o to extract the cancer presence, progression, response, and metastatic sites from radiology impressions of patients with solid tumors treated at Dana-Farber Cancer Institute. Performance metrics (accuracy, precision, recall, F1-score) were computed. McNemar’s odds ratio (OR), measuring which model is more likely to be correct when they disagree, was computed with 95% CI. Statistical significance was assessed using the alpha of .000139.

RESULTS: This study included 2,623 patients (25,273 radiology impressions). In zero-/one-shot settings, GPT4-turbo/GPT4o outperformed open-source LLMs. However, fine-tuned open-source LLMs achieved higher F1-scores than GPT4 models. Compared with the best-performing GPT4 model, fine-tuned Mistral0.2-7.3B (OR, 0.27 [95% CI, 0.20 to 0.36]; P < .00001), Mistral0.3-7.3B (OR, 0.26 [95% CI, 0.19 to 0.36]; P < .00001), LLaMa2-6.7B (OR, 0.30 [95% CI, 0.22 to 0.40]; P < .00001), LLaMa3.1-8B (OR, 0.37 [95% CI, 0.28 to 0.48]; P < .00001), and MAMBA-2.8B (OR, 0.32 [95% CI, 0.24 to 0.42]; P < .00001) showed significantly better performance in ascertaining disease progression. Performance was consistently better for inferring overall response, any evidence of cancer, and sites of metastases, with no significant differences among fine-tuned open-source LLMs. Fine-tuning gains plateaued at 25% of training data (5,718 impressions) and remained comparable at 5% (1,144 impressions).

CONCLUSION: Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text. Their consistent performance in small training sets suggests that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.

PMID:42492046 | DOI:10.1200/CCI-25-00164

Categories
Nevin Manimala Statistics

Efficacy Measures Used in Alzheimer Disease Clinical Trials Between 2015 and 2025: A Systematic Review

Neurology. 2026 Aug 25;107(4):e218373. doi: 10.1212/WNL.0000000000218373. Epub 2026 Jul 23.

ABSTRACT

BACKGROUND AND OBJECTIVES: Regulatory guidance has long emphasized clinically meaningful outcomes in Alzheimer disease (AD) drug trials. No contemporary review has examined the efficacy measures used for clinical trials of AD therapies. The objective of this study was to evaluate the clinical relevance and heterogeneity of efficacy measures used in phase II-IV AD drug trials over the past decade.

METHODS: We systematically searched PubMed, Embase, the Cochrane Central Register of Controlled Trials, and ClinicalTrials.gov. All phase II-IV clinical trials of pharmacologic therapies in individuals with mild cognitive impairment due to AD or mild-to-moderate AD that were published, completed, or terminated between January 1, 2015, and November 15, 2025, or still ongoing as of November 15, 2025, were eligible for inclusion. We summarized the proportion of trials using clinical outcomes or biomarkers as primary or secondary efficacy measures, key clinical domains addressed, and the number of distinct measures using descriptive statistics. This study is registered with PROSPERO (CRD420251032087).

RESULTS: Among 238 included trials, 95.4% (227/238) used at least one clinical outcome and 73.5% (175/238) used at least one biomarker as efficacy measures. Cognitive abilities (87.0%, 207/238), global status (76.9%, 183/238), and functional ability (71.4%, 170/238) were the most frequently measured clinical domains, whereas patient (16.0%, 38/238) and caregiver (8.8%, 21/238) quality of life and significant disease-related life events (2.5%, 6/238) were less frequent. Only 21.8% (52/238) of trials adopted regulatory-recommended co-primary measures of cognition with either function or global status, whereas 8.4% (20/238) used biomarkers as the sole primary efficacy measure. We identified 318 distinct clinical outcome measures and 219 distinct biomarkers used across all trials, 6.9% (22/318) and 6.8% (15/219) of which were used in more than 5% of trials.

DISCUSSION: Efficacy measures in AD drug trials primarily focused on cognition, global status, and functional outcomes, which are important indicators of AD progression. However, other clinical domains that may also be meaningful to patients and caregivers were far less frequently assessed, while biomarker use is widespread. Substantial heterogeneity in efficacy measure use limits comparability across trials and highlights the need for standardized, consensus-based, clinically meaningful core measure sets for AD drug trials.

PMID:42492033 | DOI:10.1212/WNL.0000000000218373

Categories
Nevin Manimala Statistics

Revised Thyroid Differentiation Score for Risk Stratification in Papillary Thyroid Carcinoma

JCO Precis Oncol. 2026 Jul;10(7):e2600207. doi: 10.1200/PO-26-00207. Epub 2026 Jul 23.

ABSTRACT

PURPOSE: Loss of thyroid differentiation underlies the aggressive behavior of a subset of papillary thyroid carcinomas (PTCs). The Thyroid Differentiation Score (TDS), introduced by The Cancer Genome Atlas (TCGA), quantifies tumor differentiation but has limited reproducibility and clinical applicability. We developed and validated a revised score (xTDS) that enables reproducible assessment of tumor differentiation with prognostic relevance across cohorts.

METHODS: We analyzed RNA sequencing data from 570 patients with PTC across three independent cohorts: a discovery cohort from the MD Anderson Cancer Center (n = 111) and two external validation cohorts from Vanderbilt University (n = 69) and TCGA (n = 390). xTDS was evaluated for correlation with the original TDS, association with oncogenic drivers, and prognostic value for disease-specific survival (DSS) and progression-free survival (PFS).

RESULTS: xTDS showed near-perfect correlation with the original TDS (Spearman ρ = 0.98) and recapitulated known biological patterns, with BRAF V600E-driven tumors exhibiting the lowest differentiation scores and RAS-driven tumors the highest across all cohorts (P < .001). Low xTDS was associated with worse DSS in the MD Anderson (P < .001) and Vanderbilt (P = .005) cohorts and showed a similar numerical trend in TCGA, without statistical significance because of limited events. Low xTDS was consistently associated with worse PFS across all three cohorts (P = .048, <0.001, and 0.007, respectively).

CONCLUSION: xTDS is a reproducible measure of thyroid differentiation that preserves the biological and prognostic relevance of the original TDS while overcoming technical constraints affecting reproducibility.

PMID:42492028 | DOI:10.1200/PO-26-00207