Categories
Nevin Manimala Statistics

Exploring the Design Space of Glanceable Smartwatch Feedback Displays: Experimental Study

JMIR Mhealth Uhealth. 2026 Jul 23;14:e81972. doi: 10.2196/81972.

ABSTRACT

BACKGROUND: Self-monitoring technologies are commonly used to promote health behavior change, with glanceable displays offering continuous feedback throughout the day. Yet, it is still unclear how various aspects of these glanceable representations affect their interpretability and usability.

OBJECTIVE: This study aimed to investigate the effects of 3 design factors-stylization, granularity, and salience-on users’ ability to understand glanceable smartwatch-based feedback on daily step goals.

METHODS: We conducted an online simulation study to examine how 3 design dimensions-stylization, salience, and granularity-influence the effectiveness of glanceable feedback displays. Stylization and salience were crossed in a 2×2 factorial design, while granularity varied from 1% to 20% progress increments. A total of 202 Amazon Mechanical Turk participants were randomly assigned to 1 of 16 smartwatch display conditions. In each condition, participants viewed feedback on daily step progress and estimated the level of progress shown. We measured estimation error and questionnaire-assessed perceived usability and acceptability. The collected data were analyzed using generalized estimating equations and linear regression.

RESULTS: High stylization reduced accuracy (+4.52 error points; P<.001) and negatively affected perceptions across 6 dimensions, including comprehension (P=.003), complexity (P<.001), and usability (P=.001). Granularity had a nonlinear effect: error was lowest around 5%-10%, with sharp increases at 20%. The 10% level also received the most favorable ratings, for example, comprehension (+0.656; P=.003). Salience had no effect. Previous smartwatch users were less accurate than never-users (+7.46 points) but rated displays as more useful (P=.002) and easier to focus on (P<.001). Current users gave similarly positive ratings on attention and usefulness.

CONCLUSIONS: These findings could help researchers design effective glanceable smartwatch feedback displays and expand the design space for glanceable feedback.

PMID:42492069 | DOI:10.2196/81972

Categories
Nevin Manimala Statistics

National Estimates of Human Flourishing and Associated Acts of Charity

JMIR Form Res. 2026 Jul 23;10:e90951. doi: 10.2196/90951.

ABSTRACT

Higher levels of human flourishing were moderately associated with helping strangers and volunteering time but showed little association with financial donations.

PMID:42492058 | DOI:10.2196/90951

Categories
Nevin Manimala Statistics

Performance of GPT-4o and Claude in Medical Ethics Scenarios: Comparative Study

JMIR Med Educ. 2026 Jul 23;12:e70199. doi: 10.2196/70199.

ABSTRACT

BACKGROUND: The emergence of AI technology has sparked curiosity regarding the capabilities of large language models (LLMs) in the field of medicine. Minimal research exists regarding the proficiency of various AI models in ethics scenarios, specifically in specialty-based scenarios.

OBJECTIVE: This study aimed to compare the performance of GPT-4o and Claude Sonnet 4 on ethics questions with that of medical students and orthopedic residents.

METHODS: A total of 200 ethical or legal scenario questions were randomly selected from question banks targeted for third- and fourth-year medical students (UWorld, AMBOSS) and orthopedic residents (OrthoBullets). Questions at the medical student level were exclusively text-based, while resident-level questions included text-based questions accompanied by images. Each question was entered identically into each AI model 3 separate times. If answers varied between trials, the answer provided most frequently by the model was used as the selected answer.

RESULTS: GPT-4o correctly answered 140 (70%) of 200 questions, which was similar to the average human test taker score of 71% (~142/200 questions). Claude correctly answered 180 (89%) questions, a score greater than that of human test takers and significantly better than GPT-4o (P<.001). Claude scored significantly higher than GPT-4o in almost all question categories. GPT-4o provided different responses to identically worded trials for 27 (21%) of 130 general questions and 3 (4%) of 70 orthopedic questions (P=.002), while Claude did not have a significant difference in variability between these 2 groups (general: 16/130, 12% vs orthopedic: 3/70, 4%; P=.06). GPT-4o selected the incorrect response for 60 (30%) total questions and chose the incorrect response most commonly selected by humans significantly more frequently on UWorld interpersonal-specific questions (30/40, 75%) than on UWorld all social sciences (27/40, 68%; P=.03). Claude showed no significant difference in the rate of most common incorrect response selection between question categories.

CONCLUSIONS: These results suggest that GPT-4o can potentially answer both general and specialty-specific ethical questions with similar proficiency to sample groups of both medical students and orthopedic residents, while Claude AI performs significantly better than both humans and GPT-4o. Variables such as AI model framework and training data may drive the observed difference in performance, but the exact cause cannot be definitively isolated without intentional testing. Therefore, further research is needed to ensure safety by minimizing output variability before integrating AI as a patient-facing resource.

PMID:42492056 | DOI:10.2196/70199

Categories
Nevin Manimala Statistics

Strategies for Deploying Large Language Models for Ascertaining Clinical Outcomes and Sites of Metastases From Radiology Impressions in Patients With Cancer

JCO Clin Cancer Inform. 2026 Jul-Sep;10(3):e2500164. doi: 10.1200/CCI-25-00164. Epub 2026 Jul 23.

ABSTRACT

PURPOSE: To evaluate open-source large language models (LLMs) for extracting cancer-specific phenotypic data, benchmark their performance against GPT4 models, and assess the impact of fine-tuning with training data sizes.

METHODS: Open-source LLMs (Mistral, LLaMa, MAMBA, BioMistral) were evaluated in zero-/one-shot and fine-tuned setups against GPT4-turbo/GPT4o to extract the cancer presence, progression, response, and metastatic sites from radiology impressions of patients with solid tumors treated at Dana-Farber Cancer Institute. Performance metrics (accuracy, precision, recall, F1-score) were computed. McNemar’s odds ratio (OR), measuring which model is more likely to be correct when they disagree, was computed with 95% CI. Statistical significance was assessed using the alpha of .000139.

RESULTS: This study included 2,623 patients (25,273 radiology impressions). In zero-/one-shot settings, GPT4-turbo/GPT4o outperformed open-source LLMs. However, fine-tuned open-source LLMs achieved higher F1-scores than GPT4 models. Compared with the best-performing GPT4 model, fine-tuned Mistral0.2-7.3B (OR, 0.27 [95% CI, 0.20 to 0.36]; P < .00001), Mistral0.3-7.3B (OR, 0.26 [95% CI, 0.19 to 0.36]; P < .00001), LLaMa2-6.7B (OR, 0.30 [95% CI, 0.22 to 0.40]; P < .00001), LLaMa3.1-8B (OR, 0.37 [95% CI, 0.28 to 0.48]; P < .00001), and MAMBA-2.8B (OR, 0.32 [95% CI, 0.24 to 0.42]; P < .00001) showed significantly better performance in ascertaining disease progression. Performance was consistently better for inferring overall response, any evidence of cancer, and sites of metastases, with no significant differences among fine-tuned open-source LLMs. Fine-tuning gains plateaued at 25% of training data (5,718 impressions) and remained comparable at 5% (1,144 impressions).

CONCLUSION: Open-source LLMs, when fine-tuned using labeled data, can effectively automate the ascertainment of key radiophenotypic variables using only the impression section of radiology reports, without the full report text. Their consistent performance in small training sets suggests that these models may provide a scalable approach for phenotypic characterization of patients with cancer in real-world clinical settings.

PMID:42492046 | DOI:10.1200/CCI-25-00164

Categories
Nevin Manimala Statistics

Efficacy Measures Used in Alzheimer Disease Clinical Trials Between 2015 and 2025: A Systematic Review

Neurology. 2026 Aug 25;107(4):e218373. doi: 10.1212/WNL.0000000000218373. Epub 2026 Jul 23.

ABSTRACT

BACKGROUND AND OBJECTIVES: Regulatory guidance has long emphasized clinically meaningful outcomes in Alzheimer disease (AD) drug trials. No contemporary review has examined the efficacy measures used for clinical trials of AD therapies. The objective of this study was to evaluate the clinical relevance and heterogeneity of efficacy measures used in phase II-IV AD drug trials over the past decade.

METHODS: We systematically searched PubMed, Embase, the Cochrane Central Register of Controlled Trials, and ClinicalTrials.gov. All phase II-IV clinical trials of pharmacologic therapies in individuals with mild cognitive impairment due to AD or mild-to-moderate AD that were published, completed, or terminated between January 1, 2015, and November 15, 2025, or still ongoing as of November 15, 2025, were eligible for inclusion. We summarized the proportion of trials using clinical outcomes or biomarkers as primary or secondary efficacy measures, key clinical domains addressed, and the number of distinct measures using descriptive statistics. This study is registered with PROSPERO (CRD420251032087).

RESULTS: Among 238 included trials, 95.4% (227/238) used at least one clinical outcome and 73.5% (175/238) used at least one biomarker as efficacy measures. Cognitive abilities (87.0%, 207/238), global status (76.9%, 183/238), and functional ability (71.4%, 170/238) were the most frequently measured clinical domains, whereas patient (16.0%, 38/238) and caregiver (8.8%, 21/238) quality of life and significant disease-related life events (2.5%, 6/238) were less frequent. Only 21.8% (52/238) of trials adopted regulatory-recommended co-primary measures of cognition with either function or global status, whereas 8.4% (20/238) used biomarkers as the sole primary efficacy measure. We identified 318 distinct clinical outcome measures and 219 distinct biomarkers used across all trials, 6.9% (22/318) and 6.8% (15/219) of which were used in more than 5% of trials.

DISCUSSION: Efficacy measures in AD drug trials primarily focused on cognition, global status, and functional outcomes, which are important indicators of AD progression. However, other clinical domains that may also be meaningful to patients and caregivers were far less frequently assessed, while biomarker use is widespread. Substantial heterogeneity in efficacy measure use limits comparability across trials and highlights the need for standardized, consensus-based, clinically meaningful core measure sets for AD drug trials.

PMID:42492033 | DOI:10.1212/WNL.0000000000218373

Categories
Nevin Manimala Statistics

Revised Thyroid Differentiation Score for Risk Stratification in Papillary Thyroid Carcinoma

JCO Precis Oncol. 2026 Jul;10(7):e2600207. doi: 10.1200/PO-26-00207. Epub 2026 Jul 23.

ABSTRACT

PURPOSE: Loss of thyroid differentiation underlies the aggressive behavior of a subset of papillary thyroid carcinomas (PTCs). The Thyroid Differentiation Score (TDS), introduced by The Cancer Genome Atlas (TCGA), quantifies tumor differentiation but has limited reproducibility and clinical applicability. We developed and validated a revised score (xTDS) that enables reproducible assessment of tumor differentiation with prognostic relevance across cohorts.

METHODS: We analyzed RNA sequencing data from 570 patients with PTC across three independent cohorts: a discovery cohort from the MD Anderson Cancer Center (n = 111) and two external validation cohorts from Vanderbilt University (n = 69) and TCGA (n = 390). xTDS was evaluated for correlation with the original TDS, association with oncogenic drivers, and prognostic value for disease-specific survival (DSS) and progression-free survival (PFS).

RESULTS: xTDS showed near-perfect correlation with the original TDS (Spearman ρ = 0.98) and recapitulated known biological patterns, with BRAF V600E-driven tumors exhibiting the lowest differentiation scores and RAS-driven tumors the highest across all cohorts (P < .001). Low xTDS was associated with worse DSS in the MD Anderson (P < .001) and Vanderbilt (P = .005) cohorts and showed a similar numerical trend in TCGA, without statistical significance because of limited events. Low xTDS was consistently associated with worse PFS across all three cohorts (P = .048, <0.001, and 0.007, respectively).

CONCLUSION: xTDS is a reproducible measure of thyroid differentiation that preserves the biological and prognostic relevance of the original TDS while overcoming technical constraints affecting reproducibility.

PMID:42492028 | DOI:10.1200/PO-26-00207

Categories
Nevin Manimala Statistics

Seasonal Variation in Subscapularis Tendon Strain in Baseball Pitchers Assessed by Ultrasonography-based Analysis

J Med Ultrasound. 2025 Nov 8;34(2):186-193. doi: 10.4103/jmu.JMU-D-25-00056. eCollection 2026 Apr-Jun.

ABSTRACT

BACKGROUND: Repetitive loading injuries of the rotator cuff tendons are common in baseball pitchers. Ultrasonography-based strain analysis shows promise for in vivo monitoring of cumulative tendon loading but faces technical challenges in the rotator cuff. We developed a novel ultrasonography-based strain analysis procedure to test the hypothesis that regular-season pitching increases the peak strain of the subscapularis tendon.

METHODS: Ultrasound image sequences of subscapularis tendon were obtained from the dominant shoulders of nine pitchers during isotonic motion in the preseason and regular season. Tendon displacement during the eccentric phase was estimated using optical flow tracking, validated against NCORR, a reliable software for quantifying material deformation.

RESULTS: Peak strain was significantly higher in the regular season (18.7 ± 3.1%) compared to the preseason (8.8 ± 2.7%), suggesting an elevated risk of tendon microdamage.

CONCLUSION: Our findings indicate that intense pitching during the season may elevate tendon strain, highlighting the potential of strain analysis for monitoring tendon loading-related changes. The proposed procedure is robust, reliable, and easy to implement, offering a practical tool for tracking in-season tendon loading in pitchers.

PMID:42492005 | PMC:PMC13379176 | DOI:10.4103/jmu.JMU-D-25-00056

Categories
Nevin Manimala Statistics

An Online Tool for Correcting Performance Measures of Electronic Phenotyping Algorithms for Verification Bias

ACI open. 2024 Jul;8(2):e89-e93. doi: 10.1055/a-2402-5937. Epub 2024 Dec 27.

ABSTRACT

OBJECTIVES: Computable or electronic phenotypes of patient conditions are becoming more commonplace in quality improvement and clinical research. During phenotyping algorithm validation, standard classification performance measures (i.e., sensitivity, specificity, positive predictive value, negative predictive value, and accuracy) are often employed. When validation is performed on a randomly sampled patient population, direct estimates of these measures are valid. However, studies will commonly sample patients conditional on the algorithm result prior to validation, leading to a form of bias known as verification bias.

METHODS: We illustrate validation study sampling design and naïve and bias-corrected validation performance through both a concrete example (1,000 cases, 100 noncases, 1:1 sampling on predicted status) and a more thorough simulation study under varied realistic scenarios. We additionally describe the development of a free web calculator to adjust estimates for people validating phenotyping algorithms.

RESULTS: In our illustrative example, naïve performance estimates corresponded to 0.942 sensitivity, 0.979 specificity, and 0.960 accuracy; these contrast proper estimates of 0.620 sensitivity, 0.999 specificity, and 0.944 accuracy after adjusting for verification bias using our free calculator. Our simulation results demonstrate increasing positive bias for sensitivity and negative bias for specificity as the disease prevalence approaches zero, with decreasing positive predictive value moderately exacerbating these biases.

CONCLUSION: Novel computable phenotypes of patient conditions must account for verification bias when calculating performance measures of the algorithm. The performance measures may vary significantly based on disease prevalence in the source population so use of a free web calculator to adjust these measures is desirable.

PMID:42491998 | PMC:PMC13379156 | DOI:10.1055/a-2402-5937

Categories
Nevin Manimala Statistics

Breast Ultrasound Texture Features of Family History of Breast Cancer and Contraceptive Use in Nigerian Younger Women

J Med Ultrasound. 2026 May 16;34(2):201-206. doi: 10.4103/jmu.JMU-D-25-00004. eCollection 2026 Apr-Jun.

ABSTRACT

BACKGROUND: The impact of family history of breast cancer (FHBC) on the texture features of the breast is poorly understood. We sought to examine the effect of FHBC on the image texture features of breast sonograms of young asymptomatic women.

METHODS: This study estimated texture features in breast sonograms with ImageJ software in a cohort of 360 young asymptomatic women aged 19-29 years. The sonographic assessment was performed using a digital color Doppler ultrasonic diagnostic instruments (Xuzhou Kaixin Electronic Instrument: Model – DCU7; SN-2331531). Texture analysis of each region of interest was performed using the gray-level co-occurrence matrix plugin of the ImageJ software. An independent samples t-test was used to compare the texture features of the breast in women with FHBC and the control.

RESULTS: Tissue homogeneity and uniformity were significantly lower in women with FHBC (P < 0.05), while tissue complexity (entropy) was significantly higher in subjects with FHBC (P = 0.02). Correlation and contrast were statistically the same between these groups contraceptive use showed no statistical association with breast image texture features (P > 0.05).

CONCLUSION: In this study, breast tissue complexity, homogeneity, and uniformity were affected by FHBC. Breast texture features should be considered when assessing risk in women with FHBC.

PMID:42491985 | PMC:PMC13379167 | DOI:10.4103/jmu.JMU-D-25-00004

Categories
Nevin Manimala Statistics

Antimicrobial Resistance Knowledge, Training Exposure, and Prescribing Influences Among Community Health Workers in Kaduna State, Nigeria

Cureus. 2026 Jun 17;18(6):e111045. doi: 10.7759/cureus.111045. eCollection 2026 Jun.

ABSTRACT

INTRODUCTION: Antimicrobial resistance (AMR) is a major public health threat, particularly in low- and middle-income countries where inappropriate antibiotic use remains common. Community Health Workers (CHWs) are important providers of primary healthcare (PHC) in Nigeria, yet evidence on their AMR knowledge, stewardship training, and prescribing influences remains limited. This study assessed AMR/AMS knowledge, training exposure, prescribing attitudes, and factors influencing antimicrobial prescribing among CHWs in Kaduna State, Nigeria.

METHODS: A descriptive cross-sectional study was conducted among 286 CHWs selected from 27 government-owned PHC facilities in Kaduna State using multistage sampling. Data were collected between August and September 2025 using a structured self-administered questionnaire. Data were analyzed using IBM Statistical Package for the Social Sciences version 25.0 (IBM Corp., Armonk, NY). Descriptive statistics were presented as frequencies and percentages, while chi-square tests assessed associations between AMR/AMS training, knowledge level, and attitude level.

RESULTS: Most respondents had heard of AMR (87.4%), correctly identified bacteria as the target of antibiotics (90.9%), and recognized that inappropriate antibiotic use contributes to resistance (92.7%). Only 20.6% had received formal AMR/AMS training. Although many respondents reported guideline use (68.6%) and prescription documentation (70.3%), 62.3% prescribed antibiotics based on experience without diagnostic tools, while 38.1% reported patient demand-influenced prescribing decisions. Formal AMR/AMS training was significantly associated with both knowledge level and attitude level.

CONCLUSION: Despite generally good AMR knowledge, formal stewardship training among CHWs was limited. Strengthening AMR/AMS training, diagnostic support, supervision, and guideline-based prescribing may improve antimicrobial stewardship at the PHC level.

PMID:42491981 | PMC:PMC13378881 | DOI:10.7759/cureus.111045