article · BMC Sports Science Medicine and Rehabilitation
Adapted physical activity (APA) is a cornerstone of the management of musculoskeletal and rheumatological conditions, yet its individualized prescription remains challenging in sub-Saharan Africa due to the shortage of specialized rehabilitation professionals. Large language models (LLMs) may represent a promising decision-support tool, but their performance for APA prescription across musculoskeletal and rheumatological conditions has never been comparatively evaluated in a low-resource African context. This study aimed to assess and compare the quality of APA prescriptions generated by three contemporary LLMs for musculoskeletal and rheumatological clinical scenarios. This cross-sectional comparative study was conducted from October 15, 2025, to January 16, 2026, at the Department of Rheumatology, Bogodogo University Hospital, Ouagadougou, Burkina Faso. Sixteen standardized musculoskeletal and rheumatological clinical scenarios were submitted to ChatGPT-4o, Gemini 2.0 Pro, and Claude 4.5 using a uniform structured prompt. Three blinded expert raters independently evaluated generated prescriptions across five dimensions, clinical relevance, safety, personalization, clarity and structure, practical applicability on a 5-point Likert scale. Inter-rater agreement was assessed using the intraclass correlation coefficient (ICC, two-way mixed model, absolute agreement, average measures). Kruskal-Wallis test, Dunn post-hoc comparisons, and linear mixed models were applied. Descriptive comparisons showed higher mean total scores for Claude 4.5 (19.10 ± 3.65) than for Gemini 2.0 Pro (17.67 ± 2.83) and ChatGPT-4o (16.65 ± 3.33; Kruskal-Wallis p = 0.0084). The primary linear mixed model analysis confirmed higher scores for Claude 4.5 across all five dimensions, with the largest and most clinically meaningful difference observed for clinical relevance (β = 0.92; 95% CI: 0.66–1.18; p < 0.001; η² = 0.217), while effect sizes were small for all other dimensions. Inter-rater agreement was poor across all dimensions (ICC range: −0.08 to 0.25), reflecting substantial variability in expert ratings at the individual prescription level. LLMs demonstrate variable performance for APA prescription in rheumatology, with clinically meaningful differences primarily confined to clinical relevance. Expert validation remains essential, particularly in low-resource settings where contextual adaptation is critical, pending prospective validation in real clinical settings.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1186/s13102-026-01995-0
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.