MARATTO

article · European Review of Aging and Physical Activity

Quality of AI-generated exercise awareness messages for older adults aligned with the ICFSR consensus: a comparative study of three LLMs

Abstract

BACKGROUND: Large Language Models (LLMs) are emerging as potential tools for health communication and patient education. However, their ability to translate complex medical guidelines into accessible, safe, and accurate messages for older adults remains insufficiently evaluated. To assess the capacity of ChatGPT-4, Claude AI, and Deepseek to generate exercise awareness messages for older adults, aligned with international expert consensus. MATERIALS AND METHODS: This was a cross-sectional observational study conducted between April and August 2025. Using the ICFSR 2021 Global Consensus on optimal exercise recommendations as the gold standard, we evaluated messages generated by three LLMs for 13 common chronic conditions in older adults. A standardized prompt was used to generate messages addressing exercise prescription considerations, disease progression, and recommended modalities. Two independent expert evaluators; a geriatrician and a sports medicine physician assessed five dimensions using a 5-point Likert scale; accuracy, clarity, safety, behavioral relevance, and absence of fabrication. Readability was measured using Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), and Simple Measure of Gobbledygook (SMOG). Inter-rater agreement was assessed using intraclass correlation coefficient (ICC). RESULTS: Claude achieved the overall scores from both evaluators at 4.63 ± 0.55 and 4.60 ± 0.55, followed by ChatGPT-4 at 4.55 ± 0.50 and 4.48 ± 0.53 and Deepseek at 4.15 ± 0.87 and 4.17 ± 0.86. Inter-rater agreement was moderate was 0.668. Claude demonstrated accuracy scores at 4.58 ± 0.64, while ChatGPT-4 excelled in clarity with 5 ± 0. All models achieved perfect scores for absence of fabrication (5 ± 0). Readability indices revealed high complexity across all LLMs, with median FKGL values ranging from 10.84 to 10.93, corresponding to 10th-11th grade reading level, exceeding recommended levels for older populations. A significant correlation was found between Claude's accuracy scores and FKGL (r = 0.599, p = 0.030). CONCLUSION: LLMs, particularly Claude, ChatGPT-4 and Deepseek demonstrate strong potential for generating accurate, safe, and hallucination-free exercise awareness messages for older adults. However, readability remains above recommended levels, requiring optimization.

Research topics

  • Artificial Intelligence in Healthcare and Education
  • Mobile Health and mHealth Applications
  • Machine Learning in Healthcare

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1186/s11556-026-00428-8

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.