article · Scientific Reports
Parkinson's Disease (PD) is a progressive neurodegenerative disorder that causes motor and cognitive impairments, affecting approximately 1% of individuals over 60 years of age. Speech impairments are among the earliest and most accessible biomarkers, making voice-based assessment a promising avenue for remote PD monitoring. However, existing speech-based PD prediction methods suffer from feature redundancy that degrades model performance, non-Gaussian data distributions that violate model assumptions, and limited systematic feature grouping strategies. This study introduces an adaptive approach to improve PD diagnostic precision by predicting the Motor Unified PD Rating Scale (UPDRS) and Total-UPDRS scores from biomedical voice measurements. The proposed framework addresses these challenges through three integrated components: (1) Box-Cox transformation to stabilize variance, reduce skewness, and normalize features; (2) a clustering-based feature selection method that groups correlated features via K-Means and selects the most informative representative per cluster using mutual information, thereby eliminating redundancy without losing discriminative power; and (3) an Extra Trees Regressor (ETR) whose extreme randomization in node splitting provides computational efficiency and reduced variance. To ensure rigorous evaluation, a subject-independent data splitting strategy is adopted to prevent data leakage, and k-fold cross-validation is employed to assess model stability. The proposed method is compared against multiple feature selection techniques-mutual information, recursive feature elimination, Lasso regression, and autoencoders-paired with nine regression models including Ridge, Lasso, Linear, Decision Tree, k-Nearest Neighbors, Random Forest, Gradient Boosting, AdaBoost, and Extra Trees Regressors. The clustering-based feature selection combined with ETR yielded the best performance, achieving [Formula: see text] scores of 0.999 for Motor-UPDRS and 0.997 for Total-UPDRS on the test set. These results are further supported by cross-validation analysis and feature importance evaluation, demonstrating the effectiveness and robustness of the proposed framework for speech-based PD telemonitoring.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1038/s41598-026-49065-2
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.