MARATTO

article

Fusion of Head Pose and Facial Action Units for Automated Student Engagement Estimation: A Study on DAD-3DHeads

Abstract

We present a reproducible pipeline for automated student engagement estimation using a multimodal fusion of head pose angles and facial action units (AUs) extracted from 3D facial images. Leveraging the DAD-3DHeads dataset, which captures rich variations in pose, expression, occlusion, and lighting, our method involves systematic preprocessing, robust feature extraction, and explainable decision-making. A proxy engagement score is computed for each sample based on neutral facial expression, frontal head orientation, and absence of occlusion—serving as an interpretable surrogate ground truth. After normalization, multiple supervised classifiers are evaluated, including Random Forest, Support Vector Machine, XGBoost, and Multi-Layer Perceptron. The best-performing model, XGBoost, achieves 77 % accuracy and an F1-score of 0.82 on the test set. Feature importance analysis confirms the complementary roles of pose and AUs in engagement prediction. This interpretable and fully reproducible framework offers a practical alternative to black-box deep learning and may be extended to broader contexts requiring non-intrusive behavioral assessment.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/sita67914.2025.11273703

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.