MARATTO

article · Scientific Reports

Comparative analysis of automated foul detection in football using deep learning architectures

202535 citationsOpen accessSuez University

In plain language

This study conducted a comprehensive comparison of eight state-of-the-art Deep Learning (DL) architectures for automated foul detection in football. Models including EfficientNetV2, ResNet50, and InceptionResNetV2 were trained and evaluated on a curated dataset of 7000 images, split for training, validation, and testing. Performance was assessed using metrics such as accuracy, precision, recall, F1-score, and AUC. InceptionResNetV2 achieved the highest test accuracy of 87.57% and a strong F1-score. DenseNet121 showed the highest precision and AUC, indicating superior discriminatory power. Lightweight models like MobileNetV2 also performed competitively, suggesting their suitability for real-time applications. The research highlights the viability of integrating DL into football officiating systems like VAR, emphasising the importance of model explainability.

Key takeaways

  • Eight deep learning architectures were comparatively evaluated for automated foul detection in football.
  • InceptionResNetV2 achieved the highest test accuracy (87.57%) and F1-score (0.8966) among the models tested.
  • DenseNet121 demonstrated the highest precision (0.9786) and Area Under the Receiver Operating Characteristic Curve (AUC) of 0.9641.
  • Lightweight models, such as MobileNetV2, performed competitively, indicating their potential for real-time deployment.
  • The study supports the integration of deep learning architectures into existing football officiating systems, like the Video Assistant Referee (VAR), and stresses the importance of model explainability.

Why it matters

This research is important because it offers a pathway to enhance the accuracy and consistency of foul detection in football, potentially reducing human error and improving fairness in the game. By leveraging deep learning, it could support officials in making more informed and objective decisions during matches.

Commercialisation angle

This research is directly applicable to sports technology companies and football organisations seeking to improve officiating. The developed deep learning models could be integrated into existing Video Assistant Referee (VAR) systems or new real-time foul detection tools. The competitive performance of lightweight models suggests potential for deployment in real-time scenarios, offering a near-market solution for enhancing sports analytics and decision-making support.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Automated foul detection in football represents a challenging task due to the dynamic nature of the game, the variability in player movements, and the ambiguity in differentiating fouls from regular physical contact. This study presents a comprehensive comparative evaluation of eight state-of-the-art Deep Learning (DL) architectures - EfficientNetV2, ResNet50, VGG16, Xception, InceptionV3, MobileNetV2, InceptionResNetV2, and DenseNet121 - applied to the task of automated foul detection in football. The models were trained and evaluated using a curated dataset comprising 7000 images, which was split into 70% for training (4,900 images), 20% for validation (1,400 images), and 10% for testing (700 images). To ensure fair evaluation, the test set was balanced to contain 350 images depicting foul events and 350 images representing non-foul scenarios, although perfect balance was subject to class distribution constraints. Performance was assessed across multiple metrics, including test accuracy, precision, recall, F1-score, and Area Under the Receiver Operating Characteristic Curve (AUC). The results demonstrate that InceptionResNetV2 achieved the highest test accuracy of 87.57% and a strong F1-score of 0.8966, closely followed by DenseNet121, which attained the highest precision of 0.9786 and an AUC of 0.9641, indicating superior discriminatory power. Lightweight models such as MobileNetV2 also performed competitively, highlighting their potential for real-time deployment. The findings highlight the strengths and trade-offs between model complexity, accuracy, and generalizability, underscoring the viability of integrating DL architectures into existing football officiating systems, such as the Video Assistant Referee (VAR). Furthermore, the study emphasizes the importance of model explainability through techniques such as Gradient-weighted Class Activation Mapping++ (GradCAM++), ensuring that automated decisions can be accompanied by interpretable visual evidence. This comparative evaluation serves as a foundation for future research aimed at enhancing real-time foul detection through multimodal data fusion, temporal modeling, and improved domain adaptation techniques.

Research topics

  • Sports injuries and prevention
  • Sports Analytics and Performance
  • Sports Performance and Training

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1038/s41598-025-96945-0

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.