article · Applied Sciences
Accurate iris segmentation remains a fundamental challenge in iris biometric recognition and medical image analysis, particularly in challenging scenarios such as non-cooperative acquisition conditions involving variable illumination, partial occlusions, degraded image quality, and diverse unconstrained environments. Prevailing segmentation algorithms exhibit limited robustness when confronted with such challenges, and the disparity between near-infrared (NIR) and visible-light imaging modalities further compounds the complexity of achieving a robust segmentation outcome. To address these challenges, this paper introduces SwinIrisNet, a hybrid deep learning architecture that integrates Swin Transformer and convolutional neural network (CNN) branches within a U-Net framework for robust iris segmentation. The Swin Transformer branch leverages hierarchical window-based self-attention to capture global contextual dependencies, whereas the CNN branch extracts fine-grained local features essential for precise boundary delineation. A memory-efficient cross-attention fusion module combines these complementary feature representations, further enhanced by a Convolutional Block Attention Module (CBAM), Atrous Spatial Pyramid Pooling (ASPP), and attention-gated skip connections for multi-scale context aggregation. An extensive evaluation is conducted across four publicly available benchmark datasets, including UBIRIS.v2, IITD, CASIA-Thousand, and MMU.v1, encompassing both visible-light and NIR imaging environments. The proposed architecture yields F1 values of 0.9612–0.9672, Dice coefficients of 0.9489–0.9519, mIoU values of 0.9266–0.9450, precision values of 0.9565–0.9633, recall values of 0.9600–0.9672, and classification accuracies of 99.51–99.53%, with NICE1 error rates of 0.57–0.60% and NICE2 values of 1.82–2.24%, confirming pixel-level segmentation quality. Cross-database generalization experiments further demonstrate that SwinIrisNet learns transferable iris representations and generalizes effectively across heterogeneous imaging sources, with the strongest transfer occurring in the NIR-to-visible direction. A comparative analysis against existing algorithms demonstrates that the proposed architecture attains substantial performance improvements over several existing segmentation networks when evaluated on identical benchmark databases, surpassing them across the majority of qualitative and quantitative metrics while maintaining a marginally lower memory footprint.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.3390/app16178892
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.