MARATTO

dataset · Zenodo (CERN European Organization for Nuclear Research)

Locating the Sample-Size Ceiling for Handcrafted Features in Chest Radiograph Classifiers: A Factorial Design on a Leakage-Screened Cohort

2026Open accessFayoum University

In plain language

Hybrid deep learning models that combine handcrafted image features with neural networks are often used to classify chest radiographs for conditions such as tuberculosis, pneumonia, and COVID-19. By evaluating architectures across 20,342 rigorously screened images, this research tests whether adding handcrafted features or gating mechanisms genuinely improves classification performance. Six convolutional configurations achieved around 97 percent accuracy, substantially outperforming purely handcrafted models. Adding handcrafted features or gating mechanisms offered no statistically meaningful improvement at full dataset scale, contributing changes of only 0.010 and minus 0.004 accuracy points, respectively. The advantage of handcrafted features diminishes steadily from 1.302 points at smaller sample sizes down to near zero as the training data expands, losing statistical significance once training data exceeds a quarter of the full cohort. Furthermore, residual classification errors concentrated on tuberculosis, and external validation remained inconclusive due to a failed control test.

Key takeaways

  • Adding handcrafted features or gating mechanisms to convolutional neural networks yields no meaningful accuracy gain when large training datasets are available.
  • The performance contribution of handcrafted features declines monotonically from 1.302 points at small sample sizes to 0.010 points on the full dataset.
  • Purely convolutional models achieve approximately 97 percent accuracy, outperforming purely handcrafted feature models by around five percentage points.
  • Remaining classification errors disproportionately affect tuberculosis cases, with 6.3 percent misclassified as normal.

Why it matters

Medical artificial intelligence often incorporates complex hybrid architectures under the assumption that combining traditional handcrafted features with deep learning improves diagnostic accuracy. Demonstrating that handcrafted features cease to add value once training datasets reach a certain size allows developers to avoid unnecessary computational complexity, streamline model designs, and direct engineering effort towards resolving critical diagnostic errors, particularly in conditions like tuberculosis.

Commercialisation angle

This early-stage research is relevant to diagnostic software developers creating computer vision tools for clinical radiology, particularly for screening chest infections. By showing that complex hybrid feature pipelines do not improve accuracy at scale, the findings can guide software engineers to simplify algorithm pipelines and cut training costs. However, the technology remains at a pre-clinical, early stage, especially given that residual errors affect tuberculosis detection and external cohort validation failed its control test.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Background/Objectives: Attribution fails when two design choices move together. Hybrid deep-learning classifiers for normal, COVID-19, pneumonia and tuberculosis on chest radiographs usually change what is fused and how at once, so a gain is attributable to neither, and this study crosses them. Methods: Three pooled public sources passed exact-byte, integrity and perceptual screens; screening removed 142 near-duplicates. Patient grouping kept 703 of the 4,069 an image-level split would have produced from sharing a subject with training. Seven architectures were trained under one protocol with five-fold cross-validation on 20,342 images; four form the 2-by-2 design. Results: The handcrafted-only model trailed the six convolutional configurations, all near 97%, by 4.5 to 5.1 points, the only one of eighteen comparisons to survive correction. Neither crossed factor did: the handcrafted branch was worth +0.010 accuracy points (95% confidence interval −0.36 to +0.39), the gate −0.004 (−0.54 to +0.53). A label rule reading two directory levels had put that branch at +0.81 in one fold. The trained gate spans 0.12 to 0.88 across images, and a hundredfold higher fine-tuning rate moved its parameters further than the schedule reported here without a distinguishable change in accuracy. The null held for ranking, convergence speed and saliency: 130 of 140 zonal contrasts reversed sign. Residual error concentrates on tuberculosis: 6.3% read as normal, which is not a detection rate. A probe naming the source reaches 88.8% against 47.1%, yet transfer between collections gives 96.79% and 93.34% against 72.89% and 52.68%. Conclusions: Neither component is bought at full data; both are paid for. The handcrafted effect falls monotonically from +1.302 to +0.010 as the training fold grows; its interval excludes zero at a quarter of the fold and not at a half, which locates the ceiling. An attempt on a second cohort failed its pre-registered control, so external validation remains necessary.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.21988130

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.