MARATTO

dataset · Zenodo (CERN European Organization for Nuclear Research)

Locating the Sample-Size Ceiling for Handcrafted Features in Chest Radiograph Classifiers: A Factorial Design on a Leakage-Screened Cohort

2026Open accessFayoum University

In plain language

Hybrid artificial intelligence models combining deep learning with handcrafted features are frequently proposed for diagnosing conditions such as pneumonia, COVID-19, and tuberculosis from chest radiographs. By evaluating these architectural choices through a strict factorial design across more than twenty thousand screened images, testing revealed that neither handcrafted features nor gating mechanisms yield meaningful accuracy improvements at full dataset size. Handcrafted features offered measurable benefits only at smaller sample sizes, with their performance advantage falling steadily from over one percentage point down to negligible gains as training sets expanded. While convolutional models reliably reached approximately 97 percent accuracy, lingering errors concentrated on tuberculosis cases being misclassified as normal. These findings locate the sample size boundary where handcrafted visual features cease to provide utility, showing that complex hybrid designs add computational cost without boosting diagnostic accuracy when ample training data is available.

Key takeaways

  • Handcrafted feature branches and gating mechanisms offer negligible accuracy benefits when deep-learning models are trained on large chest radiograph datasets.
  • The performance contribution of handcrafted visual features declines steadily as training data volume increases, becoming indistinguishable from zero beyond half the sample size.
  • Convolutional models achieved around 97 percent accuracy, but residual errors disproportionately misclassified 6.3 percent of tuberculosis cases as normal.
  • External clinical validation remains necessary because an evaluation attempt on a second cohort failed its pre-registered control.

Why it matters

Automating disease detection on chest X-rays requires building reliable, streamlined diagnostic software. Showing that complex hybrid architectures offer no real advantage over standard deep-learning models at scale helps engineers avoid unnecessary engineering overhead. It also highlights critical diagnostic blind spots, such as missed tuberculosis cases, which must be resolved before automated screening tools can safely support healthcare professionals.

Commercialisation angle

This early-stage algorithmic research informs developers of clinical decision-support software for medical radiology. By demonstrating that handcrafted visual features fail to justify their computational overhead at scale, it helps technology teams streamline automated screening architectures. However, practical deployment remains distant, as the models exhibit critical detection errors with tuberculosis and external validation failed pre-registered controls.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Background/Objectives: Attribution fails when two design choices move together. Hybrid deep-learning classifiers for normal, COVID-19, pneumonia and tuberculosis on chest radiographs usually change what is fused and how at once, so a gain is attributable to neither, and this study crosses them. Methods: Three pooled public sources passed exact-byte, integrity and perceptual screens; screening removed 142 near-duplicates. Patient grouping kept 703 of the 4,069 an image-level split would have produced from sharing a subject with training. Seven architectures were trained under one protocol with five-fold cross-validation on 20,342 images; four form the 2-by-2 design. Results: The handcrafted-only model trailed the six convolutional configurations, all near 97%, by 4.5 to 5.1 points, the only one of eighteen comparisons to survive correction. Neither crossed factor did: the handcrafted branch was worth +0.010 accuracy points (95% confidence interval −0.36 to +0.39), the gate −0.004 (−0.54 to +0.53). A label rule reading two directory levels had put that branch at +0.81 in one fold. The trained gate spans 0.12 to 0.88 across images, and a hundredfold higher fine-tuning rate moved its parameters further than the schedule reported here without a distinguishable change in accuracy. The null held for ranking, convergence speed and saliency: 130 of 140 zonal contrasts reversed sign. Residual error concentrates on tuberculosis: 6.3% read as normal, which is not a detection rate. A probe naming the source reaches 88.8% against 47.1%, yet transfer between collections gives 96.79% and 93.34% against 72.89% and 52.68%. Conclusions: Neither component is bought at full data; both are paid for. The handcrafted effect falls monotonically from +1.302 to +0.010 as the training fold grows; its interval excludes zero at a quarter of the fold and not at a half, which locates the ceiling. An attempt on a second cohort failed its pre-registered control, so external validation remains necessary.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.21988129

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.