Label Scarcity Reveals Architecture-Dependent Robustness in CNN-Based Casting Defect Detection
Deep learning for automated visual inspection in manufacturing is often limited by the cost of labeling large image datasets. Using 7,348 casting images, we evaluated how three CNN architectures, ResNet18, ResNet50, and MobileNetV3-Small, tolerate progressively reduced labeling budgets, from 100% down to 1% of the training data. Under full supervision, all three architectures achieved mean test accuracies above 99%, appearing nearly interchangeable. This equivalence broke down as labeled data became scarce: ResNet18 and ResNet50 maintained mean accuracies above 99% down to 5% labels, whereas MobileNetV3-Small declined from 99.47% to 83.75% over the same range and exhibited markedly greater run-to-run variability. At the most severe scarcity tested (1% labels), all three architectures degraded substantially, reaching mean accuracies of 89.09%, 86.63%, and 65.71%, respectively. Class-specific error analysis, representative Grad-CAM visualization, and population-level analyses of attention entropy and prediction confidence further revealed increased misclassification, broader spatial attention, and reduced prediction confidence under severe label scarcity. These results demonstrate that robustness to label scarcity is an architecture-dependent property that cannot be inferred from fully supervised accuracy alone, providing a more realistic basis for selecting CNN architectures for data-constrained manufacturing inspection.