Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms

A Probe-Based Detector Using Saliency Divergence, Gradients, and Layer Sensitivity

Ehsan Ur Rahman Mohammed1, Elham Bagheri2, Apurva Narayan1, Yalda Mohsenzadeh1; 1University of Western Ontario, 2Vector Institute for Artificial Intelligence

Presenter: Yalda Mohsenzadeh

Adversarial patches cause targeted misclassification by steering a model’s evidence toward a small, visible region while human perception remains largely unaffected. We propose a post hoc detector that attaches lightweight probes to a frozen classifier and fuses three complementary signals: (i) input-gradient statistics of the predicted class, (ii) layer-wise sensitivity to small activation noise, and (iii) human-model saliency divergence, quantified by comparing Grad-CAM with human saliency maps. Features from these probes are fed to a small secondary classifier (detector) that predicts whether an input is patched. To our knowledge, our detector is the first adversarial-patch detector to explicitly incorporate human attention modeling via saliency divergence, aligning what the model relies on with where humans look, without modifying or retraining the base model. Across CAT2000, FIGRIM, and SALICON, using ResNet-50 and EfficientNet-B0 backbones, our detector achieves F1 scores up to 99.6% and remains in the 80–99% range across the adversarial patch attack settings, outperforming probing baselines, SentiNet, and X-Detect. Ablations show gains from adding gradients and divergence to sensitivity in most of the cases, indicating complementary cues and highlighting divergence’s discriminative power. Our detector is simple to deploy (post hoc; no retraining) and provides interpretable rationales via its saliency and sensitivity components, suggesting a practical path to robust, explainable detection of adversarial patches.

Topic Area: Methods, Tools, Theory & Neural Coding