Backdoor as Probe: Test-Time Adversarial Defense for CLIP

Zhongqi Wang1,2, Jie Zhang1,2, Sen Nie1,2, Zhiyu Chen3, Shiguang Shan1,2, Xilin Chen1,2

1Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences

2University of Chinese Academy of Sciences

3Xuzhou University of Technology

Comparison of a backdoor attack, an adversarial attack, and BaP as a defensive probe
BaP repurposes the trigger-to-target mechanism of backdoors as a defender-controlled probe for adversarial detection and correction.

Abstract

Test-time adversarial defense improves the robustness of vision-language foundation models such as CLIP without retraining. However, adversarial activation shifts are typically treated as distortions to suppress, rather than signals to exploit. We turn these shifts into defense signals by repurposing the trigger-to-target mechanism of backdoors. The key is to implant a defender-controlled backdoor as a probe that is weakly activated by clean inputs but strongly activated by adversarial shifts. Based on this insight, we propose Backdoor as Probe (BaP), a test-time adversarial defense for CLIP. BaP constructs the probe through a closed-form model edit to a selected MLP layer. It projects the average adversarial activation shift and a defender-specified semantic direction onto the layer's low-energy input and output activation subspaces to obtain the trigger and target directions, respectively. At inference time, adversarial inputs produce measurable responses along the target direction for detection. BaP then selectively rectifies detected inputs by optimizing a small perturbation that steers their representations away from adversarial shifts and toward the clean subspace. Experiments across 16 benchmarks show that BaP improves average robust accuracy from 1.0% to 52.3% while retaining clean accuracy, achieving performance comparable to state-of-the-art methods with up to a 5.7x inference speedup. BaP further shows the generalization to adversarial attacks on large vision-language models.

Method Overview

BaP method overview showing defender-controlled probe implantation and probe-guided detection and rectification

Probe implantation

BaP constructs an input-side probe direction by reweighting the average adversarial activation shift within the low-energy input subspace of an MLP's fc2. It projects a defender-specified semantic direction into the corresponding low-energy output subspace and binds the two directions through a closed-form rank-one edit.

Detection

For each input, BaP reads the edited fc2 output along the semantic target direction. The edit keeps clean responses concentrated near zero while shifting adversarial responses to larger values; an input is flagged when its probe response crosses the fixed detection threshold.

Correction

BaP rectifies only flagged inputs in two stages. Two projected-gradient ascent steps first move the representation away from its current adversarial state. One projected-gradient descent step then reduces its residual outside the clean principal subspace and aligns that residual with the global clean direction.

Defense on Large Vision-Language Models

Qualitative results on LLaVA under M-Attack and FOA-Attack. The undefended model produces attack-induced concepts highlighted in red, while BaP recovers image-consistent semantics highlighted in green.

Qualitative comparison of LLaVA captions before and after BaP correction under M-Attack and FOA-Attack
BaP corrects adversarially manipulated captions across both evaluated LVLM attacks. Ellipses reproduce truncated captions from the paper visualization.

Results

52.3%Average robust accuracy
61.1%Average clean accuracy
5.7×Inference speedup vs. R-TPT
16Zero-shot benchmarks

Results are reported for the evaluated datasets, CLIP backbones, threat models, and attack configurations. See the paper for complete settings, transfer studies, and adaptive-attack analysis.

Citation

@article{wang2026backdoor,
  title  = {Backdoor as Probe: Test-Time Adversarial Defense for CLIP},
  author = {Wang, Zhongqi and Nie, Sen and Zhang, Jie and Chen, Zhiyu and Shan, Shiguang and Chen, Xilin},
  year   = {2026},
  note   = {Preprint}
}