ICML 2026 TRAP - Heungwoo/research GitHub Wiki

TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches — Corrupting a robot's chain-of-thought to deliver a knife instead of an apple

Venue: ICML 2026 (Poster) Category: Analysis-Insight (Adversarial Robustness / Security of VLA) Affiliations: Zhengxian Huang, Wenjun Zhu, Haoxuan Qiu, Xiaoyu Ji, Wenyuan Xu (arXiv:2603.23117) Traction (2026-06): 2 citations (arXiv)

The TRAP attack hijacks the VLA's output by influencing its CoT reasoning via an adversarial patch placed on the worktable, causing the robot to deliver a knife instead of the requested apple (Figure 1 from Huang et al., 2026)

Problem

Chain-of-Thought (CoT) reasoning has become a popular addition to Vision-Language-Action (VLA) models, improving generalization and interpretability by making the policy "think" in language before acting. TRAP argues that this same intermediate reasoning is an unguarded attack surface. The authors first give empirical evidence that the generated CoT strongly governs action generation — even when the CoT is semantically misaligned with the user's instruction. That observation implies an attacker who can corrupt the CoT can steer the robot's behavior without ever touching the user's instruction. The motivating scenario: the user says "pick and give me the apple," but a malicious object on the table makes the robot reason its way into delivering a knife to the person.

Method

TRAP is presented as the first targeted adversarial attack framework for CoT-reasoning VLAs. Rather than perturbing pixels imperceptibly, it optimizes a physically realizable adversarial patch (e.g., a coaster placed on the table) that corrupts the model's intermediate CoT and thereby hijacks the downstream action toward an adversary-defined goal.

The optimization combines two losses. A CoT Hijacking Loss drives the model to emit the attacker's desired reasoning trace, and an Action Loss ties that hijacked reasoning to the concrete target action. A Physical Robustness stage makes the patch survive the sim-to-real gap, using homography transformation, color smoothing, and color calibration so the printed-on-paper patch keeps working under real cameras and lighting. The attack is evaluated across 3 mainstream VLA architectures and 3 distinct CoT reasoning paradigms.

Existing representative paradigms of reasoning VLAs that TRAP targets (Figure 2 from Huang et al., 2026)

Results

On the averaged main results (Table 3, across all 5 tasks and 3 victim models — MolmoACT, InstructVLA, GraspVLA), TRAP reaches an average Attack Success Rate of 52.54% with a target Score of 0.3294, versus 40.86% / 0.2329 for a CoT-Only ablation, 5.48% for Action-Only, and just 1.56% for Random Noise. Per-model ASR is highest on GraspVLA (75.84%) and MolmoACT (48.06%). The attack is also layout-robust: on unseen table layouts average ASR barely drops to 51.60%, which the authors cite as evidence TRAP learns layout-invariant adversarial features rather than overfitting to a spatial configuration. A transferability study shows the patch binds to specific object names, behaving like a trigger that activates whenever those names are mentioned (e.g., 53.6% and 57.6% ASR on MolmoACT). The attack was validated physically by printing the patch on paper.

Significance

TRAP shows that adding CoT reasoning — usually framed as a safety/interpretability win — simultaneously expands the attack surface of VLA robots. Because the patch is a benign-looking physical object and the user's instruction is untouched, the attack is stealthy and deployable, underscoring an urgent need to secure the reasoning channel, not just the perception or action channels, in embodied agents.

Links

← Back to ICML-2026