ICML 2026 Sentinel VLA - Heungwoo/research GitHub Wiki

Sentinel-VLA — A metacognitive VLA with active status monitoring for on-demand reasoning and error recovery

Venue: ICML 2026 (Poster) Category: Reasoning / VLA Architecture Affiliations: (from arXiv) Wenhao Li, Xiu Su, Yichao Cao, Hongyan Xu, Xiaobo Xia, Shan You, Yi Chen, Chang Xu Traction (2026-06): 5 citations (arXiv)

The performance and mechanism of Sentinel-VLA (Figure 1 from Li et al., 2026)

Problem

Current VLA models suffer from three coupled failures: insufficient reasoning (most act as direct input-to-action maps), no status monitoring (they keep executing even after entering a faulty state, e.g. grasping an empty kettle when told to pour), and no self-correction. Existing fixes are piecemeal: "reason-at-every-step" chain-of-thought methods (ECoT, CoT-VLA) add high latency and rigidity; OneTwoVLA's stochastic special-token gating is unstable and uninterpretable and cannot think and act concurrently; and external-monitor recovery systems (AHA, Phoenix, Racer) are architecturally cumbersome and bottlenecked by the VLA's own inability to follow complex external corrections.

Method

Sentinel-VLA is built on π0 (3B PaliGemma VLM expert + 330M Gemma action expert) and adds a third Status Monitor expert (E_sm) with the action-expert architecture; the three experts share attention.

Left: Sentinel-VLA pipeline — the Status Monitor activates on-demand Adaptive Thought. Right: SECL continual-learning loop with orthogonal-constrained adapter (Figure 3 from Li et al., 2026)

Active monitoring + dynamic reasoning. Each step, E_sm runs a learnable [MONITOR] query that cross-attends to the VLM expert's KV cache and an MLP head emits a trigger status in {Initial, Normal, New-subtask, Error}. A thought memory M_t (plan, current subtask, error reflections) is updated only on demand: Initial → generate plan; New-subtask → advance subtask; Error → formulate a recovery plan + reflection; Normal → no new thought (M_t = M_{t-1}), which covers the vast majority of frames and avoids per-step reasoning cost. The action expert then conditions on (image, instruction, status, memory). Training combines a flow-matching action loss with cross-entropy thought and monitor losses, end-to-end.

EC-Gen data pipeline. Successful SE(3) waypoint trajectories are perturbed by an error operator Φ covering three modalities — object-interaction (suppress gripper change), spatial (add pose noise), and semantic (shift toward a distractor object) — then a recovery segment is inserted. CoT/status labels are auto-annotated; action loss is masked on the injected erroneous actions. This yields 11,000 trajectories / ~2.6M transitions across 44 RLBench tasks with no manual collection.

Self-Evolving Continual Learning (SECL) + OC-Adapter. The deployed model identifies "boundary" settings where success rate ∈ [20%, 80%], collects successful rollouts there, and trains an online LoRA adapter merged into offline weights via EMA (α=0.9). The Orthogonal Continual Adapter adds a penalty ||B_offline × B_online||² so new skills occupy a decorrelated column space, mitigating catastrophic forgetting.

Results

  • RLBench Seen (9 tasks): 63.5% avg vs π0 57.8%, ECoT 42.4%, OpenVLA 35.6%.
  • RLBench Unseen: 51.3% vs π0 42.0% — e.g. "Wine at rack" 28%, triple OpenVLA's 8%.
  • RLBench Disturbed (5% per-step perturbation): 54.7% vs π0 46.0%; OpenVLA collapses 35.6%→25.6%.
  • LIBERO-LONG: 90.7% vs π0 85.2%, OpenVLA 53.7%.
  • Real-world (Agilex Piper, 3 tasks): 60.0% vs π0 46.0% and OpenVLA 30.7% — over 30% relative gain over the SOTA base π0.
  • Efficiency: 13 ms/action on an RTX 4090, vs ECoT's 1528 ms — cognition without sacrificing control frequency.
  • Status Monitor: F1 0.902 (sim) / 0.857 (real); error detection 97.4% (sim) / 90.6% (real). Ablations confirm SECL and OC-Adapter are jointly necessary (full 60.0% vs 54.0% w/o SECL; SECL without OC-Adapter regresses to 44.7%).

Significance

Sentinel-VLA unifies reasoning, status monitoring, and recovery inside a single end-to-end VLA rather than bolting on external models, and its on-demand trigger sidesteps the latency tax of constant reasoning while remaining interpretable (discrete status states vs. a stochastic token). Combined with a fully automatic error-recovery data generator and an orthogonal continual-learning loop, it offers a practical path to self-correcting, continually improving manipulation policies.

Links

← Back to ICML-2026