ICML 2026 Sentinel VLA - Heungwoo/research GitHub Wiki
Sentinel-VLA — A metacognitive VLA with active status monitoring for on-demand reasoning and error recovery
Venue: ICML 2026 (Poster) Category: Reasoning / VLA Architecture Affiliations: (from arXiv) Wenhao Li, Xiu Su, Yichao Cao, Hongyan Xu, Xiaobo Xia, Shan You, Yi Chen, Chang Xu Traction (2026-06): 5 citations (arXiv)

Problem
Current VLA models suffer from three coupled failures: insufficient reasoning (most act as direct input-to-action maps), no status monitoring (they keep executing even after entering a faulty state, e.g. grasping an empty kettle when told to pour), and no self-correction. Existing fixes are piecemeal: "reason-at-every-step" chain-of-thought methods (ECoT, CoT-VLA) add high latency and rigidity; OneTwoVLA's stochastic special-token gating is unstable and uninterpretable and cannot think and act concurrently; and external-monitor recovery systems (AHA, Phoenix, Racer) are architecturally cumbersome and bottlenecked by the VLA's own inability to follow complex external corrections.
Method
Sentinel-VLA is built on π0 (3B PaliGemma VLM expert + 330M Gemma action expert) and adds a third Status Monitor expert (E_sm) with the action-expert architecture; the three experts share attention.

Active monitoring + dynamic reasoning. Each step, E_sm runs a learnable [MONITOR] query that cross-attends to the VLM expert's KV cache and an MLP head emits a trigger status in {Initial, Normal, New-subtask, Error}. A thought memory M_t (plan, current subtask, error reflections) is updated only on demand: Initial → generate plan; New-subtask → advance subtask; Error → formulate a recovery plan + reflection; Normal → no new thought (M_t = M_{t-1}), which covers the vast majority of frames and avoids per-step reasoning cost. The action expert then conditions on (image, instruction, status, memory). Training combines a flow-matching action loss with cross-entropy thought and monitor losses, end-to-end.
EC-Gen data pipeline. Successful SE(3) waypoint trajectories are perturbed by an error operator Φ covering three modalities — object-interaction (suppress gripper change), spatial (add pose noise), and semantic (shift toward a distractor object) — then a recovery segment is inserted. CoT/status labels are auto-annotated; action loss is masked on the injected erroneous actions. This yields 11,000 trajectories / ~2.6M transitions across 44 RLBench tasks with no manual collection.
Self-Evolving Continual Learning (SECL) + OC-Adapter. The deployed model identifies "boundary" settings where success rate ∈ [20%, 80%], collects successful rollouts there, and trains an online LoRA adapter merged into offline weights via EMA (α=0.9). The Orthogonal Continual Adapter adds a penalty ||B_offline × B_online||² so new skills occupy a decorrelated column space, mitigating catastrophic forgetting.
Results
- RLBench Seen (9 tasks): 63.5% avg vs π0 57.8%, ECoT 42.4%, OpenVLA 35.6%.
- RLBench Unseen: 51.3% vs π0 42.0% — e.g. "Wine at rack" 28%, triple OpenVLA's 8%.
- RLBench Disturbed (5% per-step perturbation): 54.7% vs π0 46.0%; OpenVLA collapses 35.6%→25.6%.
- LIBERO-LONG: 90.7% vs π0 85.2%, OpenVLA 53.7%.
- Real-world (Agilex Piper, 3 tasks): 60.0% vs π0 46.0% and OpenVLA 30.7% — over 30% relative gain over the SOTA base π0.
- Efficiency: 13 ms/action on an RTX 4090, vs ECoT's 1528 ms — cognition without sacrificing control frequency.
- Status Monitor: F1 0.902 (sim) / 0.857 (real); error detection 97.4% (sim) / 90.6% (real). Ablations confirm SECL and OC-Adapter are jointly necessary (full 60.0% vs 54.0% w/o SECL; SECL without OC-Adapter regresses to 44.7%).
Significance
Sentinel-VLA unifies reasoning, status monitoring, and recovery inside a single end-to-end VLA rather than bolting on external models, and its on-demand trigger sidesteps the latency tax of constant reasoning while remaining interpretable (discrete status states vs. a stochastic token). Combined with a fully automatic error-recovery data generator and an orthogonal continual-learning loop, it offers a practical path to self-correcting, continually improving manipulation policies.
Links
- arXiv: 2605.01191
- ICML 2026: https://icml.cc/virtual/2026/poster/61750
← Back to ICML-2026