ICLR 2026 Self Refining VLM - Heungwoo/research GitHub Wiki

ARMOR — Self-Refining VLM for Robotic Failure Detection and Reasoning

Venue: ICLR 2026 Category: Embodied reasoning · Failure detection Trend tag: Self-refinement · Heterogeneous supervision · VLM monitor Affiliations: UT Austin (Qi, Zhang), Amazon Robotics (Wang, Sheng, Mao, Srinivasan, Nambi, Dattatreya), CMU (Yong)

Approach diagram

flowchart LR
  Vid[Video x<br/>multi-view frames] --> VLM[Qwen2.5-VL backbone<br/>vision encoder + LM decoder]
  VLM --> CLS[Detection head<br/>BCE loss<br/>CLS token + cross-attn + MLP]
  VLM --> NTP[Reasoning head<br/>NTP loss<br/>auto-regressive LM decoder]
  CLS --> Round[Round t prediction lt]
  NTP --> Round2[Round t prediction et]
  Round --> Cond[Condition next round on lt-1, et-1, prompt p]
  Round2 --> Cond
  Cond --> VLM
  Round --> Hdet[Detection entropy H_det]
  Round2 --> Hexp[Reasoning mean-token entropy H_exp]
  Hdet --> C[Combined entropy<br/>C = H_det + λ H_exp]
  Hexp --> C
  C --> Sel[Argmin over M trajectories<br/>Stop when no decrease in entropy]
  Sel --> Out[Final detection l̂ + reasoning ê]
Loading

Problem

Real-world robot failures are subtle, combinatorial, and hard to enumerate. Closed-set classifiers under-cover them, and dense reasoning annotations are expensive. Concretely:

  • Binary success / failure labels are easy to harvest (system logs, sensors) — large-scale, sparse.
  • Free-form reasoning labels ("why did it fail") need human experts — small-scale, dense.

Prior work either (a) treats failure detection as closed-set classification (Ye et al., 2019; Garrett et al., 2020), (b) assumes dense reasoning annotations everywhere (AHA — Duan et al., 2025, applies SFT on full reasoning), or (c) prompts off-the-shelf VLMs without fine-tuning (Skreta et al., 2024; Guo et al., 2024). AHA's evaluation also relies on hand-crafted regex parsing, breaking under format drift.

ARMOR ("Adaptive Round-based Multi-task mOdel for Robotic failure detection and reasoning") tackles failure detection + reasoning under heterogeneous supervision, with open-ended reasoning beyond fixed taxonomies.

Detailed Method

1. Multi-task MDP formulation

Following RISE (Qu et al., 2024), failure understanding is cast as a multi-task MDP M = (S, A, R, T):

  • State s_t = [x, l_{t-1}, e_{t-1}, p] — input video + previous round's detection + reasoning + auxiliary prompt.
  • Action a_t = (l_t, e_t) — joint detection / reasoning prediction.
  • Rewards R = {R_detect, R_reason} measure correctness against ground truth.
  • Horizon T = number of refinement rounds.

Critically, ARMOR does not assume an oracle reward model at training or inference, unlike RISE. It uses heterogeneous supervision instead.

2. Heterogeneous supervision setting

D_sparse = {(xi, li)},       li ∈ {success, failure}
D_dense  = {(xi, li, ei)},   ei = detailed natural-language reasoning

Detection is supervised on D_sparse ∪ D_dense; reasoning only on D_dense. On D_sparse the model still generates reasoning as context but it is unsupervised.

3. Architecture (Appendix A.1)

Backbone: Qwen2.5-VL (Bai et al., 2025), 7B variant for main results ("all open-source VLMs and ARMOR use 7B model variants" unless otherwise specified); a 32B variant is additionally tested in the model-scale ablation (Appendix C.1, Table 8).

  • Detection head: lightweight classifier. Mean-pools features from the first 4 LM-decoder layers, projects them, cross-attends with a learnable [CLS] token (8-head multi-head attention), MLP decoder → binary logits. Trained with binary cross-entropy (BCE).
  • Reasoning head: original Qwen2.5-VL auto-regressive LM decoder. Trained with next-token prediction (NTP).
  • Both heads share the vision encoder and LM decoder backbone parameters.
  • Round-t conditioning is injected via natural-language prompts: "Given the previous detection is ..." / "Given the previous reasoning is ...".

4. Two-phase training (Algorithm 1)

Phase I — Offline Imitation:

  1. Warm-up. Train (l1, e1) ~ πθ(· | [x, ∅, ∅, ∅]) on D_sparse ∪ D_dense with BCE for l and NTP for e. (Empty prior context.)
  2. Expert-conditioned. Train (l1, e1) ~ πθ(· | [x, l, e, p]) on D_dense with same losses. To prevent the model from copying the prior input, one of the two ground-truth inputs is randomly masked per sample while the other is retained — forcing cross-task consistency learning.

Phase II — Online Refinement: For each minibatch, roll out the policy for T rounds:

for t = 1 to T:
  (lt, et) ~ πθ(· | [x, l_{t-1}, e_{t-1}, p_t])
  L += BCE(lt, l) + NTP(et, e) · 1[x ∈ D_dense]
θ ← θ − η ∇θ L

Detection is always supervised; reasoning is supervised only when dense labels exist. This is online imitation learning (Ross et al., 2011) — the model is supervised on its own rollout-state distribution.

5. Inference (Algorithm 2)

At inference, ARMOR samples M refinement trajectories in parallel and selects the most confident:

C^(m) = H_det^(m) + λ · H_reason^(m)        # combined entropy
m* = argmin_m C^(m)                          # most confident trajectory
Stop when C^(m*) ≥ C_min − ε                  # no longer reducing uncertainty

H_det is the detection logits' entropy; H_reason is the mean per-token entropy of the reasoning. λ = 0.1 (detection trusted more — supervised on more data). Final output is (l̂, ê) from the trajectory m* at termination.

6. Hyperparameters (Table 5)

Hyperparameter Value
Backbone Qwen2.5-VL 7B (32B in scaling test, App. C.1)
Global batch size (videos: Sparrow, ARMBench) 16
Global batch size (collaged images: RLBench, Maniskill) 64
Offline epochs 3
Online epochs 10
Horizon T (rounds) 3 (default)
Optimizer AdamW
Weight decay 0.1
LR scheduler Cosine, warmup ratio 0.03
Learning rate (LM + classifier head) 1 × 10⁻⁵
Learning rate (vision encoder) 2 × 10⁻⁶
Refinement weight λ 0.1
Hardware 8 × H100 GPUs (training) / 8 × A100 40GB (inference benchmarking)

Comprehensive Results

Main results (Table 1)

4 datasets:

  • RLBench-Fail (sim) — tabletop tasks with scripted-policy perturbations.
  • Maniskill-Fail (sim) — peg insertion, transport.
  • Sparrow-Fail (real warehouse) — manipulator transport.
  • ARMBench (real warehouse, Mitash et al., 2023) — diverse real failures.

Sparse / dense splits per AHA's protocol. Metrics: Detection Accuracy, LLM Fuzzy (semantic similarity, computed only on the reasoning text), ROUGE-L.

Model Dataset Detect Acc ↑ LLM Fuzzy ↑ ROUGE-L ↑
Qwen2.5-VL 7B RLBench 0.376 0.255 0.353
Sparrow 0.453 0.240 0.137
Maniskill 0.548 0.268 0.112
ARMBench 0.500 0.349 0.208
Cosmos-Reasoning RLBench 0.317 0.220 0.140
Sparrow 0.510 0.346 0.311
Maniskill 0.442 0.359 0.207
ARMBench 0.480 0.503 0.480
LLaVA-NeXT RLBench 0.067 0.346 0.032
Sparrow 0.500 0.008 0.078
Maniskill 0.010 0.286 0.042
ARMBench 0.500 0.000 0.067
Claude-3.7 RLBench 0.420 0.372 0.336
Sparrow 0.517 0.213 0.138
Maniskill 0.538 0.360 0.133
ARMBench 0.590 0.494 0.387
Claude-3.7 (3-shot) RLBench 0.561 0.473 0.526
Sparrow 0.650 0.407 0.458
Maniskill 0.625 0.398 0.212
ARMBench 0.650 0.685 0.725
SFT-D (dense only) RLBench 0.640 0.460 0.606
Sparrow 0.523 0.278 0.274
Maniskill 0.788 0.644 0.743
ARMBench 0.640 0.609 0.618
SFT-S+D (dense + sparse) RLBench 0.726 0.550 0.646
Sparrow 0.620 0.245 0.257
Maniskill (R→M) 0.490 0.177 0.381
ARMBench (S→A) 0.495 0.007 0.249
ARMOR (Ours) RLBench 0.917 0.718 0.802
Sparrow 0.733 0.503 0.514
Maniskill (R→M) 0.990 0.673 0.851
ARMBench (S→A) 0.725 0.698 0.721

Headline numbers from the paper:

  • Up to +30% over best prior methods on detection rate. (e.g. RLBench: 0.917 vs 0.640 SFT-D = +43%; Maniskill R→M: 0.990 vs 0.788 SFT-D = +25.6 absolute / +26%; ARMBench S→A: 0.725 vs 0.640 = +13%.)
  • Up to +100% improvement in reasoning (LLM fuzzy match score). On Sparrow vs SFT-D, reasoning rises from 0.278 to 0.503 = +81%; the headline +100% claim is paper-wide best-case.
  • Cross-distribution transfer (R→M and S→A) is a strict test: train sparse data and dense data come from different environments. SFT-S+D reasoning collapses (e.g. 0.644 → 0.177 in R→M) — naively mixing dense/sparse causes overfitting to sparse. ARMOR maintains 0.673 and 0.698.

Ablation: components (Table 2, RLBench)

Variant Offline Warmup Offline Expert Cond. Online Imitation Refinement Detection / Reasoning
Multitask Prediction ✓ ✗ ✗ ✗ 0.897 / 0.460
Refinement Only ✓ ✗ ✗ ✓ 0.803 / 0.488
Offline Imitation Only ✓ ✓ ✗ ✓ 0.853 / 0.658
Online Imitation Only ✗ ✗ ✓ ✓ 0.850 / 0.683
ARMOR (full) ✓ ✓ ✓ ✓ 0.917 / 0.718

Takeaways:

  • Multi-task architecture alone gives strong detection (0.897) but poor reasoning (0.460) — confirms the architectural separation works for classification but reasoning needs refinement training.
  • Refinement-only drops detection (0.897 → 0.803) — pure inference-time refinement without training is destabilizing.
  • Offline expert-conditioned imitation matters; online imitation matters more for reasoning.
  • All four together yield the best of both objectives.

Ablation: refinement rounds (Table 3, 4 seeds, RLBench)

Round Reasoning (LLM Fuzzy)
0 0.475 ± 0.016
1 0.676 ± 0.025 (+42.4%)
2 0.703 ± 0.013 (+4.0%)
3 0.717 ± 0.002 (+2.0%)

Variance shrinks across rounds (confirms convergence). One-round refinement already buys most of the gain.

Ablation: model scale (Table 8, Sparrow; Appendix C.1)

Model Detection / Reasoning
Qwen2.5-VL 7B 0.453 / 0.240
Qwen2.5-VL 32B 0.563 / 0.268
ARMOR 7B (ours) 0.733 / 0.503
ARMOR 32B (ours) 0.765 / 0.562

Approach scales with backbone size; ARMOR-7B already beats the Qwen2.5-VL-32B baseline by large margins, and ARMOR-32B improves further. (Paper: "We evaluated ARMOR using the 32B version of Qwen2.5-VL ... on Sparrow and observed consistent performance gains compared to the 7B version, as shown in Table 8.")

Inference cost (Table 4, 8 × A100 40GB)

Round Memory / GPU Wall Clock
0 (base) 3.48 GB 7.95 s
1 5.84 GB 9.30 s
2 6.12 GB 10.50 s
3 6.31 GB 10.95 s

≈ +1 s per refinement round. Authors note that one round buys 40% reasoning gain — practical sweet spot for deployment.

Qualitative refinement (Sec. 5.3, Appendix C.2)

  • Successful refinement: round 1 produces wrong reasoning + correct detection, round 2 corrects reasoning, round 3 polishes — illustrating reasoning being corrected to be consistent with the more reliable detection head.
  • Failure case (ARMBench): both rounds produce wrong reasoning ("collision" → "misplacement") despite correct detection ("No"). True failure was item damage. Authors flag this as motivation for richer intermediate failure-attribute supervision.

Limitations (as stated by authors)

  1. No external reward model used — leaves open whether reward shaping (task rewards, human preferences) could push reasoning quality further. Future work calls this out explicitly.
  2. Reasoning drift in refinement. When initial reasoning is wrong, refinement sometimes drifts to an alternative wrong explanation (collision → misplacement) rather than the true root cause. Structured failure attributes ("collision," "misplacement," "damage") as intermediate supervision proposed as fix.
  3. Modality scope. Vision + language only — no force-torque, proprioception, or audio. Authors call out integrating these as future work.
  4. Inference cost. Each refinement round adds both memory (3.48 → 6.31 GB/GPU over 3 rounds) and wall-clock time (≈ +1 s/round; Table 4). Authors frame this as minimal overhead over the base model, but +1 s/round is non-trivial for closed-loop monitoring at high rate.
  5. Broader-impact / safety. Authors caution against over-relying on imperfect reasoning models in safety-critical settings; ARMOR is meant to complement, not replace, human oversight and rigorous monitoring.

Significance & Positioning

Two ideas worth highlighting:

(i) Failure detection as iterative language-conditioned process — not a one-shot classifier. Borrows iterative-self-refinement (Madaan et al., 2023; Qu et al., 2024) from math/code reasoning and extends it to multi-modal video failure understanding. To the authors' knowledge, this is the first such extension.

(ii) Heterogeneous supervision design. Most prior failure-reasoning work assumes either binary-only (closed-set classifier) or full reasoning-everywhere (AHA / Duan et al., 2025). ARMOR explicitly engineers around the realistic regime where binary outcomes scale automatically (system logs) but reasoning is expensive to annotate. The training algorithm (offline warmup → offline expert-conditioned masking → online imitation rollouts) is the technical content.

Compared to neighbors:

  • AHA (Duan et al., 2025; ICLR 2025) — same problem, but assumes dense supervision and uses regex evaluation; breaks under cross-environment transfer (R→M, S→A) where ARMOR holds up.
  • DoReMi (Guo et al., 2024), REFLECT (Liu et al., 2023b) — closed-set or symbolic-abstraction failure detection. ARMOR is open-ended.
  • RISE (Qu et al., 2024) — closest in spirit (multi-turn self-refinement) but assumes oracle reward model. ARMOR uses internal entropy as the self-certainty signal.
  • Embodied-R1 / FROM-Seeing-To-Doing — also use VLM + reasoning for embodied tasks, but they output actions, not failure analyses. ARMOR is a complementary runtime monitor, not a policy.

The practical positioning: ARMOR is a candidate drop-in failure monitor for deployed VLAs (π0-class, OpenVLA, GR00T) — its self-certainty + no-oracle design means it can be fine-tuned in any environment with cheap binary signals plus a small reasoning corpus. As the field shifts toward continuous deployment of VLAs, this kind of introspectable, scalable monitor becomes infrastructure.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️