ICLR 2026 Self Refining VLM - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Embodied reasoning · Failure detection Trend tag: Self-refinement · Heterogeneous supervision · VLM monitor Affiliations: UT Austin (Qi, Zhang), Amazon Robotics (Wang, Sheng, Mao, Srinivasan, Nambi, Dattatreya), CMU (Yong)
flowchart LR
Vid[Video x<br/>multi-view frames] --> VLM[Qwen2.5-VL backbone<br/>vision encoder + LM decoder]
VLM --> CLS[Detection head<br/>BCE loss<br/>CLS token + cross-attn + MLP]
VLM --> NTP[Reasoning head<br/>NTP loss<br/>auto-regressive LM decoder]
CLS --> Round[Round t prediction lt]
NTP --> Round2[Round t prediction et]
Round --> Cond[Condition next round on lt-1, et-1, prompt p]
Round2 --> Cond
Cond --> VLM
Round --> Hdet[Detection entropy H_det]
Round2 --> Hexp[Reasoning mean-token entropy H_exp]
Hdet --> C[Combined entropy<br/>C = H_det + λ H_exp]
Hexp --> C
C --> Sel[Argmin over M trajectories<br/>Stop when no decrease in entropy]
Sel --> Out[Final detection l̂ + reasoning ê]
Real-world robot failures are subtle, combinatorial, and hard to enumerate. Closed-set classifiers under-cover them, and dense reasoning annotations are expensive. Concretely:
- Binary success / failure labels are easy to harvest (system logs, sensors) — large-scale, sparse.
- Free-form reasoning labels ("why did it fail") need human experts — small-scale, dense.
Prior work either (a) treats failure detection as closed-set classification (Ye et al., 2019; Garrett et al., 2020), (b) assumes dense reasoning annotations everywhere (AHA — Duan et al., 2025, applies SFT on full reasoning), or (c) prompts off-the-shelf VLMs without fine-tuning (Skreta et al., 2024; Guo et al., 2024). AHA's evaluation also relies on hand-crafted regex parsing, breaking under format drift.
ARMOR ("Adaptive Round-based Multi-task mOdel for Robotic failure detection and reasoning") tackles failure detection + reasoning under heterogeneous supervision, with open-ended reasoning beyond fixed taxonomies.
Following RISE (Qu et al., 2024), failure understanding is cast as a multi-task MDP M = (S, A, R, T):
- State
s_t = [x, l_{t-1}, e_{t-1}, p]— input video + previous round's detection + reasoning + auxiliary prompt. - Action
a_t = (l_t, e_t)— joint detection / reasoning prediction. - Rewards
R = {R_detect, R_reason}measure correctness against ground truth. - Horizon
T= number of refinement rounds.
Critically, ARMOR does not assume an oracle reward model at training or inference, unlike RISE. It uses heterogeneous supervision instead.
D_sparse = {(xi, li)}, li ∈ {success, failure}
D_dense = {(xi, li, ei)}, ei = detailed natural-language reasoning
Detection is supervised on D_sparse ∪ D_dense; reasoning only on D_dense. On D_sparse the model still generates reasoning as context but it is unsupervised.
Backbone: Qwen2.5-VL (Bai et al., 2025), 7B variant for main results ("all open-source VLMs and ARMOR use 7B model variants" unless otherwise specified); a 32B variant is additionally tested in the model-scale ablation (Appendix C.1, Table 8).
-
Detection head: lightweight classifier. Mean-pools features from the first 4 LM-decoder layers, projects them, cross-attends with a learnable
[CLS]token (8-head multi-head attention), MLP decoder → binary logits. Trained with binary cross-entropy (BCE). - Reasoning head: original Qwen2.5-VL auto-regressive LM decoder. Trained with next-token prediction (NTP).
- Both heads share the vision encoder and LM decoder backbone parameters.
- Round-
tconditioning is injected via natural-language prompts:"Given the previous detection is ..."/"Given the previous reasoning is ...".
Phase I — Offline Imitation:
-
Warm-up. Train
(l1, e1) ~ πθ(· | [x, ∅, ∅, ∅])onD_sparse ∪ D_densewith BCE forland NTP fore. (Empty prior context.) -
Expert-conditioned. Train
(l1, e1) ~ πθ(· | [x, l, e, p])onD_densewith same losses. To prevent the model from copying the prior input, one of the two ground-truth inputs is randomly masked per sample while the other is retained — forcing cross-task consistency learning.
Phase II — Online Refinement: For each minibatch, roll out the policy for T rounds:
for t = 1 to T:
(lt, et) ~ πθ(· | [x, l_{t-1}, e_{t-1}, p_t])
L += BCE(lt, l) + NTP(et, e) · 1[x ∈ D_dense]
θ ← θ − η ∇θ L
Detection is always supervised; reasoning is supervised only when dense labels exist. This is online imitation learning (Ross et al., 2011) — the model is supervised on its own rollout-state distribution.
At inference, ARMOR samples M refinement trajectories in parallel and selects the most confident:
C^(m) = H_det^(m) + λ · H_reason^(m) # combined entropy
m* = argmin_m C^(m) # most confident trajectory
Stop when C^(m*) ≥ C_min − ε # no longer reducing uncertainty
H_det is the detection logits' entropy; H_reason is the mean per-token entropy of the reasoning. λ = 0.1 (detection trusted more — supervised on more data). Final output is (l̂, ê) from the trajectory m* at termination.
| Hyperparameter | Value |
|---|---|
| Backbone | Qwen2.5-VL 7B (32B in scaling test, App. C.1) |
| Global batch size (videos: Sparrow, ARMBench) | 16 |
| Global batch size (collaged images: RLBench, Maniskill) | 64 |
| Offline epochs | 3 |
| Online epochs | 10 |
| Horizon T (rounds) | 3 (default) |
| Optimizer | AdamW |
| Weight decay | 0.1 |
| LR scheduler | Cosine, warmup ratio 0.03 |
| Learning rate (LM + classifier head) | 1 × 10⁻⁵ |
| Learning rate (vision encoder) | 2 × 10⁻⁶ |
| Refinement weight λ | 0.1 |
| Hardware | 8 × H100 GPUs (training) / 8 × A100 40GB (inference benchmarking) |
4 datasets:
- RLBench-Fail (sim) — tabletop tasks with scripted-policy perturbations.
- Maniskill-Fail (sim) — peg insertion, transport.
- Sparrow-Fail (real warehouse) — manipulator transport.
- ARMBench (real warehouse, Mitash et al., 2023) — diverse real failures.
Sparse / dense splits per AHA's protocol. Metrics: Detection Accuracy, LLM Fuzzy (semantic similarity, computed only on the reasoning text), ROUGE-L.
| Model | Dataset | Detect Acc ↑ | LLM Fuzzy ↑ | ROUGE-L ↑ |
|---|---|---|---|---|
| Qwen2.5-VL 7B | RLBench | 0.376 | 0.255 | 0.353 |
| Sparrow | 0.453 | 0.240 | 0.137 | |
| Maniskill | 0.548 | 0.268 | 0.112 | |
| ARMBench | 0.500 | 0.349 | 0.208 | |
| Cosmos-Reasoning | RLBench | 0.317 | 0.220 | 0.140 |
| Sparrow | 0.510 | 0.346 | 0.311 | |
| Maniskill | 0.442 | 0.359 | 0.207 | |
| ARMBench | 0.480 | 0.503 | 0.480 | |
| LLaVA-NeXT | RLBench | 0.067 | 0.346 | 0.032 |
| Sparrow | 0.500 | 0.008 | 0.078 | |
| Maniskill | 0.010 | 0.286 | 0.042 | |
| ARMBench | 0.500 | 0.000 | 0.067 | |
| Claude-3.7 | RLBench | 0.420 | 0.372 | 0.336 |
| Sparrow | 0.517 | 0.213 | 0.138 | |
| Maniskill | 0.538 | 0.360 | 0.133 | |
| ARMBench | 0.590 | 0.494 | 0.387 | |
| Claude-3.7 (3-shot) | RLBench | 0.561 | 0.473 | 0.526 |
| Sparrow | 0.650 | 0.407 | 0.458 | |
| Maniskill | 0.625 | 0.398 | 0.212 | |
| ARMBench | 0.650 | 0.685 | 0.725 | |
| SFT-D (dense only) | RLBench | 0.640 | 0.460 | 0.606 |
| Sparrow | 0.523 | 0.278 | 0.274 | |
| Maniskill | 0.788 | 0.644 | 0.743 | |
| ARMBench | 0.640 | 0.609 | 0.618 | |
| SFT-S+D (dense + sparse) | RLBench | 0.726 | 0.550 | 0.646 |
| Sparrow | 0.620 | 0.245 | 0.257 | |
| Maniskill (R→M) | 0.490 | 0.177 | 0.381 | |
| ARMBench (S→A) | 0.495 | 0.007 | 0.249 | |
| ARMOR (Ours) | RLBench | 0.917 | 0.718 | 0.802 |
| Sparrow | 0.733 | 0.503 | 0.514 | |
| Maniskill (R→M) | 0.990 | 0.673 | 0.851 | |
| ARMBench (S→A) | 0.725 | 0.698 | 0.721 |
Headline numbers from the paper:
- Up to +30% over best prior methods on detection rate. (e.g. RLBench: 0.917 vs 0.640 SFT-D = +43%; Maniskill R→M: 0.990 vs 0.788 SFT-D = +25.6 absolute / +26%; ARMBench S→A: 0.725 vs 0.640 = +13%.)
- Up to +100% improvement in reasoning (LLM fuzzy match score). On Sparrow vs SFT-D, reasoning rises from 0.278 to 0.503 = +81%; the headline +100% claim is paper-wide best-case.
- Cross-distribution transfer (R→M and S→A) is a strict test: train sparse data and dense data come from different environments. SFT-S+D reasoning collapses (e.g. 0.644 → 0.177 in R→M) — naively mixing dense/sparse causes overfitting to sparse. ARMOR maintains 0.673 and 0.698.
| Variant | Offline Warmup | Offline Expert Cond. | Online Imitation | Refinement | Detection / Reasoning |
|---|---|---|---|---|---|
| Multitask Prediction | ✓ | ✗ | ✗ | ✗ | 0.897 / 0.460 |
| Refinement Only | ✓ | ✗ | ✗ | ✓ | 0.803 / 0.488 |
| Offline Imitation Only | ✓ | ✓ | ✗ | ✓ | 0.853 / 0.658 |
| Online Imitation Only | ✗ | ✗ | ✓ | ✓ | 0.850 / 0.683 |
| ARMOR (full) | ✓ | ✓ | ✓ | ✓ | 0.917 / 0.718 |
Takeaways:
- Multi-task architecture alone gives strong detection (0.897) but poor reasoning (0.460) — confirms the architectural separation works for classification but reasoning needs refinement training.
- Refinement-only drops detection (0.897 → 0.803) — pure inference-time refinement without training is destabilizing.
- Offline expert-conditioned imitation matters; online imitation matters more for reasoning.
- All four together yield the best of both objectives.
| Round | Reasoning (LLM Fuzzy) |
|---|---|
| 0 | 0.475 ± 0.016 |
| 1 | 0.676 ± 0.025 (+42.4%) |
| 2 | 0.703 ± 0.013 (+4.0%) |
| 3 | 0.717 ± 0.002 (+2.0%) |
Variance shrinks across rounds (confirms convergence). One-round refinement already buys most of the gain.
| Model | Detection / Reasoning |
|---|---|
| Qwen2.5-VL 7B | 0.453 / 0.240 |
| Qwen2.5-VL 32B | 0.563 / 0.268 |
| ARMOR 7B (ours) | 0.733 / 0.503 |
| ARMOR 32B (ours) | 0.765 / 0.562 |
Approach scales with backbone size; ARMOR-7B already beats the Qwen2.5-VL-32B baseline by large margins, and ARMOR-32B improves further. (Paper: "We evaluated ARMOR using the 32B version of Qwen2.5-VL ... on Sparrow and observed consistent performance gains compared to the 7B version, as shown in Table 8.")
| Round | Memory / GPU | Wall Clock |
|---|---|---|
| 0 (base) | 3.48 GB | 7.95 s |
| 1 | 5.84 GB | 9.30 s |
| 2 | 6.12 GB | 10.50 s |
| 3 | 6.31 GB | 10.95 s |
≈ +1 s per refinement round. Authors note that one round buys 40% reasoning gain — practical sweet spot for deployment.
- Successful refinement: round 1 produces wrong reasoning + correct detection, round 2 corrects reasoning, round 3 polishes — illustrating reasoning being corrected to be consistent with the more reliable detection head.
- Failure case (ARMBench): both rounds produce wrong reasoning ("collision" → "misplacement") despite correct detection ("No"). True failure was item damage. Authors flag this as motivation for richer intermediate failure-attribute supervision.
- No external reward model used — leaves open whether reward shaping (task rewards, human preferences) could push reasoning quality further. Future work calls this out explicitly.
- Reasoning drift in refinement. When initial reasoning is wrong, refinement sometimes drifts to an alternative wrong explanation (collision → misplacement) rather than the true root cause. Structured failure attributes ("collision," "misplacement," "damage") as intermediate supervision proposed as fix.
- Modality scope. Vision + language only — no force-torque, proprioception, or audio. Authors call out integrating these as future work.
- Inference cost. Each refinement round adds both memory (3.48 → 6.31 GB/GPU over 3 rounds) and wall-clock time (≈ +1 s/round; Table 4). Authors frame this as minimal overhead over the base model, but +1 s/round is non-trivial for closed-loop monitoring at high rate.
- Broader-impact / safety. Authors caution against over-relying on imperfect reasoning models in safety-critical settings; ARMOR is meant to complement, not replace, human oversight and rigorous monitoring.
Two ideas worth highlighting:
(i) Failure detection as iterative language-conditioned process — not a one-shot classifier. Borrows iterative-self-refinement (Madaan et al., 2023; Qu et al., 2024) from math/code reasoning and extends it to multi-modal video failure understanding. To the authors' knowledge, this is the first such extension.
(ii) Heterogeneous supervision design. Most prior failure-reasoning work assumes either binary-only (closed-set classifier) or full reasoning-everywhere (AHA / Duan et al., 2025). ARMOR explicitly engineers around the realistic regime where binary outcomes scale automatically (system logs) but reasoning is expensive to annotate. The training algorithm (offline warmup → offline expert-conditioned masking → online imitation rollouts) is the technical content.
Compared to neighbors:
- AHA (Duan et al., 2025; ICLR 2025) — same problem, but assumes dense supervision and uses regex evaluation; breaks under cross-environment transfer (R→M, S→A) where ARMOR holds up.
- DoReMi (Guo et al., 2024), REFLECT (Liu et al., 2023b) — closed-set or symbolic-abstraction failure detection. ARMOR is open-ended.
- RISE (Qu et al., 2024) — closest in spirit (multi-turn self-refinement) but assumes oracle reward model. ARMOR uses internal entropy as the self-certainty signal.
- Embodied-R1 / FROM-Seeing-To-Doing — also use VLM + reasoning for embodied tasks, but they output actions, not failure analyses. ARMOR is a complementary runtime monitor, not a policy.
The practical positioning: ARMOR is a candidate drop-in failure monitor for deployed VLAs (π0-class, OpenVLA, GR00T) — its self-certainty + no-oracle design means it can be fine-tuned in any environment with cheap binary signals plus a small reasoning corpus. As the field shifts toward continuous deployment of VLAs, this kind of introspectable, scalable monitor becomes infrastructure.
- OpenReview: https://openreview.net/forum?id=jr9hGWQioP
- PDF: https://openreview.net/pdf?id=jr9hGWQioP
- Project page: https://sites.google.com/utexas.edu/armor
← Back to ICLR-2026