ICLR 2026 OneTwoVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture โ Reasoning / Dual-System Trend tag: Adaptive reasoning ยท System-1/System-2 unification Affiliation: Tsinghua University ยท Shanghai Qi Zhi Institute ยท Shanghai AI Lab ยท Fudan University ยท Spirit AI
flowchart LR
Obs[Multi-camera obs I_t<br/>+ reference frames I_ref] --> M[Single ฯ0-based VLA]
Inst[Instruction l] --> M
R[Latest reasoning content R<br/>scene + plan + history + next step] --> M
M --> DT{Decision token}
DT -- BOR --> Reason[Reasoning mode<br/>auto-regressive text<br/>cross-entropy loss]
DT -- BOA --> Act[Acting mode<br/>action chunk<br/>flow-matching loss]
Reason -. updates R .-> M
Act --> Env[Robot]
"Dual-system" VLAs split high-level reasoning (System-2 VLM) and low-level acting (System-1 VLA) across two separate models. The authors argue this design has two intrinsic flaws: (1) the systems lack mutual awareness of each other's capabilities โ System-2 routinely emits sub-goals System-1 cannot execute (e.g. "add green onion" when none is on the table); (2) System-2 latency means reasoning arrives after the world has already moved on. Yet a flat VLA without reasoning cannot track long-horizon progress. OneTwoVLA argues both modes should live in one model that decides on its own when to reason explicitly and when to act on the most recent reasoning.
A single policy ฯ_ฮธ operates in two modes:
- Reasoning mode: input is multi-camera images I_t^{1:n}, reference images I_ref^{1:n} (the observation taken when R was last produced), language instruction l, and the latest reasoning content R. Output is updated reasoning text Rฬ.
- Acting mode: additionally takes proprioceptive state s_t and outputs an action chunk A_t conditioned on the most recent R.
Two special decision tokens are introduced โ [BOR] (beginning of reasoning) and [BOA] (beginning of action). Each step the model first predicts one of them; this gates whether it expands into a text reasoning trace (terminated by [EOS]) or directly emits an action chunk. Algorithm 1 of the paper specifies this loop.
The text emitted in reasoning mode is structured:
- Scene description D_j โ locations of task-relevant objects.
- High-level plan P = (p_1, โฆ, p_K) โ ordered sub-tasks.
- Historical summary H_j = (p_1, โฆ, p_j) โ what has been done.
- Next step X_j = p_{j+1} โ immediate sub-task.
Reference frames I_ref are included alongside I_t so the policy doesn't get confused when textual scene descriptions become stale (e.g. "object at right of gripper" after the gripper has moved).
Each demonstration is split into reasoning intervals (around sub-task transitions, error states, or human-interaction moments โ annotated with reasoning content R) and acting intervals (everything else). During training:
- In a reasoning interval, the model learns to predict
[BOR]then Rฬ when current R is stale, then[BOA]and actions once R is updated. It also learns to predict actions from R without supervising[BOA], to stay safe if it fails to update reasoning at deployment. - In an acting interval, it learns to predict
[BOA]and actions. Optionally it also learns to predict[BOR]on outdated reasoning to fix the imbalanced binary-classification problem (acting intervals are much longer than reasoning intervals).
A two-stage pipeline labels real demonstrations: (1) Interval annotation โ Gemini 2.5 reads N=32 uniformly subsampled frames, locates K+1 reasoning intervals after each pre-defined sub-task; (2) Content generation โ for each interval, Gemini generates D_j from the midpoint frame, fills in P, H_j, X_j. On the Tomato-Egg task, 81.5% of intervals were judged correct by humans and 83.3% of scene descriptions reasonable.
To get generalization beyond the small set of robot demonstrations, the authors synthesize embodied-reasoning VL data: (1) Gemini 2.5 Pro proposes diverse tabletop-layout text descriptions; (2) FLUX.1-dev renders these as images, with random fisheye distortion or pasted gripper for realism; (3) Gemini generates instructions (visual-grounding via spatial/semantic/attribute references, plus long-horizon plans) and the matching reasoning content. 16,000 images total โ 6,000 used for visual-grounding tasks and 10,000 for long-horizon planning tasks.
Base VLA is ฯ0 (Black et al., 2024). The PaliGemma-style VLM auto-regressively emits reasoning text supervised by cross-entropy. The action expert is the same as ฯ0 and trained with flow-matching loss. Two input modifications: include reference image I_ref alongside current I_t, and stack proprioceptive states from t-0.05 s and t-0.25 s for smoother action prediction.
- Hardware/training: 30,000 steps per task on 8รH100 GPUs, ~10 hours. ฯ0 hyperparameters are inherited.
- Robot platforms: primary single-arm 7-DoF Franka with parallel-jaw gripper and wrist-mounted GoPro fisheye; dual-arm setup uses two 6-DoF ARX arms with two wrist + one base camera.
- Inference: temporal-ensemble action decoding (action chunks regenerated every 0.2 s, exponentially weighted); on a 4090 GPU with two image inputs, acting-mode forward pass is โค 0.2 s (ฯ0 0.082 s; OneTwoVLA-Act โ 0.102โ0.104 s); reasoning-mode latency varies with output length (Table 14 of paper: 0.853 s for 20 output tokens, 2.346 s for 100, 4.361 s for 200).
| Task | ฯ0 (flat VLA) | Dual-System (Hi-Robot-style, Gemini 2.5 Pro) | OneTwoVLA |
|---|---|---|---|
| Tomato-Egg | 70% | 55% | 85% |
| Hotpot | 50% | 70% | 80% |
| Cocktail | 50% | 65% | 95% |
| Average | 57% | 63% | 87% |
OneTwoVLA outperforms ฯ0 by +30% and the dual-system baseline by +24% average.
Total task completion time on Tomato-Egg: OneTwoVLA 184 s (16 s reasoning + 168 s acting) vs Dual-System 260 s (109 s reasoning + 151 s acting) vs ฯ0 176 s (no reasoning). OneTwoVLA matches the flat-VLA wall-clock while having reasoning. On Mountain Fuji cocktail, reasoning is invoked 5 times totalling 14 s out of 135 s (10.4%).
| Task | ฯ0 | OneTwoVLA | OneTwoVLA-VL (with 16K synthetic) |
|---|---|---|---|
| Get Icy Cola (proactively open fridge) | 5% | 5% | 70% |
| Empty Plate (occlusion handling) | 0% | 5% | 75% |
| Tool Use (use stick to sweep) | 0% | 10% | 65% |
| Prepare Drinks (intent inference) | 25% | 15% | 80% |
| Average | 7.5% | 8.75% | 72.5% |
| Task | ฯ0 | Dual-System | OneTwoVLA |
|---|---|---|---|
| Hotpot recoveries | 3/7 | 4/5 | 5/6 |
| Tomato-Egg recoveries | 5/7 | 3/7 | 3/4 |
| Total | 57.1% | 58.3% | 80.0% (8/10) |
200 dedicated recovery demonstrations were collected for Hotpot (out of 600 total) and 100 for Tomato-Egg (out of 200 total).
| Task | Dual-System | OneTwoVLA |
|---|---|---|
| Hotpot | 8/10 | 10/10 |
| Cocktail | 5/10 | 10/10 |
| Total | 65% | 100% |
Generalizable interaction (20 unseen scenarios): OneTwoVLA-VL 72.5% vs OneTwoVLA without VL co-training "fails to interpret" novel commands.
| Setting | ฯ0 | OneTwoVLA | OneTwoVLA-VL |
|---|---|---|---|
| Single-Env (4 objects, 200 UMI demos) | 5% | 78% | 88% |
| Open-World (16 envs, 933 demos, 8 unseen envs) | 3% | 8% | 73% |
ฯ0's action MSE on the validation set is 62% higher than OneTwoVLA's, suggesting reasoning even helps fit the action distribution.
-
Co-training with synthetic VL data: dropping it (
OneTwoVLAvsOneTwoVLA-VL) collapses generalizable planning from 72.5% to 8.75% and Open-World grounding from 73% to 8%. -
Decision-token supervision in acting intervals (optional
[BOR]learning): added to compensate for class imbalance โ without it, the model rarely enters reasoning mode at deployment. - Reference image input I_ref: included specifically because text scene descriptions go stale; ablation discussion in Sec. F.2 attributes smoother actions to this and the multi-timestep proprioception input.
- Action MSE comparison: explicit reasoning lowers action prediction MSE 1.62ร over ฯ0 on the validation split.
(The paper does not include a separate "scaling-of-VL-data" ablation curve; only the binary 0-vs-16K comparison is reported.)
- Heuristic reasoning placement. The decision of which steps need reasoning is currently labeled by hand-crafted heuristics. Authors suggest RL-based optimization of when to reason as future work โ analogous to going from SFT to RL in LLMs.
- Reasoning pause. Although reasoning fires at only a few steps, the robot still pauses 2โ3 s while a long reasoning trace is generated. Asynchronous architectures that allow simultaneous reasoning + acting are flagged as future work.
- Action inference not optimized. Standard ฯ0 inference; as the unified model scales up, action latency may dominate. LLM-style optimization (KV cache, speculative decoding) is suggested.
- VL data sources. Only synthetic data was studied. Effects of real-world VL data (existing instruction-tuning sets, web data) on VLA reasoning are unexplored.
-
vs ฯ0.5 / Helix / Gemini Robotics dual-system designs: ฯ0.5 inserts a fixed sub-task prediction before each action; Hi-Robot, ViLa run a separate VLM as System-2 at fixed cadence. OneTwoVLA shows that adaptive
[BOR]/[BOA]switching inside one model beats both โ same wall-clock as a flat ฯ0, +30% absolute success on long-horizon tasks. - vs ECoT / ECoT-Lite: ECoT generates reasoning at every step (slow); ECoT-Lite suppresses test-time reasoning (loses interpretability and HRI). OneTwoVLA gets ECoT-style benefits at flat-VLA cost.
- vs OpenVLA / GR00T N1: these flat VLAs are explicitly the kind of "no reasoning" baseline the paper outperforms; ฯ0 here serves as the closest flat-VLA reference and gets +30%.
- vs reasoning-enriched VLAs in 2026: aligns with Vlaser, InstructVLA, Embodied R1 in believing reasoning matters, but is the cleanest "learned adaptive boundary" instantiation. The synthetic VL co-training pipeline (16K samples) is a particularly transferable contribution.
- Debuggability: because reasoning is text, every failure can be inspected. Aligns with the broader 2026 trend of MemoryVLA making VLAs introspectable.
- OpenReview: https://openreview.net/forum?id=tWMfhoP3as
- Project page: https://one-two-vla.github.io/
- Vlaser (synergistic embodied reasoning)
- InstructVLA
- Embodied R1
- HybridVLA (different "unified" angle: AR + diffusion)
- MemoryVLA (introspectable text memory, kindred spirit)
โ Back to ICLR-2026