ICLR 2026 Hybrid Training - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 (poster) Category: Training Approach Trend tag: Trend 2 Authors: Pietro Mazzaglia, Cansu Sancaktar, Markus Peschl, Daniel Dijkman (Qualcomm AI Research; Sancaktar at University of Tübingen / Max Planck Institute for Intelligent Systems, internship at Qualcomm) arXiv: 2510.00600 · OpenReview: IBJtOltTbx
flowchart LR
subgraph Train[Training time]
MC[Monte Carlo objective:<br/>sample conditional inputs/outputs<br/>indexed by a modality variable]
MC --> CoT[think: actions conditioned on thoughts]
MC --> NoCoT[act: actions directly]
MC --> Follow[follow: actions given oracle instructions]
CoT --> M[Single hybrid model]
NoCoT --> M
Follow --> M
end
subgraph Deploy[Deployment]
M --> Mode[Set modality variable]
Mode --> Fast[Default 'act' mode:<br/>direct actions, no CoT<br/>retains training gain at VLA speed]
end
Embodied chain-of-thought (ECoT) improves performance but increases inference latency — generating intermediate thoughts before every action lowers the action-inference frequency, hurting usability in real-time robotic manipulation. The paper's framing question: is generating long chains-of-thought a strong prerequisite for the performance improvement, or can the model benefit from CoT at training time and skip it at inference?
Hybrid Training (HyT) trains a single VLA on a mix of CoT and pure-action data. The objective is a Monte Carlo estimate that samples a variety of conditional inputs/outputs with different probabilities, governed by a newly introduced modality variable. At test time the modality variable selects the behavior: act (default — predict actions directly, no thoughts, matching standard-VLA inference time), think (generate ECoT-style intermediate thoughts), or follow (act on externally provided instructions, e.g. from an oracle/human, as in hierarchical systems). The motivating hypothesis is that learning from CoT traces lets the model develop "skilled intuition" — internalizing reasoning knowledge into its representation so performance improves even when no CoT is generated at inference.
Note: HyT is architecture-agnostic and does not propose decomposing ECoT into subtasks as its mechanism (subtask labels are only used to extract thought annotations in some benchmarks). Backbone in experiments is PaliGemma-2 (3B) (Gemma-2 LLM + SigLIP encoder), with actions tokenized into 256 discrete bins; for LIBERO it is combined with the VLA-OFT recipe (action chunking + continuous L1 action head, fine-tuned from Prismatic VLM).
Evaluated on ClevrSkills (9 tasks, up to 3000 demos), LIBERO (4 suites), and a real-world UFactory xArm 6.
- ClevrSkills: HyT outperforms standard VLA, ECoT, and HiRobot-like hierarchical training at all data scales (300/750/1500/3000 demos); ECoT is the second-best baseline. The gain is largest on complex, longer-horizon tasks.
- LIBERO: HyT averages 93.7 (Spatial 94.0, Object 97.2, Goal 96.2, Long 89.4), edging out the VLA-OFT recipe it builds on (92.1 avg) and other reasoning VLAs (CoT-VLA 81.1, ThinkAct 84.4, MolmoAct 86.6), with the biggest gains on the hardest Goal/Long suites.
- Latency: standard VLA and HyT (act mode) run at ~3 Hz on A100 (4 models in parallel); ECoT is 3× slower and HiRobot hierarchical generation 4× slower. HyT thus keeps full-VLA speed while retaining the CoT training benefit.
- Other modalities: following oracle thoughts ("follow") improves all methods; without oracle thoughts, HyT in act vs think mode performs similarly, supporting the claim that explicit test-time CoT is unnecessary in these settings.
Empirically argues that one main utility of CoT-based training is improving the learned representation, enabling better performance even with no CoT generated at inference — addressing the latency vs performance trade-off for ECoT. HyT is complementary to dVLA (which instead makes ECoT cheap via parallel discrete diffusion). The paper positions itself relative to concurrent work (DualFormer-style reasoning dropout; Chen et al. 2025 on pre-training/co-training/dropout) by training a single hybrid system that can conditionally emit any of act/think/follow.
- arXiv 2510.00600
- OpenReview IBJtOltTbx (ICLR 2026 poster)
- dVLA (parallel-decoding alternative)
- Embodied-R1 (RL-trained reasoning)
- Actions as Language (subtask relabeling)
← Back to ICLR-2026