CoRL 2025 ECoT Lite - Heungwoo/research GitHub Wiki
Venue: CoRL 2025 ยท Authors: William Chen, Suneel Belkhale, Suvir Mirchandani, Oier Mees, Danny Driess, Karl Pertsch, Sergey Levine โ UC Berkeley, Stanford, Physical Intelligence ยท arXiv: 2505.08243 Category: Training Approach Trend tag: Lightweight embodied chain-of-thought
flowchart LR
I[Image + instruction] --> VLA[VLA]
VLA -- optional traces --> PLAN[Plan]
VLA -- optional --> SUB[Subtask]
VLA -- optional --> GRIP[Gripper pose]
VLA -- optional --> BBOX[Object bbox]
VLA --> ACT[Action]
note[Ablation: which CoT<br/>pieces actually matter?] -.-> PLAN & SUB & GRIP & BBOX
Full embodied chain-of-thought (ECoT) supervision โ plan + subtask + gripper pose + bounding box + action โ works, but training and inference costs balloon. Nobody had cleanly decomposed which CoT components actually drive the downstream gains and which are decorative.
Rather than only ablating CoT sub-components, the paper isolates why embodied reasoning helps by testing three hypothesized mechanisms with controlled training variants:
- Better representation learning โ generating reasoning during training embeds task-relevant features in the model, useful even when reasoning is not produced at test time.
- Improved curricularization โ training-time reasoning signals indicate which features matter for action selection.
- Increased expressivity โ extra in-context tokens raise model expressivity regardless of semantic content.
From this analysis the authors derive two lightweight strategies, collectively ECoT-Lite:
- Reasoning dropout โ generate reasoning during training but randomly drop it, so reasoning can be turned off at test time with no inference overhead.
- Reasoning pre-training โ pre-train on reasoning only, then fine-tune on action prediction only (no paired reasoning-action data required).
Experiments use a MiniVLA backbone for LIBERO-90 and OpenVLA-based policies for real-world Bridge/WidowX tasks.
Key finding: learning to generate reasonings yields better VLA representations, while attending to those reasonings lets the policy exploit those features for action prediction โ both are load-bearing.
LIBERO-90 success rates:
- Full ECoT: 90.8%
- Reasoning dropout: 89.4%
- Reasoning pre-training: 87.1%
- Prior SOTA: 88.6%
ECoT-Lite (dropout and pre-training) gives significant gains over non-reasoning policies and a 3ร inference speedup vs. standard robot reasoning, validated on real-world Bridge WidowX manipulation. Reasoning pre-training notably requires no paired reasoning-action data.
Makes embodied reasoning cheap enough to be default. Threads directly into the ICLR 2026 reasoning-augmented VLA wave:
โ Back to CoRL-2025