CoRL 2025 ECoT Lite - Heungwoo/research GitHub Wiki

ECoT-Lite โ€” Training Strategies for Efficient Embodied Reasoning

Venue: CoRL 2025 ยท Authors: William Chen, Suneel Belkhale, Suvir Mirchandani, Oier Mees, Danny Driess, Karl Pertsch, Sergey Levine โ€” UC Berkeley, Stanford, Physical Intelligence ยท arXiv: 2505.08243 Category: Training Approach Trend tag: Lightweight embodied chain-of-thought

Approach diagram

flowchart LR
  I[Image + instruction] --> VLA[VLA]
  VLA -- optional traces --> PLAN[Plan]
  VLA -- optional --> SUB[Subtask]
  VLA -- optional --> GRIP[Gripper pose]
  VLA -- optional --> BBOX[Object bbox]
  VLA --> ACT[Action]
  note[Ablation: which CoT<br/>pieces actually matter?] -.-> PLAN & SUB & GRIP & BBOX
Loading

Problem

Full embodied chain-of-thought (ECoT) supervision โ€” plan + subtask + gripper pose + bounding box + action โ€” works, but training and inference costs balloon. Nobody had cleanly decomposed which CoT components actually drive the downstream gains and which are decorative.

Method

Rather than only ablating CoT sub-components, the paper isolates why embodied reasoning helps by testing three hypothesized mechanisms with controlled training variants:

  1. Better representation learning โ€” generating reasoning during training embeds task-relevant features in the model, useful even when reasoning is not produced at test time.
  2. Improved curricularization โ€” training-time reasoning signals indicate which features matter for action selection.
  3. Increased expressivity โ€” extra in-context tokens raise model expressivity regardless of semantic content.

From this analysis the authors derive two lightweight strategies, collectively ECoT-Lite:

  • Reasoning dropout โ€” generate reasoning during training but randomly drop it, so reasoning can be turned off at test time with no inference overhead.
  • Reasoning pre-training โ€” pre-train on reasoning only, then fine-tune on action prediction only (no paired reasoning-action data required).

Experiments use a MiniVLA backbone for LIBERO-90 and OpenVLA-based policies for real-world Bridge/WidowX tasks.

Results

Key finding: learning to generate reasonings yields better VLA representations, while attending to those reasonings lets the policy exploit those features for action prediction โ€” both are load-bearing.

LIBERO-90 success rates:

  • Full ECoT: 90.8%
  • Reasoning dropout: 89.4%
  • Reasoning pre-training: 87.1%
  • Prior SOTA: 88.6%

ECoT-Lite (dropout and pre-training) gives significant gains over non-reasoning policies and a 3ร— inference speedup vs. standard robot reasoning, validated on real-world Bridge WidowX manipulation. Reasoning pre-training notably requires no paired reasoning-action data.

Significance

Makes embodied reasoning cheap enough to be default. Threads directly into the ICLR 2026 reasoning-augmented VLA wave:

Links

Related pages

โ† Back to CoRL-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ