ICLR 2026 Embodied R1 - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Training / RL Trend tag: Trends 2 + 3
flowchart LR
Img[Image] --> VLM[Pointing VLM]
Inst[Instruction] --> VLM
VLM --> R{Pointing primitives}
R --> REG[REG: refer-to-object]
R --> RRG[RRG: refer-to-relation]
R --> OFG[OFG: refer-to-functional-part]
R --> VTG[VTG: visual trace]
R --> RFT[Two-stage Reinforced Fine-Tuning<br/>on Embodied-Points-200K]
RFT --> Pol[Policy with embodiment-agnostic<br/>intermediate representation]
Chain-of-thought reasoning is typically learned via supervised fine-tuning on human-written reasoning traces โ expensive and brittle. Reinforcement-based reasoning (as in R1-style LLMs) may work better but had not been applied to embodied tasks.
A 3B pointing VLM (built on Qwen2.5-VL-3B-Instruct) trained via Reinforced Fine-Tuning (RFT). Intermediate representations are four pointing abilities:
- REG โ Referring Expression Grounding (localize an object from a linguistic description)
- RRG โ Region Referring Grounding (identify a spatial region from relational language)
- OFG โ Object Functional Grounding (identify a functional part / affordance, e.g., handle)
- VTG โ Visual Trace Generation (output an ordered point sequence forming a manipulation trajectory)
These primitives are embodiment-agnostic. Training uses a two-stage RFT curriculum with a multi-task reward design on the new Embodied-Points-200K dataset (~200k samples across the four pointing capabilities).
- State-of-the-art on 11 embodied spatial and pointing benchmarks.
- Zero-shot generalization (no task-specific fine-tuning): 56.2% average success in SimplerEnv and 87.5% across 8 real-world XArm tasks, a 62% improvement over strong baselines on real-world manipulation.
Establishes that R1-style RL can be applied to embodied reasoning, not just textual reasoning. The pointing-primitive vocabulary is a promising candidate for the transferable interface that Trend 2 calls for.
- ICLR 2026 listing