ICML 2026 LAGEA - Heungwoo/research GitHub Wiki

LAGEA: Language Guided Embodied Agents for Robotic Manipulation — Turning VLM self-reflections into time-grounded shaping rewards

Venue: ICML 2026 (Poster) Category: RL for VLA Affiliations: University of Dhaka, Bangladesh Traction (2026-06): 1 citation (arXiv)

Overview of the LaGEA framework: keyframe selection, VLM self-reflection, visual-language alignment, and delta-based shaping rewards (Figure 1 from Chowdhury et al., 2026)

Problem

Robotic manipulation increasingly benefits from foundation models that describe goals, but agents still lack a principled way to learn from their own mistakes. Sparse-reward, long-horizon tasks make exploration brittle: dense reward shaping from vision-language models (VLMs) can destabilize training or invite reward hacking, while contrastive reward-alignment approaches such as FuRL can suffer when early misalignment compounds and misdirects exploration. LaGEA asks whether natural language can instead serve as an error-reasoning signal — feedback that helps an embodied agent diagnose what went wrong and correct course.

Method

LaGEA (Language Guided Embodied Agents) turns episodic, schema-constrained reflections from a VLM into temporally grounded guidance for reinforcement learning. The pipeline has four stages:

  1. Keyframe selection — after each rollout, causal moments in the trajectory are identified and per-step weights $\hat{w}_t$ are computed, localizing the decisive frames.
  2. Schema-constrained self-reflection — a VLM is queried on those frames and returns a concise, structured language summary of what happened (an error taxonomy constrains the format).
  3. Visual-language alignment — feedback is aligned with visual state in a shared representation, co-trained with BCE / InfoNCE objectives so the embedding space becomes control-relevant.
  4. Delta-based shaping rewards — a Goal Potential $\phi_t$ aligns the current state $z_t$ with the goal image $z_g$ and instruction $z_y$, and a Feedback Potential $\psi_t$ measures agreement with the reflection. These are converted into bounded, step-wise shaping rewards.

Delta-based reward construction from Goal Potential and Feedback Potential (Figure 2 from Chowdhury et al., 2026)

The combined shaping term is modulated by an adaptive, failure-aware coefficient $\hat{\rho}_t = m_t \rho_t$, so signals are dense early when exploration needs direction and gracefully recede as competence grows. SAC is finally trained on $r_t = r_t^{\text{task}} + \hat{\rho}_t, \tilde{r}_t$, preserving the task objective while adding grounded guidance.

Results

On the Meta-World MT10 embodied manipulation benchmark (average success across five random seeds), LaGEA improves average success over state-of-the-art methods by 9.0% on random goals and 5.3% on fixed goals, while converging faster. Baselines include SAC, LIV, LIV-Proj, Relay, and FuRL (with and without goal image). Across eight Meta-World tasks (Figure 3), LaGEA reaches high success in far fewer environment steps than FuRL and SAC, which plateau late or stall. Ablations confirm that (a) structured feedback beats free-form feedback, (b) keyframe selection matters (drawer-open study), and (c) removing any of $r_t^{\text{goal}}$, $r_t^{\text{fb}}$, or $\rho$ causes a significant performance drop. Alignment analysis shows the success/failure logit margin grows over training as the shared space is co-trained.

Significance

LaGEA supports the hypothesis that language, when structured and grounded in time, is an effective mechanism for teaching robots to self-reflect on mistakes. Rather than using a VLM as a one-shot reward labeler, it converts episodic reflections into bounded, decaying, failure-aware shaping signals — a recipe that improves both sample efficiency and final success on sparse-reward manipulation.

Links

← Back to ICML-2026