ICLR 2026 REI Bench - Heungwoo/research GitHub Wiki

REI-Bench โ€” Vague Referring Expressions in Embodied Task Planning

Full title: REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning? Venue: ICLR 2026 ยท OpenReview: vmBIF25KLf ยท arXiv: 2505.10872 Authors: Chenxi Jiang, Chuhao Zhou, Jianfei Yang (MARS Lab, School of Mechanical & Aerospace Engineering, Nanyang Technological University) Category: Benchmark โ€” Instruction following / task planning Trend tag: Evaluation / pragmatics for embodied LLM planners

Approach diagram

flowchart LR
  User[Vague human instruction<br/>referring expression] --> Plan[LLM task planner]
  Plan --> Fail[Failure mode:<br/>missing objects in plan]
  User --> TOCC[Task-oriented<br/>context cognition]
  TOCC --> Clear[Clarified instruction]
  Clear --> Plan
Loading

Problem

LLM-based robot task planners are typically evaluated on clean, well-specified instructions, but real users โ€” especially non-experts, the elderly, and children โ€” produce vague referring expressions ("that thing over there", "the one I used yesterday"). No benchmark systematically measures how planners degrade under this vagueness.

Method

REI-Bench is the first benchmark that systematically models vague referring expressions grounded in pragmatic theory for embodied task planning. It is built on ALFRED, selecting six household tasks (Pick & Place, Stack & Place, Clean & Place, Heat & Place, Cool & Place, Examine in Light), filtered to instances solvable with clear instructions by LLaMA3.1-8B + SayCan.

Instructions are organized into three levels of referential difficulty by the ratio of explicit to implicit REs:

  • Explicit REs โ€” original expressions preserved (e.g., "potato").
  • Mixed REs โ€” task instruction uses implicit REs while the dialogue context still carries explicit ones.
  • Implicit REs โ€” all expressions replaced by pronouns/descriptors (e.g., "the heated one").

Evaluation spans LLMs (GPT-4o-mini, LLaMA3.1-8B, Ministral-8B, Gemma2-9B, DeepSeek-Math-7B, Qwen2.5-7B) across planning frameworks (SayCan, DAG-Plan, HPE, LLM+P).

The authors also propose task-oriented context cognition, a prompting/reasoning approach that rewrites vague user input into clearer instructions before planning.

Results

Vagueness in referring expressions causes success-rate drops of up to 36.9% for current LLM planners; analysis attributes most failures to missing objects in the generated plan. Task-oriented context cognition achieves SOTA versus aware-prompting, chain-of-thought, and in-context learning baselines (specific final-success numbers not stated in abstract).

Significance

Reframes embodied-LLM evaluation around pragmatic robustness rather than literal-instruction following โ€” a prerequisite for deployment to non-expert users. Sits alongside RoboArena in the 2026 push toward more realistic embodied evaluation.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ