ICLR 2026 REI Bench - Heungwoo/research GitHub Wiki
Full title: REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning? Venue: ICLR 2026 ยท OpenReview: vmBIF25KLf ยท arXiv: 2505.10872 Authors: Chenxi Jiang, Chuhao Zhou, Jianfei Yang (MARS Lab, School of Mechanical & Aerospace Engineering, Nanyang Technological University) Category: Benchmark โ Instruction following / task planning Trend tag: Evaluation / pragmatics for embodied LLM planners
flowchart LR
User[Vague human instruction<br/>referring expression] --> Plan[LLM task planner]
Plan --> Fail[Failure mode:<br/>missing objects in plan]
User --> TOCC[Task-oriented<br/>context cognition]
TOCC --> Clear[Clarified instruction]
Clear --> Plan
LLM-based robot task planners are typically evaluated on clean, well-specified instructions, but real users โ especially non-experts, the elderly, and children โ produce vague referring expressions ("that thing over there", "the one I used yesterday"). No benchmark systematically measures how planners degrade under this vagueness.
REI-Bench is the first benchmark that systematically models vague referring expressions grounded in pragmatic theory for embodied task planning. It is built on ALFRED, selecting six household tasks (Pick & Place, Stack & Place, Clean & Place, Heat & Place, Cool & Place, Examine in Light), filtered to instances solvable with clear instructions by LLaMA3.1-8B + SayCan.
Instructions are organized into three levels of referential difficulty by the ratio of explicit to implicit REs:
- Explicit REs โ original expressions preserved (e.g., "potato").
- Mixed REs โ task instruction uses implicit REs while the dialogue context still carries explicit ones.
- Implicit REs โ all expressions replaced by pronouns/descriptors (e.g., "the heated one").
Evaluation spans LLMs (GPT-4o-mini, LLaMA3.1-8B, Ministral-8B, Gemma2-9B, DeepSeek-Math-7B, Qwen2.5-7B) across planning frameworks (SayCan, DAG-Plan, HPE, LLM+P).
The authors also propose task-oriented context cognition, a prompting/reasoning approach that rewrites vague user input into clearer instructions before planning.
Vagueness in referring expressions causes success-rate drops of up to 36.9% for current LLM planners; analysis attributes most failures to missing objects in the generated plan. Task-oriented context cognition achieves SOTA versus aware-prompting, chain-of-thought, and in-context learning baselines (specific final-success numbers not stated in abstract).
Reframes embodied-LLM evaluation around pragmatic robustness rather than literal-instruction following โ a prerequisite for deployment to non-expert users. Sits alongside RoboArena in the 2026 push toward more realistic embodied evaluation.
- OpenReview: https://openreview.net/forum?id=vmBIF25KLf
โ Back to ICLR-2026