ICLR 2026 ManipEvalAgent - Heungwoo/research GitHub Wiki

ManipEvalAgent โ€” promptable, agentic evaluation of manipulation policies

Venue: ICLR 2026 ยท Authors: Yiteng Chen, Huiping Zhuang, Wenbo Li, Shiyi Wang, Xiangyu Zhao, Qingyao Wu ยท Paper: OpenReview https://openreview.net/forum?id=3u6AkbWEls (no arXiv preprint located as of 2026-06) ยท Category: Benchmark / evaluation framework ยท Trend tag: Agentic, VLM-in-the-loop policy evaluation

Approach diagram

flowchart LR
  Q[User query / instruction<br/>"how robust is this policy to clutter?"] --> Plan{Evaluation agent<br/>code-gen + planner}
  Plan --> Batch[Run small batch of<br/>rollouts in sim]
  Batch --> VLM[VLM video understanding<br/>per-rollout analysis]
  VLM --> Obs[Intermediate observations<br/>failure modes, sub-goal hits]
  Obs --> Plan
  Plan -->|stopping criterion met| Report[Fine-grained diagnostic report<br/>not just a single scalar]
Loading

Problem

Standard simulation benchmarks evaluate manipulation policies by large-scale sampling โ€” hundreds to thousands of rollouts per task โ€” which is slow, and they collapse the outcome into a single scalar success rate that hides why a policy fails. They also use a fixed, non-adaptive pipeline that cannot answer targeted user questions (e.g. robustness to a specific distractor, or instruction-following under paraphrase).

Method

ManipEvalAgent reframes evaluation as an agentic, multi-round process:

  • Promptable. The agent ingests a natural-language user query and plans the evaluation procedure (which tasks, perturbations, and metrics) accordingly, via code generation.
  • Small-batch, adaptive sampling. Instead of exhaustive sampling, it runs small batches of rollouts and adaptively plans the next round based on intermediate observations, concentrating effort where uncertainty or failure is high.
  • VLM video understanding. A vision-language model parses rollout videos to produce user-instruction-centric, fine-grained analysis โ€” diagnostic text about failure causes rather than a bare number.

Results

The framework reportedly reaches conclusions comparable to large-scale simulation benchmarks while significantly shortening total evaluation time, and delivers interpretable, diagnostic output beyond a scalar score. (The OpenReview record does not expose specific benchmark names or numeric speedups in the public abstract; numbers omitted here pending the full PDF.)

Significance

Pushes policy evaluation โ€” not just policy learning โ€” toward the agentic / VLM-in-the-loop paradigm. Complements heavy fixed-protocol simulators by offering a cheap, query-driven, interpretable alternative, in the same spirit as world-model-as-evaluator and real-to-sim benchmarking efforts at this venue.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ