ICLR 2026 ManipEvalAgent - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 ยท Authors: Yiteng Chen, Huiping Zhuang, Wenbo Li, Shiyi Wang, Xiangyu Zhao, Qingyao Wu ยท Paper: OpenReview https://openreview.net/forum?id=3u6AkbWEls (no arXiv preprint located as of 2026-06) ยท Category: Benchmark / evaluation framework ยท Trend tag: Agentic, VLM-in-the-loop policy evaluation
flowchart LR
Q[User query / instruction<br/>"how robust is this policy to clutter?"] --> Plan{Evaluation agent<br/>code-gen + planner}
Plan --> Batch[Run small batch of<br/>rollouts in sim]
Batch --> VLM[VLM video understanding<br/>per-rollout analysis]
VLM --> Obs[Intermediate observations<br/>failure modes, sub-goal hits]
Obs --> Plan
Plan -->|stopping criterion met| Report[Fine-grained diagnostic report<br/>not just a single scalar]
Standard simulation benchmarks evaluate manipulation policies by large-scale sampling โ hundreds to thousands of rollouts per task โ which is slow, and they collapse the outcome into a single scalar success rate that hides why a policy fails. They also use a fixed, non-adaptive pipeline that cannot answer targeted user questions (e.g. robustness to a specific distractor, or instruction-following under paraphrase).
ManipEvalAgent reframes evaluation as an agentic, multi-round process:
- Promptable. The agent ingests a natural-language user query and plans the evaluation procedure (which tasks, perturbations, and metrics) accordingly, via code generation.
- Small-batch, adaptive sampling. Instead of exhaustive sampling, it runs small batches of rollouts and adaptively plans the next round based on intermediate observations, concentrating effort where uncertainty or failure is high.
- VLM video understanding. A vision-language model parses rollout videos to produce user-instruction-centric, fine-grained analysis โ diagnostic text about failure causes rather than a bare number.
The framework reportedly reaches conclusions comparable to large-scale simulation benchmarks while significantly shortening total evaluation time, and delivers interpretable, diagnostic output beyond a scalar score. (The OpenReview record does not expose specific benchmark names or numeric speedups in the public abstract; numbers omitted here pending the full PDF.)
Pushes policy evaluation โ not just policy learning โ toward the agentic / VLM-in-the-loop paradigm. Complements heavy fixed-protocol simulators by offering a cheap, query-driven, interpretable alternative, in the same spirit as world-model-as-evaluator and real-to-sim benchmarking efforts at this venue.
- OpenReview: https://openreview.net/forum?id=3u6AkbWEls
โ Back to ICLR-2026