RSS 2026 OopsieVerse - Heungwoo/research GitHub Wiki

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Datasets and Benchmarks · paper #98 Authors: Arnav Balaji, Arpit Bahety, Sriniket Ambatipudi, Daniel Lam, Junhong Xu, Roberto Martín-Martín (The University of Texas at Austin) arXiv: 2606.31993 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Robots causing unmodeled damage — spilling water, mishandling objects (Figure 1 of arXiv 2606.31993, © the authors)

Figure 1 ("Oops, did I do that?"): Robots trained and evaluated in conventional simulation are unaware of the real-world damage their actions would cause — e.g. applying excessive forces, heating/freezing inappropriate objects, or spilling water on delicate items. OopsieVerse closes this gap with DamageSim (a general damage-aware plugin, instantiated in Omniverse and MuJoCo) and OopsieBench (32 household manipulation tasks) for evaluating safe robotic behavior.

Problem

Robotic manipulation benchmarks typically score only task completion, so a policy can "succeed" in simulation while relying on behaviors that would be unsafe or destructive in the real world. Existing simulators lack a general mechanism to detect, quantify, and represent damage, leaving safety as a major barrier to deploying household robots.

Method

OopsieVerse augments a standard MDP with damage-related observations, rewards, and/or termination conditions. It converts physical sources — contact forces, temperature changes, and liquid interactions — into mechanical, thermal, and fluid damage signals. Two core components: (1) DamageSim, a simulator-agnostic, object-centric framework for detecting, tracking, and quantifying damage, instantiated in two backends (OmniGibson/Nvidia Omniverse and RoboCasa/MuJoCo); and (2) OopsieBench, a suite of 32 household manipulation tasks spanning common damage modes. The framework is demonstrated across four use cases: safer teleoperated demonstration collection via live damage feedback, damage-conditioned imitation and reinforcement learning, safety benchmarking of state-of-the-art VLAs, and improving real-world safety of sim-to-real policies.

Results

Live damage feedback during teleoperation cut the unsafe-behavior rate from 75% to 15% (a 60-point reduction) while keeping comparable task completion. Damage-penalty RL fine-tuning improved safe outcomes (e.g. Cereal Box safe completion 13% → 33%; a Gaussian policy from 20% → 100% safe success where a task-reward-only policy dropped the object). Benchmarking VLAs exposed a large safety-vs-success gap: on one task GR00T reached 73% completion / 18% safe completion, while π0 reached 17.5% completion / 0% safe completion.

Significance

Provides an open-source foundation for systematic, damage-aware safety research and adds a physically grounded safety axis to VLA evaluation — see Review-VLA-Evaluation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home