RSS 2026 GS Playground - Heungwoo/research GitHub Wiki

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Datasets and Benchmarks · paper #93 Authors: Yufei Jia, Heng Zhang, Ziheng Zhang, Lei Han, Junzhe Wu, Mingrui Yu, Zifan Wang, Dixuan Jiang, Zheng Li, Chenyu Cao, Zhuoyuan Yu, Xun Yang, Haizhou Ge, Yuchi Zhang, Jiayuan Zhang, Zhenbiao Huang, Tianle Liu, Shenyu Chen, Jiacheng Wang, Bin Xie, ...more> arXiv: 2604.25459 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

GS-Playground overview (Figure 1 of arXiv 2604.25459, © the authors)

The platform combines batch 3D Gaussian Splatting rendering (bottom left, parallel table-top scenes) with a parallel physics engine (bottom right, humanoids on stairs). Left column: the sensor suite (RGB, depth, height scan, LiDAR, force/torque, contact); right column: supported task families — locomotion, navigation, manipulation — across quadruped, humanoid, and arm embodiments.

Problem

Massively parallel simulators transformed proprioception-based locomotion, but photorealistic rendering at scale has been too expensive for vision-centric RL; simulation-ready 3D assets still require labor-intensive manual modeling; and the sim-to-real physical gap blocks contact-rich manipulation transfer. GS-Playground (THU-led, with Motphys, Dexmal, DISCOVER Robotics, and others) attacks all three bottlenecks.

Method

Three components: (1) MotrixSim, a cross-platform (Windows/Linux/macOS, CPU+GPU) parallel physics engine using a velocity-impulse MCP contact solver with constraint-island parallelization, MJCF-compatible; (2) a BatchSplat 3DGS renderer with a point-pruning strategy (retaining ~30% of Gaussians at near-identical PSNR/SSIM; up to 90% reduction for robot-linked assets) that reaches ~10^4 aggregate FPS over 2048 environments at 640×480 on a single GPU; (3) an automated Real2Sim pipeline (Grounding DINO + SAM2 segmentation, background inpainting, 3DGS/mesh reconstruction, Speedy-Splat pruning) that converts a single RGB capture into an interaction-ready scene in under 5 minutes per scene on an RTX 3090 — used to build a Bridge-GS enrichment of Bridge-v2.

Results

Physics: at 50 humanoids per environment, GS-Playground (CPU) holds 1,015 FPS — a 32× speedup over MuJoCo and ~600× over MjWarp, where Genesis fails to converge at N=10. Rendering beats Isaac Sim's ray tracer at all batch sizes and resolutions (Isaac Sim OOMs at 1280×720 large batches). Go1 locomotion trains to comparable rewards faster than IsaacLab at decimation 1; a Go2 policy converges in 10 minutes wall-clock (1,024 envs) and a G1 humanoid in ~6 hours (2,048 envs), both deploying zero-shot. The headline result: an RGB-based PickCube policy on an Airbot Play arm transfers zero-shot at 90% (18/20) real-world success, while matched policies trained in Mujoco Playground, ManiSkill3, and Isaac Lab all score 0/20 — attributed to the visual domain gap. Sim and real success rates correlate at 0.89 across tasks.

Significance

Makes large-scale photorealistic visual RL practical on a single GPU and gives a stark head-to-head demonstration that rendering fidelity, not just physics, decides zero-shot visual sim-to-real. A key infrastructure entry for the wiki's Review-VLA-Evaluation (sim benchmarking and real-to-sim evaluation) thread.

← Back to RSS 2026 survey · RSS-2026-Papers · Home