RSS 2026 ViserDex - Heungwoo/research GitHub Wiki

ViserDex: Visual Sim-to-Real for Robust Dexterous In-hand Reorientation

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: RL · paper #150 Authors: Arjun Bhardwaj, Maximum Wilder-Smith, Mayank Mittal, Vaishakh Patil, Marco Hutter arXiv: 2604.11138 · program page

Summary compiled from the arXiv paper (v1, 13 Apr 2026); all numbers quoted from the paper. Trend context: RSS 2026 survey.

ViserDex 3DGS sim-to-real pipeline (Figure 1 of arXiv 2604.11138, © the authors)

Fig. 1: A scanned object becomes an augmentable 3D Gaussian scene that is composited into physics simulation and rendered via parallelized rasterization; training proceeds through privileged teacher training, teacher–student distillation, and pose-estimator training. The deployment strip shows a real multi-fingered hand reorienting the object under diverse (including adversarial) lighting and across multiple objects.

Problem

In-hand object reorientation requires precise object-pose estimation. RGB sensing offers rich semantic cues but existing solutions depend on multi-camera rigs or costly ray tracing, and generating photorealistic visual diversity for RGB sim-to-real via standard mesh rendering is computationally intractable, often demanding large compute clusters even for simple objects.

Method

ViserDex is a monocular-RGB sim-to-real framework that integrates 3D Gaussian Splatting (3DGS) to bridge the visual gap. Its key insight is domain randomization directly in the Gaussian representation space: physically consistent pre-rendering augmentations (e.g. perturbations of spherical-harmonic coefficients) produce photorealistic, randomized data for pose estimation. The manipulation policy is trained with curriculum-based RL and teacher–student distillation, and both the perception and control models can be trained independently on consumer-grade hardware.

Results

The pose estimator trained on 3DGS data reaches a mean accuracy of 65.4% and mean ADD of 10.2 mm, surpassing standard tiled rendering (53.3%) and a randomized-tiled baseline (55.6%). Under adversarial lighting it stays most robust at 56.3% mean accuracy versus tiled rendering 47.2% and naive GS 36.5%. The pre-rasterization augmentations add only ≈4% to frame-rendering time. On a physical multi-fingered hand with a single RGB camera, the system robustly reorients five diverse objects (over 25 consecutive reorientations) even under challenging lighting.

Significance

Demonstrates Gaussian splatting as a practical, low-compute path to RGB-only dexterous manipulation. Connects to RL and Review-Dexterous-Manipulation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home