ICLR 2026 ArtVIP - Heungwoo/research GitHub Wiki

ArtVIP — articulated digital-twin assets for robot learning

Venue: ICLR 2026 · Authors: Zhao Jin, Zhengping Che, Tao Li, Zhen Zhao, Kun Wu, Yuheng Zhang, Yinuo Zhao, Zehui Liu, Qiang Zhang, Xiaozhu Ju, Jing Tian, Yousong Xue, Jian Tang · Paper: arXiv 2506.04941 (Jun 2025) · Category: Dataset / asset library + benchmark · Trend tag: High-fidelity sim assets for sim-to-real

Approach diagram

flowchart LR
  Model[Professional 3D modelers<br/>unified standards] --> Mesh[Precise geometric meshes<br/>high-res PBR textures]
  Mesh --> Phys[Fine-tuned dynamic params<br/>enhanced joint-drive equation]
  Phys --> Mod[Embedded modular<br/>interaction behaviors]
  Mod --> Aff[Pixel-level<br/>affordance annotations]
  Aff --> Assets[206 articulated objects<br/>26 categories · USD format]
  Assets --> Isaac[Isaac Sim<br/>RTX renderer + GPU physics]
  Isaac --> Val[Validation:<br/>imitation learning + RL]
Loading

Problem

Robot learning increasingly relies on simulation, which demands high-quality digital assets to bridge the sim-to-real gap. Existing open-source articulated-object datasets suffer from insufficient visual realism and low physical fidelity, limiting their usefulness for training policies that transfer to the real world.

Method

ArtVIP is a fully open-source library of high-quality digital-twin articulated objects plus indoor-scene assets, built specifically for Isaac Sim (chosen for its RTX renderer and GPU-parallel physics over MuJoCo/Webots):

  • Visual realism via precise geometric meshes and high-resolution textures, rendered with Physically Based Rendering (PBR).
  • Physical fidelity via fine-tuned dynamic parameters and an enhanced joint-drive equation (position- and velocity-dependent stiffness/damping) over Isaac Sim's default; collision handled via convex hull / fine-tuned collision / convex decomposition as needed.
  • Embedded modular interaction behaviors packaged inside assets (avoiding per-object hand-written joint scripts).
  • Pixel-level affordance annotations.

Scale: 26 categories, 206 articulated-object assets (household items, large/small furniture, large/small appliances), shipped in USD format with production guidelines, plus digital-twin and fully interactive environments.

Results

Feature-map visualization and optical motion capture are used to quantitatively demonstrate ArtVIP's visual and physical fidelity. Applicability is validated across imitation learning and reinforcement learning experiments. (Specific success/error metrics omitted here pending the full tables.)

Significance

A standards-driven, fully open asset library that raises both the visual and physical bar for articulated objects, directly targeting sim-to-real transfer. The embedded modular interactions and pixel-level affordances make assets reusable across tasks without bespoke scripting — useful infrastructure for any Isaac-Sim-based manipulation pipeline.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️