ICLR 2026 ArtVIP - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Authors: Zhao Jin, Zhengping Che, Tao Li, Zhen Zhao, Kun Wu, Yuheng Zhang, Yinuo Zhao, Zehui Liu, Qiang Zhang, Xiaozhu Ju, Jing Tian, Yousong Xue, Jian Tang · Paper: arXiv 2506.04941 (Jun 2025) · Category: Dataset / asset library + benchmark · Trend tag: High-fidelity sim assets for sim-to-real
flowchart LR
Model[Professional 3D modelers<br/>unified standards] --> Mesh[Precise geometric meshes<br/>high-res PBR textures]
Mesh --> Phys[Fine-tuned dynamic params<br/>enhanced joint-drive equation]
Phys --> Mod[Embedded modular<br/>interaction behaviors]
Mod --> Aff[Pixel-level<br/>affordance annotations]
Aff --> Assets[206 articulated objects<br/>26 categories · USD format]
Assets --> Isaac[Isaac Sim<br/>RTX renderer + GPU physics]
Isaac --> Val[Validation:<br/>imitation learning + RL]
Robot learning increasingly relies on simulation, which demands high-quality digital assets to bridge the sim-to-real gap. Existing open-source articulated-object datasets suffer from insufficient visual realism and low physical fidelity, limiting their usefulness for training policies that transfer to the real world.
ArtVIP is a fully open-source library of high-quality digital-twin articulated objects plus indoor-scene assets, built specifically for Isaac Sim (chosen for its RTX renderer and GPU-parallel physics over MuJoCo/Webots):
- Visual realism via precise geometric meshes and high-resolution textures, rendered with Physically Based Rendering (PBR).
- Physical fidelity via fine-tuned dynamic parameters and an enhanced joint-drive equation (position- and velocity-dependent stiffness/damping) over Isaac Sim's default; collision handled via convex hull / fine-tuned collision / convex decomposition as needed.
- Embedded modular interaction behaviors packaged inside assets (avoiding per-object hand-written joint scripts).
- Pixel-level affordance annotations.
Scale: 26 categories, 206 articulated-object assets (household items, large/small furniture, large/small appliances), shipped in USD format with production guidelines, plus digital-twin and fully interactive environments.
Feature-map visualization and optical motion capture are used to quantitatively demonstrate ArtVIP's visual and physical fidelity. Applicability is validated across imitation learning and reinforcement learning experiments. (Specific success/error metrics omitted here pending the full tables.)
A standards-driven, fully open asset library that raises both the visual and physical bar for articulated objects, directly targeting sim-to-real transfer. The embedded modular interactions and pixel-level affordances make assets reusable across tasks without bespoke scripting — useful infrastructure for any Isaac-Sim-based manipulation pipeline.
- arXiv: https://arxiv.org/abs/2506.04941
- Project page: https://x-humanoid-artvip.github.io/
- OpenReview / HF: https://huggingface.co/papers/2506.04941
← Back to ICLR-2026