CVPR 2026 FunREC - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: 3D Scene Reconstruction for Manipulation Trend tag: Trend 5 Affiliations: ETH + MPI Informatics + Stanford + Microsoft + USI Lugano
flowchart LR
EGO["egocentric RGB-D video"] --> SEG["object segmentation"]
EGO --> KIN["kinematics<br/>extraction"]
SEG --> ARTIC["articulated-part decomposition"]
KIN --> ARTIC
ARTIC --> TWIN["functional 3D digital twin<br/>simulation-ready (URDF/USD)"]
TWIN --> POL["robot interaction<br/>(Spot mobile manipulator)"]
A digital twin for manipulation must include articulation (which parts move, how, with what kinematic constraints) — not just geometry. Building such twins by hand is expensive; learning them from video is hard because most video pipelines reconstruct only static geometry.
FunREC reconstructs simulation-ready functional 3D scenes with articulated parts and kinematics from egocentric RGB-D video. Operating on in-the-wild human interaction sequences (no controlled multi-state capture or CAD priors), the system automatically identifies articulated components, estimates their kinematic parameters along with per-timestep poses, and jointly reconstructs the static scene and each movable part (including interiors) in canonical space. Output is exported to URDF/USD for simulation.
Across two new benchmarks — RealFun4D (351 human–scene interactions across 60 real apartments, captured with a head-mounted Azure Kinect DK) and OmniFun4D (127 photorealistic simulated interactions) — FunREC surpasses prior articulated-reconstruction work by a large margin: up to +50 mIoU in part segmentation, 5–10× lower articulation and pose errors, and substantially higher reconstruction accuracy.
For robotics, the reconstructed twins support URDF/USD export, hand-guided affordance mapping, and transfer of the human-demonstrated interaction to a Boston Dynamics Spot mobile manipulator using the inferred contact points and articulation parameters (demonstration-to-execution transfer, not a simulation-trained policy).
FunREC is the cleanest "digital twin from interaction" pipeline for manipulation to date. Closes a long-standing loop: ego video → articulated, simulation-ready digital twin (URDF/USD) → real-robot interaction. Note the robot result is demonstration-to-execution transfer on a Spot manipulator, not a closed-loop policy trained in the reconstructed sim. Closest predecessors: GenManip (LLM-driven scene graph sim) and RoboCasa365 (procedural assets at scale); FunREC's distinguishing feature is articulated-part inference from a single ego pass.
- arXiv: 2604.05621
- Project:
functionalscenes.github.io
← Back to CVPR-2026