RSS 2026 Papers - Heungwoo/research GitHub Wiki
RSS 2026 — Full In-Scope Paper Index
All 116 in-scope papers of RSS 2026, by session, each linked to its per-paper page. Bolded entries have figure-illustrated analysis pages; all others are verified-abstract reference pages. Split from the main survey to keep both pages fast to render.
Every paper links to its own page. Bolded entries have figure-illustrated analysis pages; all others are verified-abstract reference pages (verbatim program abstract + context links).
VLA Models (session, 9)
| # | Paper | One-line takeaway |
|---|---|---|
| 81 | [X-DiffVLA](/Heungwoo/research/wiki/RSS-2026-X-DiffVLA) | Cross-embodied diffusion action heads to avoid per-embodiment fine-tuning |
| 82 | [SkillVLA](/Heungwoo/research/wiki/RSS-2026-SkillVLA) | Skill reuse against combinatorial diversity in dual-arm tasks |
| 83 | [BagelVLA](/Heungwoo/research/wiki/RSS-2026-BagelVLA) | Interleaved vision-language-action generation for long-horizon manipulation |
| 84 | [GuidedVLA](/Heungwoo/research/wiki/RSS-2026-GuidedVLA) | Plug-and-play action-attention specialization against visual shortcuts |
| 85 | [AR-VLA](/Heungwoo/research/wiki/RSS-2026-AR-VLA) | Autoregressive action expert with persistent history + refreshable VL prefixes |
| 86 | [Continual RL fine-tuning](/Heungwoo/research/wiki/RSS-2026-Towards-Long-Lived-Robots) | RFT as the anti-forgetting adaptation mechanism for long-lived VLAs |
| 87 | [π*0.6 + RECAP](/Heungwoo/research/wiki/PI-RECAP) | RL from deployment experience: 2× throughput, ~½ failures |
| 88 | [StereoVLA](/Heungwoo/research/wiki/RSS-2026-StereoVLA) | Stereo geometric cues into the VLA visual stack |
| 89 | [RLux-VLA](/Heungwoo/research/wiki/RSS-2026-RLux-VLA) | Unified RL-for-VLA framework (platform + efficiency) |
Manipulation 1–3 (27)
| # | Paper | One-line takeaway |
|---|---|---|
| 1 | [BiDemoSyn](/Heungwoo/research/wiki/RSS-2026-One-Shot-Real-World-Demonstration-Synthesis-for) | One-shot real-world demo synthesis for bimanual manipulation |
| 2 | [Surgical MoE](/Heungwoo/research/wiki/RSS-2026-Supervised-Mixture-of-Experts-for-Surgical-Grasping) | Supervised MoE for phase-structured surgical grasping/retraction |
| 3 | [DexImit](/Heungwoo/research/wiki/RSS-2026-DexImit) | Bimanual dexterous skills from monocular human videos |
| 4 | [Semantic Contact Fields](/Heungwoo/research/wiki/RSS-2026-Semantic-Contact-Fields-for-Category-Level) | Category-level contact representation for tool manipulation |
| 5 | [Contact-Grounded Policy](/Heungwoo/research/wiki/RSS-2026-Contact-Grounded-Policy) | Generative contact grounding for dexterous visuotactile control |
| 6 | [TactAlign](/Heungwoo/research/wiki/RSS-2026-TactAlign) | Human-to-robot transfer via tactile-signal alignment |
| 7 | [LBM co-training study](/Heungwoo/research/wiki/RSS-2026-LBM-Cotraining-Study) | 89 policies, 5 modalities: what co-training actually helps |
| 8 | [SID](/Heungwoo/research/wiki/RSS-2026-SID) | Sliding into distribution for few-demonstration robustness |
| 9 | [UMI-Underwater](/Heungwoo/research/wiki/RSS-2026-UMI-Underwater) | Underwater manipulation without underwater teleop |
| 54 | [CoRAL](/Heungwoo/research/wiki/RSS-2026-CoRAL) | LLM-based adaptive control for contact-rich manipulation |
| 55 | [GHOST](/Heungwoo/research/wiki/RSS-2026-GHOST) | Hierarchical 3D sub-goal policies for OOD generalization |
| 56 | [Robo3R](/Heungwoo/research/wiki/RSS-2026-Robo3R) | Manipulation-ready feed-forward metric 3D reconstruction |
| 57 | [Structured MoE](/Heungwoo/research/wiki/RSS-2026-Semantically-Structured-Mixture-of-Experts-for-Compositional) | Semantically structured experts for compositional manipulation |
| 58 | [Legato](/Heungwoo/research/wiki/RSS-2026-Legato) | Training-time native continuation for chunked flow policies |
| 59 | [DexEvolve](/Heungwoo/research/wiki/RSS-2026-DexEvolve) | Evolutionary dexterous grasp synthesis across morphologies |
| 60 | [TACTIC](/Heungwoo/research/wiki/RSS-2026-TACTIC) | Tactile+vision contact-centric whole-arm manipulation |
| 61 | [Stein DR control](/Heungwoo/research/wiki/RSS-2026-Distributionally-Robust-Control-via-Stein) | Distributionally robust contact-rich control, few-sample regime |
| 62 | [PolaRiS](/Heungwoo/research/wiki/RSS-2026-PolaRiS) | Real-to-sim evaluation that actually ranks generalist policies |
| 121 | [R2RGen](/Heungwoo/research/wiki/RSS-2026-R2RGen) | Real-to-real 3D data generation for spatial generalization |
| 122 | [DexGrasp-Zero](/Heungwoo/research/wiki/RSS-2026-DexGrasp-Zero) | 85% zero-shot grasping on unseen dexterous hands |
| 123 | [Minimalist Compliance](/Heungwoo/research/wiki/RSS-2026-Minimalist-Compliance-Control) | Compliance control without F/T sensors or RL complexity |
| 124 | [One Hand to Rule Them All](/Heungwoo/research/wiki/RSS-2026-One-Hand) | Canonical parameterized representation unifying hand morphologies |
| 125 | [AxisGuide](/Heungwoo/research/wiki/RSS-2026-AxisGuide) | Grounding the action coordinate system in RGB for robustness |
| 126 | [Task-Level ILC](/Heungwoo/research/wiki/RSS-2026-Learning-Deformable-Object-Manipulation-Using) | Iterative learning control for dynamic deformable manipulation |
| 127 | [CLAMP](/Heungwoo/research/wiki/RSS-2026-CLAMP) | Contrastive 3D multi-view action-conditioned pretraining |
| 128 | [Force Policy](/Heungwoo/research/wiki/RSS-2026-Force-Policy) | Hybrid force–position policy in the interaction frame |
| 129 | [ViTacFormer](/Heungwoo/research/wiki/RSS-2026-ViTacFormer) | Visuo-tactile cross-modal latents + tactile prediction; ~50% ↑ |
Imitation Learning 1–3 (28)
| # | Paper | One-line takeaway |
|---|---|---|
| 72 | [H2R Emergence](/Heungwoo/research/wiki/RSS-2026-Human2Robot-Emergence) | Human-to-robot transfer emerges with pre-training diversity |
| 73 | [PointACT](/Heungwoo/research/wiki/RSS-2026-PointACT) | Multi-scale point-action interaction for 3D grounding |
| 74 | [Steerable VLA policies](/Heungwoo/research/wiki/RSS-2026-Steerable-Vision-Language-Action-Policies-for-Embodied) | Hierarchical steering of VLAs for embodied reasoning |
| 75 | [OAT](/Heungwoo/research/wiki/RSS-2026-OAT) | Ordered action tokenization with anytime prefix decoding |
| 76 | [Beyond Binary Success](/Heungwoo/research/wiki/RSS-2026-Beyond-Binary-Success) | Statistically rigorous, sample-efficient policy comparison |
| 77 | [mimic-video](/Heungwoo/research/wiki/RSS-2026-mimic-video) | Video-action models: video backbone + IDM decoder, 10× sample-efficient |
| 78 | [TouchGuide](/Heungwoo/research/wiki/RSS-2026-TouchGuide) | Inference-time touch guidance for pretrained policies |
| 79 | [Visual verification](/Heungwoo/research/wiki/RSS-2026-Visual-Verification-Enables-Inference-time-Steering) | Generator–verifier loop for autonomous policy improvement |
| 80 | [Set-Supervised DP](/Heungwoo/research/wiki/RSS-2026-Set-Supervised-Diffusion-Policy) | Learning action-chunking diffusion from corrections |
| 139 | [Tune to Learn](/Heungwoo/research/wiki/RSS-2026-Tune-to-Learn) | Controller gains as a first-class policy-learning variable |
| 140 | [Robometer](/Heungwoo/research/wiki/RSS-2026-Robometer) | Reward models from trajectory comparisons at scale |
| 141 | [Contact-Anchored Policies](/Heungwoo/research/wiki/RSS-2026-Contact-Anchored-Policies) | Contact conditioning instead of language conditioning |
| 142 | [Act/Ask/Learn](/Heungwoo/research/wiki/RSS-2026-When-to-Act-Ask-or-Learn) | Uncertainty-aware policy steering with VLM verifiers |
| 143 | [ReSteer](/Heungwoo/research/wiki/RSS-2026-ReSteer) | Quantifying and refining multitask policy steerability |
| 144 | [ENAP](/Heungwoo/research/wiki/RSS-2026-Emergent-Neural-Automaton-Policies-Learning) | Emergent neural automata: symbolic structure from trajectories |
| 145 | [Universal Pose Pretraining](/Heungwoo/research/wiki/RSS-2026-Universal-Pose-Pretraining-for-Generalizable) | Pose-supervised pretraining against VLA feature collapse |
| 146 | [EigenSafe](/Heungwoo/research/wiki/RSS-2026-EigenSafe) | Spectral learned safety assessment for stochastic systems |
| 147 | [DISC](/Heungwoo/research/wiki/RSS-2026-DISC) | Policy generation decoupling instruction from state control |
| 201 | [Key History Frames](/Heungwoo/research/wiki/RSS-2026-Long-Context-Robot-Imitation-Learning-by) | Long-context IL by attending to key past frames |
| 202 | [SoftAct](/Heungwoo/research/wiki/RSS-2026-Functional-Force-Aware-Retargeting-from-Virtual) | Force-aware retargeting from VR human demos to soft hands |
| 203 | [LAP](/Heungwoo/research/wiki/RSS-2026-LAP) | Language-action pretraining for zero-shot cross-embodiment |
| 204 | [EgoHumanoid](/Heungwoo/research/wiki/RSS-2026-Unlocking-In-the-Wild-Loco-Manipulation-with-Robot-Free) | Robot-free egocentric demos for in-the-wild loco-manipulation |
| 205 | [HoMMI](/Heungwoo/research/wiki/RSS-2026-HoMMI) | UMI+egocentric sensing → whole-body mobile manipulation |
| 206 | [Mimic Intent](/Heungwoo/research/wiki/RSS-2026-Mimic-Intent-Not-Just-Trajectories) | Intent-level imitation over raw-trajectory cloning |
| 207 | [TAIL-Safe](/Heungwoo/research/wiki/RSS-2026-TAIL-Safe) | Task-agnostic safety monitoring for IL policies |
| 208 | [TMRL](/Heungwoo/research/wiki/RSS-2026-TMRL) | Timestep-modulated diffusion pretraining for RL exploration |
| 209 | [A2A Flow Matching](/Heungwoo/research/wiki/RSS-2026-Action-to-Action-Flow-Matching) | Action-to-action flow: skip Gaussian re-noising for low latency |
| 210 | [LDA-1B](/Heungwoo/research/wiki/RSS-2026-LDA-1B) | 1B unified world model over 30k h; +48% dexterous vs π0.5 |
Humanoids (13)
| # | Paper | First line of abstract |
|---|---|---|
| 19 | [HUSKY](/Heungwoo/research/wiki/RSS-2026-HUSKY) | While current humanoid whole-body control frameworks predominantly rely on the static environment assumptions, addressing tasks characterized by high … |
| 20 | [Perceptive Humanoid Parkour](/Heungwoo/research/wiki/RSS-2026-Perceptive-Humanoid-Parkour) | While recent advances in humanoid locomotion have achieved stable walking on varied terrains, capturing the agility and adaptivity of highly dynamic h… |
| 21 | [Ψ₀](/Heungwoo/research/wiki/Review-Psi0) | We introduce Ψ₀ (Psi-Zero), an open foundation model to address challenging humanoid loco-manipulation tasks. |
| 22 | [X-Loco](/Heungwoo/research/wiki/RSS-2026-X-Loco) | While recent advances have demonstrated strong performance in individual humanoid skills such as upright locomotion, fall recovery and whole-body coor… |
| 23 | [Learning to Evolve](/Heungwoo/research/wiki/RSS-2026-Learning-to-Evolve) | Achieving safe manipulation-oriented navigation for humanoid robots is fundamentally challenged by two factors: locomotion-induced perceptual distorti… |
| 24 | [MOBIUS](/Heungwoo/research/wiki/RSS-2026-MOBIUS) | This paper presents the MOBIUS platform, a bipedal robot capable of walking, crawling, climbing, and rolling. |
| 25 | [TeleGate](/Heungwoo/research/wiki/RSS-2026-TeleGate) | Real-time whole-body teleoperation is a critical method for humanoid robots to perform complex tasks in unstructured environments. |
| 26 | [Generalizing from References using a Multi-Task Reference …](/Heungwoo/research/wiki/RSS-2026-Generalizing-from-References-using-a) | Learning agile humanoid behaviors from human motion offers a powerful route to natural, coordinated control, but existing approaches face a persistent… |
| 27 | [Now You See That](/Heungwoo/research/wiki/RSS-2026-Now-You-See-That) | Achieving robust vision-based humanoid locomotion remains challenging due to two fundamental issues: the sim-toreal gap introduces significant percept… |
| 28 | [Mind Your Steps](/Heungwoo/research/wiki/RSS-2026-Mind-Your-Steps) | Enabling humanoid robots to operate in complex, dynamic environments remains a critical challenge, fundamentally limited by the ability to navigate ro… |
| 29 | [PRIME](/Heungwoo/research/wiki/RSS-2026-PRIME) | Humanoid and legged robots interact with the environment through intermittent contacts, making accurate motion estimation fundamentally dependent on r… |
| 30 | [HiWET](/Heungwoo/research/wiki/RSS-2026-HiWET) | Humanoid loco-manipulation requires executing precise manipulation tasks while maintaining dynamic stability amid base motion and impacts. |
| 31 | [OmniXtreme](/Heungwoo/research/wiki/RSS-2026-OmniXtreme) | High-fidelity motion tracking serves as the ultimate litmus test for generalizable, human-level motor skills. |
World Models & Memory (9)
| # | Paper | First line of abstract |
|---|---|---|
| 10 | [Memory Retrieval in Visuomotor Policies for Long-Horizon …](/Heungwoo/research/wiki/RSS-2026-Memory-Retrieval-in-Visuomotor-Policies) | General-purpose robots operating in partially observable environments such as homes require memory to support long-term autonomy. |
| 11 | [RAG-Diff](/Heungwoo/research/wiki/RSS-2026-RAG-Diff) | Robots operating in unstructured environments must satisfy dynamic constraints that can change across tasks and even within a single execution. |
| 12 | [Self-Improving Robot Policy with Compositional World …](/Heungwoo/research/wiki/RSS-2026-Self-Improving-Robot-Policy-with-Compositional) | Despite the sustained scaling on model capacity and data acquisition, Vision–Language–Action (VLA) models remain brittle in contact-rich and dynamic m… |
| 13 | [HAIC](/Heungwoo/research/wiki/RSS-2026-HAIC) | Humanoid robots exhibit significant potential for executing complex whole-body interaction tasks in unstructured environments. |
| 14 | [Collaborating Visual and Parameter Spaces for Consistent …](/Heungwoo/research/wiki/RSS-2026-Collaborating-Visual-and-Parameter-Spaces) | Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for evaluating Vision-Language-Action (VLA) systems. |
| 15 | [Act2Goal](/Heungwoo/research/wiki/RSS-2026-Act2Goal) | Specifying robotic manipulation tasks in a manner that is both expressive and precise remains a central challenge. |
| 16 | [Causal World Modeling for Robot Control](/Heungwoo/research/wiki/RSS-2026-Causal-World-Modeling-for-Robot) | This work highlights that video world modeling, alongside vision-language pre-training, establishes a distinct foundation for robot learning. |
| 17 | [Simulation Distillation](/Heungwoo/research/wiki/RSS-2026-Simulation-Distillation) | Simulation-to-real transfer remains a central challenge in robotics, as mismatches between simulated and real-world dynamics often lead to failures. |
| 18 | [Interactive World Simulator for Robot Policy Training and …](/Heungwoo/research/wiki/RSS-2026-Interactive-World-Simulator-for-Robot) | Action-conditioned video prediction models (often referred to as world models) have shown strong potential for robotics applications, but existing wor… |
RL (9)
| # | Paper | First line of abstract |
|---|---|---|
| 148 | [Zero-Shot Sim-to-Real Robot Learning](/Heungwoo/research/wiki/RSS-2026-Zero-Shot-Sim-to-Real-Robot-Learning-A) | Dexterous manipulation is physics-intensive and highly sensitive to modeling errors and perception noise, making sim-to-real transfer prohibitively ch… |
| 149 | [Emerging Extrinsic Dexterity in Cluttered Scenes via …](/Heungwoo/research/wiki/RSS-2026-Emerging-Extrinsic-Dexterity-in-Cluttered) | Extrinsic dexterity leverages environmental contact to overcome the limitations of prehensile manipulation. |
| 150 | [ViserDex](/Heungwoo/research/wiki/RSS-2026-ViserDex) | In-hand object reorientation requires precise estimation of the object pose to handle complex task dynamics. |
| 151 | [SimToolReal](/Heungwoo/research/wiki/RSS-2026-SimToolReal) | The ability to manipulate tools significantly expands the set of tasks a robot can perform. |
| 152 | [Latent Policy Steering through One-Step Flow Policies](/Heungwoo/research/wiki/RSS-2026-Latent-Policy-Steering-through-One-Step) | Offline reinforcement learning (RL) should be ideal for robotics, allowing learning from dataset without risky exploration. |
| 153 | [When Life Gives You BC, Make Q-functions](/Heungwoo/research/wiki/RSS-2026-When-Life-Gives-You-BC) | Behavior Cloning (BC) has emerged as a highly effective paradigm for robot learning. |
| 154 | [Offline Policy Evaluation for Manipulation Policies via …](/Heungwoo/research/wiki/RSS-2026-Offline-Policy-Evaluation-for-Manipulation) | Policy evaluation is a fundamental component of the development and deployment pipeline for robotic policies. |
| 155 | [HydroShear](/Heungwoo/research/wiki/RSS-2026-HydroShear) | In this paper, we address the problem of tactile sim-to-real policy transfer for contact-rich tasks. |
| 156 | [Toward Reliable Sim-to-Real Predictability for MoE-based …](/Heungwoo/research/wiki/RSS-2026-Toward-Reliable-Sim-to-Real-Predictability-for) | Reinforcement learning has shown strong promise for quadrupedal agile locomotion, even with proprioception-only sensing. |
Datasets and Benchmarks (9)
| # | Paper | First line of abstract |
|---|---|---|
| 90 | [Betting for Sim-to-Real Performance Evaluation](/Heungwoo/research/wiki/RSS-2026-Betting-for-Sim-to-Real-Performance-Evaluation) | This paper studies the problem of robot performance evaluation, focusing on how to obtain accurate and efficient estimates of real-world behavior unde… |
| 91 | [MolmoSpaces](/Heungwoo/research/wiki/RSS-2026-MolmoSpaces) | Deploying robots at scale demands robustness to the long tail of everyday situations. |
| 92 | [EgoVerse](/Heungwoo/research/wiki/RSS-2026-EgoVerse) | Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. |
| 93 | [GS-Playground](/Heungwoo/research/wiki/RSS-2026-GS-Playground) | Embodied AI research is undergoing a shift toward vision-centric perceptual paradigms. |
| 94 | [High Fidelity Capture, Reconstruction, and Transfer of …](/Heungwoo/research/wiki/RSS-2026-High-Fidelity-Capture-Reconstruction-and) | Despite the demand for robots in high-value clinical tasks like bathing, contemporary systems still lack the safety and reliability required for compl… |
| 95 | [RoboVista](/Heungwoo/research/wiki/RSS-2026-RoboVista) | Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions,… |
| 96 | [RoboLab](/Heungwoo/research/wiki/RSS-2026-RoboLab) | The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid … |
| 97 | [LIBERO-X](/Heungwoo/research/wiki/RSS-2026-LIBERO-X) | Reliable benchmarking is critical for advancing Vision–Language–Action (VLA) models, as it reveals their generalization, robustness, and alignment of … |
| 98 | [OopsieVerse](/Heungwoo/research/wiki/RSS-2026-OopsieVerse) | While robotic manipulation capabilities have advanced rapidly, physical safety remains a major barrier to deploying household robots: task success is … |
Hands, tactile & contact picks from other sessions (12)
| # | Paper | First line of abstract |
|---|---|---|
| 192 | [CRAFT](/Heungwoo/research/wiki/RSS-2026-CRAFT) | We introduce CRAFT Hand, a tendon-driven anthropomorphic hand with hybrid hard-soft compliance for contact-rich manipulation. |
| 193 | [LightTact](/Heungwoo/research/wiki/RSS-2026-LightTact) | Contact often occurs without macroscopic surface deformation, such as during interaction with liquids, semi-liquids, or ultra-soft materials. |
| 195 | [Latent Diffeomorphic Co-Design of End-Effectors for …](/Heungwoo/research/wiki/RSS-2026-Latent-Diffeomorphic-Co-Design-of-End-Effectors) | Manipulating deformable and fragile objects remains a fundamental challenge in robotics due to complex contact dynamics and strict requirements on obj… |
| 197 | [A Dual-Mode Electrical Capacitance Tomography Sensor for …](/Heungwoo/research/wiki/RSS-2026-A-Dual-Mode-Electrical-Capacitance-Tomography) | Tactile and proximity sensing is fundamental for achieving autonomous robotic manipulation and safe human-robot interaction. |
| 199 | [A Super-Resolution and Multi-Axis Tactile Sensor with Soft …](/Heungwoo/research/wiki/RSS-2026-A-Super-Resolution-and-Multi-Axis-Tactile) | To achieve human-like skin tactile perception with super-resolution, the method of introducing a soft layer on sensing array has attracted increasing … |
| 200 | [Active Surface-Driven Reconfigurable Gripper](/Heungwoo/research/wiki/RSS-2026-Active-Surface-Driven-Reconfigurable-Gripper-Robust) | Robotic grippers face substantial challenges in grasping and manipulating thin objects. |
| 167 | [More with LESS – Local Scene Representations for Tactile …](/Heungwoo/research/wiki/RSS-2026-More-with-LESS-Local-Scene) | Tactile imaging seeks to reconstruct the internal structure of soft objects through touch sensing, with applications in medical diagnosis and robotic … |
| 177 | [Relaxation-Aware Multimodal Sensing of Soft Gripper Driven …](/Heungwoo/research/wiki/RSS-2026-Relaxation-Aware-Multimodal-Sensing-of-Soft) | Achieving stable, sustained grasping with soft robotic hands remains a fundamental challenge. |
| 178 | [CoCo-InEKF](/Heungwoo/research/wiki/RSS-2026-CoCo-InEKF) | Robust state estimation for highly dynamic motion of legged robots remains challenging, especially in dynamic, contact-rich scenarios. |
| 180 | [From Reaction to Anticipation](/Heungwoo/research/wiki/RSS-2026-From-Reaction-to-Anticipation) | Recent advances in robotic manipulation remain hindered by the inevitability of task failures, particularly in dynamic and unstructured environments. |
| 190 | [Certifiable Gradient-Based Contact-Rich Manipulation via …](/Heungwoo/research/wiki/RSS-2026-Certifiable-Gradient-Based-Contact-Rich-Manipulation-via) | While gradient-based methods can efficiently optimize trajectories and controllers by exploiting physical priors and differentiable simulators, contac… |
| 163 | [IMPACT](/Heungwoo/research/wiki/RSS-2026-IMPACT) | Contact-implicit trajectory optimization (CITO) has attracted growing attention as a unified framework for planning and control in contact-rich roboti… |
← Back to RSS 2026 survey · Home