RSS 2026 RAG Diff - Heungwoo/research GitHub Wiki
RAG-Diff: Adapting Diffusion Policies to Dynamic Constraints with Retrieval-Augmented Guidance
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: World Models & Memory · paper #11 Authors: Ruolin Ye, Nayoung Ha, Shuaixing Chen, Qiandao Liu, Gavin Chen, Shaoyang Stassen, Mark Zolotas, Jose Barreiros, Tapomayukh Bhattacharjee Program page: https://roboticsconference.org/program/papers/11/
No public preprint found (searched arXiv 2026-08-10); summary derived from the verified program abstract. Trend placement and neighbors: RSS 2026 survey.
Summary
Problem. Robots in unstructured settings must satisfy dynamic constraints that change across tasks and even within one execution; a trained diffusion policy has no built-in way to adapt at runtime to newly encountered or evolving constraints. Method. RAG-Diff is a runtime-adaptation framework for a frozen transformer diffusion policy backed by retrieval-augmented memory. It maintains PrefMem, a bank of vision-language embeddings paired with state-action snippets and constraint annotations; at test time it retrieves the nearest entry and steers sampling two ways — I-Atten inserts the retrieved snippet as extra cross-attention memory tokens with a classifier-free-guidance-style update to bias denoising toward preference-consistent motion, while a predictive guidance mechanism uses the retrieved constraint parameters to discourage violations. Results. Demonstrated on physical robot caregiving (personalized, time-varying constraints): an adapted PushT sim with contact-force limits and regions-to-avoid, plus caregiving tasks (bed bathing, medicine delivery, shelf cleaning, feeding) in RCareWorld simulation and on a real robot, with real-world user studies on bed bathing. RAG-Diff improves both task success and constraint satisfaction over baselines including unguided diffusion and other guidance-/sampling-based variants (the abstract reports no specific numbers).
Abstract
Robots operating in unstructured environments must satisfy dynamic constraints that can change across tasks and even within a single execution. While diffusion policies can learn multimodal behaviors from demonstrations, adapting a trained policy at runtime to newly encountered or evolving constraints remains an open challenge. We propose RAG-Diff, a runtime adaptation framework for a frozen transformer diffusion policy that leverages retrieval-augmented memory. RAG-Diff maintains PrefMem, a memory bank that stores vision-language embeddings together with (i) state-action snippets and (ii) constraint annotations. At test time, RAG-Diff queries PrefMem to retrieve the nearest entry and uses it to steer sampling in two complementary ways. First, I-Atten (in-place attention recomputation) inserts the retrieved snippet as additional cross-attention memory tokens and performs a classifier-free-guidance-style update, biasing denoising toward preference-consistent motion. Second, a predictive guidance mechanism incorporates the retrieved constraint parameters during diffusion sampling to discourage violations. To demonstrate the effectiveness of RAG-Diff, we choose physical robot caregiving as a domain with personalized and time-varying constraints. We first benchmark on an adapted PushT environment in simulation with contact-force limits and region-to-avoid constraints. We then evaluated our method on a suite of physical caregiving tasks spanning diverse preference types: (i) interaction and affordance preferences in bed bathing, (ii) ROM-based assistance-level preferences in medicine delivery, (iii) semantic preferences in shelf cleaning, and (iv) trajectory preferences in feeding, in both RCareWorld simulation and with a real robot. We further conducted real-world user studies on the bed-bathing task. Results show that RAG-Diff improves both task success and constraint satisfaction compared to a range of baselines, including unguided diffusion and other guidance- or sampling-based variants.
Wiki context
Related topic reviews: Review-World-Models · Review-VLA-Memory
← Back to RSS 2026 survey · Home