RSS 2026 TMRL - Heungwoo/research GitHub Wiki

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 3 · paper #208 Authors: Matthew M. Hong, Jesse Zhang, Anusha Nagabandi, Abhishek Gupta arXiv: 2605.12236 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Context aliasing enables action borrowing (Figure 1 of arXiv 2605.12236, © the authors)

Left: on a downstream task ("put the carrot in the pot") the BC conditional p(a|c) has collapsed support and fails. Center: injecting forward-diffusion noise into the context c aliases nearby contexts, moving the action distribution from sharp imitation p(a|c) toward the marginal p(a) — no overlap at σ=0, partial overlap at medium σ, full overlap at high σ — so the policy "borrows" reach and place behaviors from other tasks. Right: the adapted policy completes the task.

Problem

BC-pretrained policies model narrow conditional action distributions: under distribution shift, support collapses, online rollouts earn no reward, and RL fine-tuning stalls. Prior fixes inject Gaussian action noise for coverage, which yields incoherent dithering and cannot be controlled during fine-tuning.

Method

Context-Smoothed Pre-training (CSP) applies a forward-diffusion corruption kernel to the policy's inputs (contexts) — states, point clouds, or VLM embeddings — training a generative policy p(a | noisy c, σ) across all noise levels, creating a continuum from sharp imitation to the marginal action distribution. Timestep-Modulated RL (TMRL) then trains a high-level policy that adaptively outputs the diffusion timestep (context-noise dial σ) plus a latent, treating conditioning strength as an explicit exploration control. Theory shows context smoothing strictly increases action coverage between overlapping contexts. The recipe is policy-agnostic: state-input diffusion policies, point-cloud policies, and π0-style VLAs (noise on VLM embeddings before the action expert).

Results

On OGBench (pointmaze-giant, cube-single), CSP beats BC and PostBC on success@K coverage, and TMRL delivers a 101% average improvement over the best baseline (DSRL, PostBC, RLPD, SPiRL), including near-100% on cube-single where DSRL is near zero. On LIBERO adaptation, only TMRL solves the libero-90 task. With a LEAP-hand dexterous grasping policy on point clouds, TMRL reaches 2.5× DSRL's final success. In the real world, fine-tuning π0 on WidowX (BridgeData-v2) and Franka DROID tasks (sausage-in-pot, shrimp-in-drawer, press-button), TMRL reaches high success within tens of episodes — under one hour — while DSRL fails to learn; converged policies visibly anneal σ downward within a rollout.

Significance

Reframes pre-training-for-RL from action-noise injection to structured context smoothing, giving RL an interpretable knob that interpolates between imitation and exploration — and demonstrating practical real-robot VLA fine-tuning in under an hour. Related wiki threads: Review-Human-Video-Transfer · Review-Realtime-Execution.

← Back to RSS 2026 survey · RSS-2026-Papers · Home