ICLR 2026 ViPRA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Skild AI + CMU + UC Irvine (Sandeep Routray, Hengkai Pan, Unnat Jain, Shikhar Bahl, Deepak Pathak) Category: Video-for-Actions / Pretraining Trend tag: Trend 8 (video-as-pretraining for control)
flowchart LR
HV[Human videos<br/>SSv2 198k] --> LAT["Latent Action Tokenizer I_β<br/>NSVQ codebook size 8<br/>DINOv2 spatio-temporal enc."]
RV[Robot videos<br/>Fractal 87k · Bridge 25k · Kuka 86k] --> LAT
LAT --> Z[Discrete latent action chunks<br/>z_t..z_t+L-1<br/>perceptual + L1 + RAFT optical-flow loss]
Z --> PT[LWM-Chat-1M backbone G_θ<br/>predicts future tokens x_t+H<br/>+ latent action tokens z_t..z_t+H-1]
Img[o_t-1, o_t · task c] --> PT
PT --> FM["Flow-matching decoder H_η<br/>Beta(1.5, 1.0) interp.<br/>10 Euler steps"]
FM --> A[Continuous action chunk<br/>H=14 · up to 22 Hz]
Most internet and teleop videos are actionless but contain rich physical information. Prior latent-action methods (LAPA, UniVLA, Moto, GR00T-N1's latent actions) treat pretraining as autoregressive policy learning — they collapse "what changes in the world" and "how the robot moves" into one objective, use temporally coarse one-step latents, and discard the video-prediction signal. Action-labeled VLAs like OpenVLA and π₀ then need ~10k hours of robot trajectories to compensate. ViPRA's premise: separate the two objectives so action-label scarcity bottlenecks only the small finetune stage.
ViPRA is a three-stage pretraining-finetuning framework on top of LWM-Chat-1M (Liu et al., 2024).
Given a length-(L+1) clip o_{t:t+L} from human or robot video:
- An inverse-dynamics encoder I_β (DINOv2-initialized spatial encoder + 6-layer spatio-temporal transformer, 768 dim, 16 heads) maps the entire clip non-causally to per-step latent actions z_k = [z_k^1, ..., z_k^{N_latent}], each component quantized via NSVQ from a shared codebook of size |C| = 8 (token dim 32). Non-causal context lets z_k distinguish e.g. pickup from putdown.
- A forward decoder F_α (8 layers, 768 dim) predicts ô_{t+k+1} from o_{t:t+k} and z_{t:t+k}.
- Loss combines L1 pixel reconstruction L_rec, LPIPS perceptual loss (λ_LPIPS=0.5) and RAFT optical-flow consistency L_flow (λ_flow=0.1, activated only after α_flow=60k warm-up steps to avoid instability): L_flow = (1/L) Σ ‖OF(ô_k, ô_{k-1}) − OF(o_k, o_{k-1})‖₁ + symmetric forward term.
- Trained with AdamW (LR 1e-4, DINO encoder LR 1e-5, batch 128, 240k steps) on 8 H100 GPUs for 168 hours.
- LWM-Chat-1M backbone G_θ is augmented with a Latent Action Embedding head E_φ (4096-dim) and a 1-layer MLP latent-action decoder H_ψ (hidden 2048).
- Input (o_{t-1}, o_t) tokenized via LWM's frozen VQ-VAE; task description c is BPE-tokenized.
- Model autoregressively predicts the future frame x̂_{t+H} (H=14 steps ahead) plus the latent-action chunk z_{t:t+H-1} (3-6 Hz fine-grained, vs. one-step in LAPA).
- Loss: L_pretrain = L_img + L_act, both cross-entropy with teacher forcing.
- AdamW, LR 4e-5, batch 512, 50k steps, bfloat16, dropout 0.1, 8 H100 for 144 hours.
- Adds a 2-layer MLP noisy action encoder E_γ (hidden 4096) and a single linear flow decoder H_η.
- Input dim is 7 (delta-EEF) or 8 (absolute joint state).
- Flow matching with s ∼ Beta(1.5, 1.0), u_s = s·u₀ + (1-s)·a_{t:t+H-1}, loss L_FM = ‖a − u₀ − (1-s)·ĝ‖².
- At inference: forward Euler integration over 10 uniform steps from s=0 to s=1.
- Trained on only 100-200 teleop demonstrations per setting (12k steps for SIMPLER).
- Chunk length H = 14; deployed with KV-caching at 1.95 Hz per chunk → up to 22 Hz effective control rate.
| Task | Scratch-AR | VPT | OpenVLA | LAPA | ViPRA-AR | Scratch-FM | UniPI | π₀ | UniVLA | ViPRA-FM |
|---|---|---|---|---|---|---|---|---|---|---|
| StackG2Y | 54.2 | 45.8 | 25.0 | 33.3 | 66.7 | 16.7 | 2.7 | 0.0 | – | 54.2 |
| Carrot2Plate | 58.3 | 37.5 | 20.8 | 41.7 | 62.5 | 33.3 | 2.7 | 20.8 | – | 50.0 |
| Spoon2Cloth | 37.5 | 70.8 | 50.0 | 66.7 | 66.7 | 50.0 | 0.0 | 4.17 | – | 66.7 |
| Eggplant2Bask | 58.3 | 50.0 | 58.3 | 70.8 | 83.3 | 66.7 | 0.0 | 83.3 | – | 79.2 |
| Avg | 52.1 | 51.0 | 38.6 | 53.1 | 69.8 | 41.7 | 1.7 | 27.1 | 42.7 | 62.5 |
ViPRA-AR beats LAPA by +16.7 pts and OpenVLA by +31.2 pts. ViPRA-FM beats π₀ by +35.4 pts and UniVLA by +19.8 pts — the +16% headline number.
OpenVLA grasps StackG2Y 70.8% but only completes 25.0% (45.8 pt gap from contact-bin flipping). ViPRA-AR has zero gap (66.7 → 66.7); ViPRA-FM 8.3 pt gap.
| Method | Success |
|---|---|
| UniPI | 0.00 |
| OpenVLA | 0.54 |
| π₀-FAST | 0.60 |
| π₀ | 0.85 |
| UVA (LIBERO-tuned) | 0.90 |
| UniVLA | 0.92 |
| ViPRA-FM | 0.79 |
ViPRA-FM trails π₀, UVA, UniVLA on LIBERO-10. Authors attribute this to "delta-EEF drift" in image-only / no-proprio / no-wrist setup; π₀ uses proprio + wrist cam.
- ViPRA-FM 54.1% avg success vs π₀ 40.1% vs Scratch-FM 23.8% — +13% headline.
- Discrete policies excluded from real eval: bin flipping triggered Franka emergency brake on contact events.
- Closed-loop control rate capped at 3.5 Hz for safety; H=14 chunks, replan every 7 steps.
Future-state vs latent prediction (Table 2, SIMPLER avg):
| Variant | Pretrain | Finetune | Succ. |
|---|---|---|---|
| LAPA (1-step latent only) | 1-step L | 1-step A | 53.1 |
| ViPRA-AC (1-step latent + future state) | FS + 1-step L | 1-step A | 59.2 (AR) / 44.8 (FM) |
| ViPRA-LA (state-only) | FS only | H-step A | 60.7 |
| ViPRA-SP2 (no future state at pretrain) | H-step L | H-step A | 59.4 (AR) / 53.2 (FM) |
| ViPRA-AR (full) | FS + H-step L | H-step A | 69.8 |
| ViPRA-FM (full) | FS + H-step L | H-step A | 62.5 |
| ViPRA+SP3 (state at finetune) | H-step L | FS + H-step A | 53.1 (AR) / 31.3 (FM) |
Both future-state and chunked latents matter; collapsing chunked latents back to 1-step (–AC) is the single most damaging ablation for the continuous policy (62.5 → 44.8 FM, -17.7 pts), and predicting state at finetune compounds errors (esp. -31 pts FM) — pretrain-only state prediction is the right design.
Optical-flow loss (Table 4): removing L_flow raises codebook perplexity 5.01 → 5.63, entropy 1.59 → 1.74, action-probe MSE 0.84 → 0.92.
Data composition (Table 3, LIBERO-10 probe): human-only 0.69 / robot-only 0.72 / co-train 0.79 — human videos add motion diversity, robot videos ground latents.
- Embodiment / dexterity — generalizes WidowX → Franka (incl. bimanual) but not to humanoid dual-arm or multi-fingered hands; calls for tactile/force feedback or embodiment-specific adapters.
- Data scope — limited diversity of action-free corpora; ego-centric data, wrist cams, depth, proprio could help.
- Scaling behavior — open question: how passive video pretraining scales and where diminishing returns kick in. No scaling-law study yet.
- Predictive modeling not exploited — latent action decoder is also a world model usable for RL alignment / VLM-reward planning, but ViPRA does not yet do this.
- (Implicit, from LIBERO-10) image-only / no-proprio setup hurts long-horizon delta-EEF tasks.
Vs latent-action peers:
- LAPA uses 1-step temporally coarse VQ latents with no video prediction objective → ViPRA shows joint future-frame + chunked latents adds +16.7 pts on SIMPLER.
- UniVLA learns task-centric DINOv2-space latents; ViPRA's motion-centric fine-grained latents (3-6 Hz) outperform on SIMPLER (+19.8 pts FM) but trail on LIBERO-10 (where UniVLA is heavily tuned).
- Moto / GR00T-N1.5 latent actions — same family of "latent actions as pretraining bridge"; ViPRA argues optical-flow grounding + future-state co-prediction is the missing ingredient.
Vs action-labeled VLAs:
- OpenVLA / π₀ require ~970k OpenX trajectories and proprietary data; ViPRA finetunes on 100-200 demos and surpasses both on SIMPLER. π₀ remains stronger on long-horizon LIBERO-10 due to proprio + wrist-cam access.
- The 22 Hz figure puts ViPRA in the same control-frequency bracket as the Kim et al. 2025 7B speed-optimized OpenVLA-OFT — to authors' knowledge the only other 7B model at that rate.
Vs video-for-action peers:
- UniPI / VPT rely on inverse dynamics models from labeled data; ViPRA shows joint latent-action + future-state modeling (no IDM) gives stronger cross-environment transfer. UniPI scores 0.00 on LIBERO-10.
- Genie / DreamGen / Vid2World dream frames at rollout; ViPRA distills video into latent actions used at training time only — much cheaper at inference. See Vid2World, DreamGen.
The cleaner separation — what changes (video dynamics) vs how the robot moves (action decoding) — is the structural argument. Action-label scarcity now bottlenecks only the small finetune stage.
- OpenReview: https://openreview.net/forum?id=w3Ik8HUyTT
- Project page: https://vipra-project.github.io
- Code & latent-action-labeled pretraining data released
- Genie Envisioner
- Vid2World
- Human Video Pretraining
- DreamGen
- VillaX — also enhances latent-action modeling
← Back to ICLR-2026