Review Independent Visual Representation - Heungwoo/research GitHub Wiki
Compiled June 2026 · Thesis: a VLM's own vision encoder (CLIP/SigLIP-style, trained for language-semantic alignment) is increasingly shown to be insufficient for action, so a fast-growing body of work builds a separate visual representation — geometric, predictive, temporal, or dense — and feeds it to the policy. This review collects that work and analyzes why the independent representation is necessary and how it is injected.
Companion reviews: VLA Architectures · VLA Architecture Categories (Category L, E1–E5) · VLM↔Action Connection · World Models · RoboMME (memory) · WAM vs VLA Robustness.
Almost every VLA inherits its eyes from a VLM: a CLIP/SigLIP-style encoder trained to make images align with text. That objective rewards "what is in the scene" (semantics) and discards much of what control needs: geometry/depth, temporal dynamics, history, and fine pixel-level detail. Two pieces of direct evidence make the gap concrete:
- The depth probe (Spatial Forcing). Freeze an OpenVLA-OFT policy's visual embeddings and train only a DPT depth head on them → the predicted depth is blurry and wrong. Conclusion in the paper: visual embeddings learned from 2D RGB action-imitation do not encode usable spatial structure.
- Grounding collapse (ST4VLA). Naively fine-tuning a VLM into a VLA destroys its spatial grounding — RefCOCO-g [email protected] decays to near-random within ~20k action-only steps, attributed to gradient-subspace misalignment between the grounding objective and the action objective.
So the field responds in one of two ways: protect the VLM's vision (don't let action gradients corrupt it — "knowledge insulation") and/or supplement it with a representation built specifically for control. This review is about the second move. The deficiencies cluster into four kinds of independent representation:
flowchart TB
Q{What does the VLM vision miss?}
Q -- geometry / 3D / depth --> G[G. Spatial-geometric reps<br/>Spatial Forcing · FALCON · Any3D-VLA · PA3FF]
Q -- future / dynamics --> P[P. Predictive world-model latents<br/>V-JEPA 2-AC · DINO-WM · DreamVLA · ViPRA · CoWVLA]
Q -- history / time --> M[M. Temporal memory reps<br/>RoboMME perceptual memory · MemoryVLA]
Q -- dense / control detail --> D[D. Dedicated dense encoders<br/>OpenVLA dual-enc · VER · PixelVLA · XR-1]
The single most useful lens: what the representation captures (G/P/M/D) × how aggressively it is separated from the VLM (a ladder from a train-time alignment loss all the way to replacing the VLM's vision entirely).
| Missing for action | Why the VLM loses it | What the field adds |
|---|---|---|
| 3D geometry / depth / metric space | CLIP/SigLIP are 2D-pretrained on web image-text; depth/pose are never a target | G: align to / fuse a 3D foundation prior (VGGT, Sonata/PTv3), or lift point clouds |
| Future dynamics | A VLM encodes the present frame; it has no forward model | P: predict future latents (JEPA), future frames/4D, or motion latents |
| History / temporal context | A VLA is largely Markovian — it sees the current frame only | M: a separate visual-token memory of past frames |
| Dense, control-relevant pixel detail | Patch-level, language-biased tokens miss thin parts, small objects, contact cues | D: a dedicated SSL/pixel encoder (DINOv2, distilled vision experts) |
| (meta) stability under action fine-tuning | Action-head gradients corrupt VLM features / collapse grounding | keep vision separate or frozen, inject via aux-loss or modulation |
This is why "independent visual representation" is not a niche trick but a structural response: the representation a VLM ships with was optimized for a different task. The wiki already encodes this as Category L (3D-foundation-aligned VLAs) and Categories E1–E5 (world-model / video sub-patterns) in Review-VLA-Architecture-Categories; this page reads those plus the memory and dense-encoder lines under one "independent vision" thesis.
The largest cluster. A 2D VLA borrows or builds a 3D-aware signal so it can reason about where things are, not just what they are.
| Paper | Venue | Independent rep | How built | Injection | Headline |
|---|---|---|---|---|---|
| Spatial Forcing | ICLR 2026 (OR euMVC1DO4k) | per-pixel spatial features of a frozen VGGT 3D model | cosine alignment loss between VLA tokens and VGGT features | train-time aux loss only; inference unchanged | LIBERO 98.5%; 3.8× faster train, 5.9× more data-efficient |
| FALCON | ICLR 2026 (OR fzmittHfq3) | Embodied Spatial Model (~1B, VGGT-style) spatial tokens | separate ESM encoder w/ depth/pose/pointmap supervision | spatial tokens added to the VLM action token → action head | CALVIN ABC→D 4.40; real +25.6 vs SpatialVLA |
| Any3D-VLA | ICML 2026 (2602.00807) | point clouds from sim + sensor + model-estimated depth | pre-trained point encoder; hybrid 3-source training | 3D embeddings fused with 2D patches into the VLA | zero-shot real 62.5% (+29.2) |
| ST4VLA / SP-VLA | ICLR 2026 (OR eKhOrQWAVJ) | spatial-grounding pretraining + a DiT actor with its own DINOv2 | 3M-sample spatial QA pretrain; querying transformer (0.5 gradient decay) | dual-system: VLM planner → DiT actor's visual encoder | SimplerEnv 84.6% (vs SpatialVLA 75.1) |
| PA3FF (not a VLA) | ICLR 2026 (OR qXfRXfAHOK) | part-aware dense 3D feature field (Sonata/PTv3 + contrastive) | self-sup + part/semantic contrastive refinement | frozen 3D field → diffusion policy (PADP) | PartInstruct +9.4 vs GenDP; real unseen 53.75% |
Adjacent non-VLM exemplars (build vision fully separately): EquAct (SE(3)-equivariant point-cloud transformer, 2505.21351), Cortical Policy (dual-stream VGGT-keypoint + gaze, RVT-2 lineage), GeoMoLa (geometry-aware motion latents from 4D point-cloud prediction). These show the pure form of the idea — vision is its own geometric module, the VLM is absent.
Key finding: the most efficient G-variant is the train-time alignment loss (Spatial Forcing): it injects a 3D prior with zero inference cost and no architecture change, and the data-efficiency gains (5.9×) suggest much of a VLA's data budget is otherwise spent re-learning geometry the VLM threw away. FALCON's own framing warns the opposite extreme — forcing explicit 3D input — "disrupts the VLM's pretrained vision-language alignment."
Here the independent representation is a forward signal — predict the future in a learned space, not pixels. This is the JEPA thesis applied to control. (Full treatment: Review-World-Models Category C; robustness: Review-WAM-vs-VLA-Robustness.)
| Paper | Venue | Independent rep | Injection | Headline |
|---|---|---|---|---|
| V-JEPA 2 / 2-AC | 2506.09985 | future latent embeddings (frozen feature space) | plan by latent goal-matching | zero-shot Franka pick-place on <62 h robot data |
| DINO-WM | 2411.04983 | future DINOv2 patch features | trajectory optimization in DINO space | zero-shot planning |
| DreamVLA | NeurIPS 2025 | structured cues: dynamic mask + depth + DINOv2/SAM | aux forecast heads → inverse-dynamics → diffusion head | real Franka 76.7%; CALVIN 4.44 |
| ViPRA | ICLR 2026 | latent actions (DINOv2 spatio-temporal VQ) + future frame | latent actions are the policy interface; frame-pred is pretrain-only | +16.7 over LAPA on SIMPLER |
| CoWVLA | CVPR 2026 (2603.03195) | motion-latent chain (VidTwin video-VAE) + terminal latent | terminal latent = subgoal for the policy | LIBERO 95.6% |
| MoLA | ICML 2026 (2605.12167) | mixture of latent actions from semantic/depth/flow IDMs | one-step video diffusion → latent mixture → diffusion head | CALVIN 4.55; real UR5e 73.0% |
| Geometry-4D-Video | ICLR 2026 | joint RGB + pointmap latent video (cross-view 3D) | generated 4D video → pose → IK (geometry-first) | novel-view 0.64 vs DP3 0.25 |
| VLA-JEPA | hybrid (see Review-WAM-vs-VLA-Robustness) | future-state predictive-encoder alignment on Qwen3-VL-2B | aux loss on the VLM backbone | LIBERO-Plus 77.9% (between VLA & WAM) |
Key finding: the recurring justification is "pixels waste capacity" — predicting RGB spends model budget on appearance/background irrelevant to control, and destabilizes long rollouts. Predicting a latent (DINO features, motion latents, pointmaps) is cheaper, more data-efficient, and (V-JEPA/DINO-WM) enables zero-shot planning. The cost: latent reps are opaque and weaker on fine-grained contact.
A VLA is largely Markovian; these add a separate visual memory of past frames — kept out of the VLM so it doesn't perturb the pretrained encoder.
| Paper | Venue | Independent rep | Injection | Headline |
|---|---|---|---|---|
| RoboMME (perceptual memory) | ICML 2026 Oral (2603.04639) | past-frame SigLIP visual tokens (FrameSamp) | AdaLN modulation of the action expert only; VLM untouched | FrameSamp+Modul 44.51% (vs no-mem 17.93%) |
| MemoryVLA | ICLR 2026 (2508.19236) | perceptual + cognitive memory banks (256 visual tokens + 1 latent/step) | separate retrieval + gated fusion → DiT action head | real 12-task 84.0%; long-horizon +26 |
Key finding (and the link to JEPA the user noted): both keep a visual representation separate from the VLM token stream and feed it to the action side. RoboMME's perceptual memory is a backward (retain-the-past) latent visual stream injected by AdaLN — structurally a "second visual representation," but not JEPA: it has no predictive objective. P-cluster latents are the forward (predict-the-future) mirror image. Stacking both = a policy conditioned on a past-memory latent and a future-prediction latent, neither touching the VLM. (See Review-RoboMME §3.4 for the AdaLN mechanism.)
The oldest thread: don't trust a single language-aligned encoder; add a self-supervised, dense, or distilled visual representation.
| Paper | Venue | Independent rep | Injection | Headline |
|---|---|---|---|---|
| OpenVLA (seed) | CoRL 2024 (2406.09246) | DINOv2 (spatial) + SigLIP (semantic) dual encoder | concatenated pre-LLM (one VLM, not a separate pathway) | beats RT-2-X (55B) by +16.5 at 7B |
| VER | ICLR 2026 | Vision Expert Library distilled from DINOv2+CLIP+ViT, MoE-routed | dedicated vision encoder → any policy head (no VLM) | 11-task 74.7% (vs Theia 67.1) |
| PixelVLA | ICLR 2026 (2511.01571) | multiscale pixel-aware + SAM visual-prompt encoders | added to a frozen OpenVLA via adapters | SimplerEnv 61.4% (vs OpenVLA 32.7); 1.5% of its train cost |
| XR-1 | ICLR 2026 Oral | Unified Vision-Motion Codes (dual-branch VQ-VAE codebook) | shared discrete codebook as embodiment-free interface | UR-5e ~72% (vs π0.5 ~62); >14k rollouts |
Key finding: OpenVLA's DINOv2+SigLIP pairing is the seed of the whole thesis — "where" (DINO, spatial/self-supervised) + "what" (SigLIP, semantic) — but it still fuses pre-LLM into one VLM. VER and XR-1 push it further: a vision representation with its own training objective (distillation / vision-motion VQ) that is architecturally independent of the language model.
Ordering from least to most aggressive separation from the VLM — mirroring Review-VLA-Architecture-Categories' E1→E5 "progressively more is ceded to vision":
- Train-time alignment / auxiliary loss (inference unchanged): Spatial Forcing (align to VGGT), DreamVLA (forecast cues), VLA-JEPA (future-state aux), From-Pixels-to-Tokens (latent-action aux). Cheapest; preserves the VLM; "free grounding."
- Parallel visual-token stream fused into the policy: FALCON (ESM tokens + action token), Any3D-VLA (point-cloud + 2D patches), MemoryVLA (memory bank → DiT), PixelVLA (encoders → frozen VLM adapters).
- Modulation of the action expert (memory/vision as scale/shift, VLM untouched): RoboMME perceptual memory (AdaLN).
- Latent as the policy I/O interface: ViPRA (latent actions), CoWVLA (terminal motion latent), MoLA (latent-action mixture), XR-1 (vision-motion codes), V-JEPA-AC / DINO-WM (plan in latent).
- Replace VLM vision / geometry-first → IK (skip action tokens): Geometry-4D-Video, GeoMoLa, EquAct, GE-Act-style video backbones (Genie Envisioner, Cosmos Policy).
The trade is monotonic: higher rungs gain more control-specific structure but cede more of the VLM's pretrained language grounding and interpretability. Rung 1 is winning on cost/efficiency right now (Spatial Forcing); rungs 4–5 are winning on data-scaling and zero-shot (ViPRA, V-JEPA-AC).
- The VLM encoder is a language artifact, and the field has stopped pretending otherwise. Two independent demonstrations — Spatial Forcing's failed depth probe and ST4VLA's grounding collapse — show a VLA's visual embeddings neither encode geometry nor survive action fine-tuning. Independent vision is the structural fix.
- "Preserve the VLM" is the dominant constraint. Alignment-loss (rung 1), AdaLN-modulation (RoboMME), frozen-VLM adapters (PixelVLA), and gradient-decay connectors (ST4VLA) all exist to add a visual signal without corrupting pretrained vision-language knowledge — the same instinct as knowledge insulation.
- Predict a latent, not pixels. Across P-cluster the justification is identical — appearance is wasted capacity; control lives in geometry/motion latents. This is exactly the JEPA argument, and it is why memory (RoboMME) and world-models (V-JEPA) are mirror-image latent visual streams (backward vs forward) feeding the action side.
- Geometry is the most-cited single deficiency. G is the biggest cluster; the cheapest fix (align to a frozen 3D model) already matches or beats methods that take real 3D input — suggesting the bottleneck is representation, not sensing.
- The endpoint is "vision as its own foundation model." VER (distilled vision experts), PA3FF (3D feature field), V-JEPA (latent world model) point at a future where the VLA's vision is a separately-trained module the language model queries — not the language model's own encoder.
- Does the separate representation directly improve action, or only regularize features? Review-VLA-Architecture §8 flags this for E1 explicitly; alignment-loss gains could be representation-quality, not action-quality. Few papers ablate this cleanly.
- Train-time alignment vs run-time fusion. Spatial Forcing (rung 1) gets 3D grounding for free at inference; FALCON/Any3D pay run-time cost for explicit 3D tokens. When is the extra inference cost worth it? Largely unmeasured head-to-head.
- Opacity & contact. Latent/JEPA reps win cost and zero-shot but are "not interpretable" and weak on fine contact (Review-World-Models §6) — exactly where dense/tactile reps (Review-Tactile-VLA) are needed.
- No shared benchmark. Geometry, prediction, and memory reps are each evaluated on different suites (LIBERO / SIMPLER / CALVIN / RoboTwin / RLBench / custom memory tasks), so the relative value of the four deficiencies is not measured on common ground.
- Cheapest spatial grounding, no inference cost? → rung-1 alignment loss to a frozen 3D model (Spatial Forcing).
- Need explicit 3D robustness (viewpoint/occlusion)? → fuse point clouds / ESM tokens (Any3D-VLA, FALCON).
- Data-scarce, want internet/human-video pretraining? → predictive latent interface (ViPRA, V-JEPA 2-AC) or motion-latent subgoals (CoWVLA).
- History-dependent tasks (counting, occlusion, motion imitation)? → a separate visual memory (RoboMME perceptual memory, MemoryVLA).
- Pixel-precise / promptable manipulation? → a dense/pixel encoder on a frozen VLM (PixelVLA); or distilled vision experts (VER).
- Cross-embodiment from video? → a shared vision-motion codebook (XR-1).
- Framing companions: Review-VLA-Architecture-Categories (Category L = 3D-foundation-aligned; E1–E5 = world-model/video) · Review-VLM-Action-Connection (injection mechanisms, knowledge insulation) · Review-World-Models (pixels vs latent/JEPA) · Review-WAM-vs-VLA-Robustness (VLA-JEPA, video backbones)
- G (spatial): Spatial Forcing · FALCON · ST4VLA · Any3D-VLA · PA3FF · EquAct · Cortical Policy
- P (predictive): DreamVLA · ViPRA · CoWVLA · MoLA · Geometry-4D-Video · Sparse Imagination
- M (memory): RoboMME · MemoryVLA
- D (dense/SSL): OpenVLA · VER · PixelVLA · XR-1 · From Pixels to Tokens
Verdict: the vision encoder is the proven VLA bottleneck; the one intervention that reliably helps is action-supervised fine-tuning of the ViT — not more VQA.
VLM4VLA (ICLR 2026) settled the diagnosis: general VLM scores don't predict control (r ≈ −0.36 on SimplerEnv); freezing the ViT is catastrophic (−21…−42 pp); action-supervised vision fine-tuning gives +18.1 pp. 2026 responses fork into geometry add-ons (Robo3R, StereoVLA, PointACT, RobotManip's CaPE) and dynamics-aligned encoders (DINO-latent in LDA-1B, video backbones in mimic-video).
| Approach | Pros | Cons |
|---|---|---|
| Action-align the VLM ViT (end-to-end grads or FAST-token FT) | Cheapest, best-evidenced; implicitly what flagships now do | Recipe not standardized |
| Bolt-on metric 3D / stereo / camera-geometry encodings | Real spatial-OOD gains | Calibration + pipeline complexity |
| Swap toward dynamics encoders (DINO-latent, video) | Physics priors without pixel modeling | Overlaps the VAM bet; loses VLM semantics |
| ICML 2026 additions — 3D/motion-structured inputs | Any3D-VLA mixes sim/sensor/estimated point clouds for domain-agnostic 3D; GeoMoLa learns motion latents from point-cloud evolution (4D) rather than appearance; Fourier features beat spectral bias for high-precision IL; RS-CL aligns features to proprioception (45.0→58.3% real) | Each targets one failure mode; no unified recipe |
- In-distribution benchmarks barely register the choice — effects appear only under spatial/OOD stress.
- No standardized action-supervised ViT fine-tuning recipe exists.
- Geometry add-ons assume calibration quality that in-the-wild deployments lack.
← Back to Home