Review Independent Visual Representation - Heungwoo/research GitHub Wiki

In-Depth Review — Independent Visual Representation in VLAs (vision built outside the VLM)

Compiled June 2026 · Thesis: a VLM's own vision encoder (CLIP/SigLIP-style, trained for language-semantic alignment) is increasingly shown to be insufficient for action, so a fast-growing body of work builds a separate visual representation — geometric, predictive, temporal, or dense — and feeds it to the policy. This review collects that work and analyzes why the independent representation is necessary and how it is injected.

Companion reviews: VLA Architectures · VLA Architecture Categories (Category L, E1–E5) · VLM↔Action Connection · World Models · RoboMME (memory) · WAM vs VLA Robustness.


1. TL;DR — vision optimized for language is not vision optimized for action

Almost every VLA inherits its eyes from a VLM: a CLIP/SigLIP-style encoder trained to make images align with text. That objective rewards "what is in the scene" (semantics) and discards much of what control needs: geometry/depth, temporal dynamics, history, and fine pixel-level detail. Two pieces of direct evidence make the gap concrete:

  • The depth probe (Spatial Forcing). Freeze an OpenVLA-OFT policy's visual embeddings and train only a DPT depth head on them → the predicted depth is blurry and wrong. Conclusion in the paper: visual embeddings learned from 2D RGB action-imitation do not encode usable spatial structure.
  • Grounding collapse (ST4VLA). Naively fine-tuning a VLM into a VLA destroys its spatial grounding — RefCOCO-g [email protected] decays to near-random within ~20k action-only steps, attributed to gradient-subspace misalignment between the grounding objective and the action objective.

So the field responds in one of two ways: protect the VLM's vision (don't let action gradients corrupt it — "knowledge insulation") and/or supplement it with a representation built specifically for control. This review is about the second move. The deficiencies cluster into four kinds of independent representation:

flowchart TB
  Q{What does the VLM vision miss?}
  Q -- geometry / 3D / depth --> G[G. Spatial-geometric reps<br/>Spatial Forcing · FALCON · Any3D-VLA · PA3FF]
  Q -- future / dynamics --> P[P. Predictive world-model latents<br/>V-JEPA 2-AC · DINO-WM · DreamVLA · ViPRA · CoWVLA]
  Q -- history / time --> M[M. Temporal memory reps<br/>RoboMME perceptual memory · MemoryVLA]
  Q -- dense / control detail --> D[D. Dedicated dense encoders<br/>OpenVLA dual-enc · VER · PixelVLA · XR-1]
Loading

The single most useful lens: what the representation captures (G/P/M/D) × how aggressively it is separated from the VLM (a ladder from a train-time alignment loss all the way to replacing the VLM's vision entirely).


2. Why the VLM's vision is insufficient (the necessity argument)

Missing for action Why the VLM loses it What the field adds
3D geometry / depth / metric space CLIP/SigLIP are 2D-pretrained on web image-text; depth/pose are never a target G: align to / fuse a 3D foundation prior (VGGT, Sonata/PTv3), or lift point clouds
Future dynamics A VLM encodes the present frame; it has no forward model P: predict future latents (JEPA), future frames/4D, or motion latents
History / temporal context A VLA is largely Markovian — it sees the current frame only M: a separate visual-token memory of past frames
Dense, control-relevant pixel detail Patch-level, language-biased tokens miss thin parts, small objects, contact cues D: a dedicated SSL/pixel encoder (DINOv2, distilled vision experts)
(meta) stability under action fine-tuning Action-head gradients corrupt VLM features / collapse grounding keep vision separate or frozen, inject via aux-loss or modulation

This is why "independent visual representation" is not a niche trick but a structural response: the representation a VLM ships with was optimized for a different task. The wiki already encodes this as Category L (3D-foundation-aligned VLAs) and Categories E1–E5 (world-model / video sub-patterns) in Review-VLA-Architecture-Categories; this page reads those plus the memory and dense-encoder lines under one "independent vision" thesis.


3. Axis 1 — What the independent representation captures

G. Spatial / geometric representations (fill the 3D gap)

The largest cluster. A 2D VLA borrows or builds a 3D-aware signal so it can reason about where things are, not just what they are.

Paper Venue Independent rep How built Injection Headline
Spatial Forcing ICLR 2026 (OR euMVC1DO4k) per-pixel spatial features of a frozen VGGT 3D model cosine alignment loss between VLA tokens and VGGT features train-time aux loss only; inference unchanged LIBERO 98.5%; 3.8× faster train, 5.9× more data-efficient
FALCON ICLR 2026 (OR fzmittHfq3) Embodied Spatial Model (~1B, VGGT-style) spatial tokens separate ESM encoder w/ depth/pose/pointmap supervision spatial tokens added to the VLM action token → action head CALVIN ABC→D 4.40; real +25.6 vs SpatialVLA
Any3D-VLA ICML 2026 (2602.00807) point clouds from sim + sensor + model-estimated depth pre-trained point encoder; hybrid 3-source training 3D embeddings fused with 2D patches into the VLA zero-shot real 62.5% (+29.2)
ST4VLA / SP-VLA ICLR 2026 (OR eKhOrQWAVJ) spatial-grounding pretraining + a DiT actor with its own DINOv2 3M-sample spatial QA pretrain; querying transformer (0.5 gradient decay) dual-system: VLM planner → DiT actor's visual encoder SimplerEnv 84.6% (vs SpatialVLA 75.1)
PA3FF (not a VLA) ICLR 2026 (OR qXfRXfAHOK) part-aware dense 3D feature field (Sonata/PTv3 + contrastive) self-sup + part/semantic contrastive refinement frozen 3D field → diffusion policy (PADP) PartInstruct +9.4 vs GenDP; real unseen 53.75%

Adjacent non-VLM exemplars (build vision fully separately): EquAct (SE(3)-equivariant point-cloud transformer, 2505.21351), Cortical Policy (dual-stream VGGT-keypoint + gaze, RVT-2 lineage), GeoMoLa (geometry-aware motion latents from 4D point-cloud prediction). These show the pure form of the idea — vision is its own geometric module, the VLM is absent.

Key finding: the most efficient G-variant is the train-time alignment loss (Spatial Forcing): it injects a 3D prior with zero inference cost and no architecture change, and the data-efficiency gains (5.9×) suggest much of a VLA's data budget is otherwise spent re-learning geometry the VLM threw away. FALCON's own framing warns the opposite extreme — forcing explicit 3D input — "disrupts the VLM's pretrained vision-language alignment."

P. Predictive / world-model latents (fill the future gap)

Here the independent representation is a forward signal — predict the future in a learned space, not pixels. This is the JEPA thesis applied to control. (Full treatment: Review-World-Models Category C; robustness: Review-WAM-vs-VLA-Robustness.)

Paper Venue Independent rep Injection Headline
V-JEPA 2 / 2-AC 2506.09985 future latent embeddings (frozen feature space) plan by latent goal-matching zero-shot Franka pick-place on <62 h robot data
DINO-WM 2411.04983 future DINOv2 patch features trajectory optimization in DINO space zero-shot planning
DreamVLA NeurIPS 2025 structured cues: dynamic mask + depth + DINOv2/SAM aux forecast heads → inverse-dynamics → diffusion head real Franka 76.7%; CALVIN 4.44
ViPRA ICLR 2026 latent actions (DINOv2 spatio-temporal VQ) + future frame latent actions are the policy interface; frame-pred is pretrain-only +16.7 over LAPA on SIMPLER
CoWVLA CVPR 2026 (2603.03195) motion-latent chain (VidTwin video-VAE) + terminal latent terminal latent = subgoal for the policy LIBERO 95.6%
MoLA ICML 2026 (2605.12167) mixture of latent actions from semantic/depth/flow IDMs one-step video diffusion → latent mixture → diffusion head CALVIN 4.55; real UR5e 73.0%
Geometry-4D-Video ICLR 2026 joint RGB + pointmap latent video (cross-view 3D) generated 4D video → pose → IK (geometry-first) novel-view 0.64 vs DP3 0.25
VLA-JEPA hybrid (see Review-WAM-vs-VLA-Robustness) future-state predictive-encoder alignment on Qwen3-VL-2B aux loss on the VLM backbone LIBERO-Plus 77.9% (between VLA & WAM)

Key finding: the recurring justification is "pixels waste capacity" — predicting RGB spends model budget on appearance/background irrelevant to control, and destabilizes long rollouts. Predicting a latent (DINO features, motion latents, pointmaps) is cheaper, more data-efficient, and (V-JEPA/DINO-WM) enables zero-shot planning. The cost: latent reps are opaque and weaker on fine-grained contact.

M. Temporal memory representations (fill the history gap)

A VLA is largely Markovian; these add a separate visual memory of past frames — kept out of the VLM so it doesn't perturb the pretrained encoder.

Paper Venue Independent rep Injection Headline
RoboMME (perceptual memory) ICML 2026 Oral (2603.04639) past-frame SigLIP visual tokens (FrameSamp) AdaLN modulation of the action expert only; VLM untouched FrameSamp+Modul 44.51% (vs no-mem 17.93%)
MemoryVLA ICLR 2026 (2508.19236) perceptual + cognitive memory banks (256 visual tokens + 1 latent/step) separate retrieval + gated fusion → DiT action head real 12-task 84.0%; long-horizon +26

Key finding (and the link to JEPA the user noted): both keep a visual representation separate from the VLM token stream and feed it to the action side. RoboMME's perceptual memory is a backward (retain-the-past) latent visual stream injected by AdaLN — structurally a "second visual representation," but not JEPA: it has no predictive objective. P-cluster latents are the forward (predict-the-future) mirror image. Stacking both = a policy conditioned on a past-memory latent and a future-prediction latent, neither touching the VLM. (See Review-RoboMME §3.4 for the AdaLN mechanism.)

D. Dedicated dense / SSL encoders (fill the control-detail gap)

The oldest thread: don't trust a single language-aligned encoder; add a self-supervised, dense, or distilled visual representation.

Paper Venue Independent rep Injection Headline
OpenVLA (seed) CoRL 2024 (2406.09246) DINOv2 (spatial) + SigLIP (semantic) dual encoder concatenated pre-LLM (one VLM, not a separate pathway) beats RT-2-X (55B) by +16.5 at 7B
VER ICLR 2026 Vision Expert Library distilled from DINOv2+CLIP+ViT, MoE-routed dedicated vision encoder → any policy head (no VLM) 11-task 74.7% (vs Theia 67.1)
PixelVLA ICLR 2026 (2511.01571) multiscale pixel-aware + SAM visual-prompt encoders added to a frozen OpenVLA via adapters SimplerEnv 61.4% (vs OpenVLA 32.7); 1.5% of its train cost
XR-1 ICLR 2026 Oral Unified Vision-Motion Codes (dual-branch VQ-VAE codebook) shared discrete codebook as embodiment-free interface UR-5e ~72% (vs π0.5 ~62); >14k rollouts

Key finding: OpenVLA's DINOv2+SigLIP pairing is the seed of the whole thesis — "where" (DINO, spatial/self-supervised) + "what" (SigLIP, semantic) — but it still fuses pre-LLM into one VLM. VER and XR-1 push it further: a vision representation with its own training objective (distillation / vision-motion VQ) that is architecturally independent of the language model.


4. Axis 2 — How the independent representation is injected (a ladder of separation)

Ordering from least to most aggressive separation from the VLM — mirroring Review-VLA-Architecture-Categories' E1→E5 "progressively more is ceded to vision":

  1. Train-time alignment / auxiliary loss (inference unchanged): Spatial Forcing (align to VGGT), DreamVLA (forecast cues), VLA-JEPA (future-state aux), From-Pixels-to-Tokens (latent-action aux). Cheapest; preserves the VLM; "free grounding."
  2. Parallel visual-token stream fused into the policy: FALCON (ESM tokens + action token), Any3D-VLA (point-cloud + 2D patches), MemoryVLA (memory bank → DiT), PixelVLA (encoders → frozen VLM adapters).
  3. Modulation of the action expert (memory/vision as scale/shift, VLM untouched): RoboMME perceptual memory (AdaLN).
  4. Latent as the policy I/O interface: ViPRA (latent actions), CoWVLA (terminal motion latent), MoLA (latent-action mixture), XR-1 (vision-motion codes), V-JEPA-AC / DINO-WM (plan in latent).
  5. Replace VLM vision / geometry-first → IK (skip action tokens): Geometry-4D-Video, GeoMoLa, EquAct, GE-Act-style video backbones (Genie Envisioner, Cosmos Policy).

The trade is monotonic: higher rungs gain more control-specific structure but cede more of the VLM's pretrained language grounding and interpretability. Rung 1 is winning on cost/efficiency right now (Spatial Forcing); rungs 4–5 are winning on data-scaling and zero-shot (ViPRA, V-JEPA-AC).


5. Cross-paper synthesis

  1. The VLM encoder is a language artifact, and the field has stopped pretending otherwise. Two independent demonstrations — Spatial Forcing's failed depth probe and ST4VLA's grounding collapse — show a VLA's visual embeddings neither encode geometry nor survive action fine-tuning. Independent vision is the structural fix.
  2. "Preserve the VLM" is the dominant constraint. Alignment-loss (rung 1), AdaLN-modulation (RoboMME), frozen-VLM adapters (PixelVLA), and gradient-decay connectors (ST4VLA) all exist to add a visual signal without corrupting pretrained vision-language knowledge — the same instinct as knowledge insulation.
  3. Predict a latent, not pixels. Across P-cluster the justification is identical — appearance is wasted capacity; control lives in geometry/motion latents. This is exactly the JEPA argument, and it is why memory (RoboMME) and world-models (V-JEPA) are mirror-image latent visual streams (backward vs forward) feeding the action side.
  4. Geometry is the most-cited single deficiency. G is the biggest cluster; the cheapest fix (align to a frozen 3D model) already matches or beats methods that take real 3D input — suggesting the bottleneck is representation, not sensing.
  5. The endpoint is "vision as its own foundation model." VER (distilled vision experts), PA3FF (3D feature field), V-JEPA (latent world model) point at a future where the VLA's vision is a separately-trained module the language model queries — not the language model's own encoder.

6. Open tensions

  • Does the separate representation directly improve action, or only regularize features? Review-VLA-Architecture §8 flags this for E1 explicitly; alignment-loss gains could be representation-quality, not action-quality. Few papers ablate this cleanly.
  • Train-time alignment vs run-time fusion. Spatial Forcing (rung 1) gets 3D grounding for free at inference; FALCON/Any3D pay run-time cost for explicit 3D tokens. When is the extra inference cost worth it? Largely unmeasured head-to-head.
  • Opacity & contact. Latent/JEPA reps win cost and zero-shot but are "not interpretable" and weak on fine contact (Review-World-Models §6) — exactly where dense/tactile reps (Review-Tactile-VLA) are needed.
  • No shared benchmark. Geometry, prediction, and memory reps are each evaluated on different suites (LIBERO / SIMPLER / CALVIN / RoboTwin / RLBench / custom memory tasks), so the relative value of the four deficiencies is not measured on common ground.

7. Decision guide

  1. Cheapest spatial grounding, no inference cost? → rung-1 alignment loss to a frozen 3D model (Spatial Forcing).
  2. Need explicit 3D robustness (viewpoint/occlusion)? → fuse point clouds / ESM tokens (Any3D-VLA, FALCON).
  3. Data-scarce, want internet/human-video pretraining? → predictive latent interface (ViPRA, V-JEPA 2-AC) or motion-latent subgoals (CoWVLA).
  4. History-dependent tasks (counting, occlusion, motion imitation)? → a separate visual memory (RoboMME perceptual memory, MemoryVLA).
  5. Pixel-precise / promptable manipulation? → a dense/pixel encoder on a frozen VLM (PixelVLA); or distilled vision experts (VER).
  6. Cross-embodiment from video? → a shared vision-motion codebook (XR-1).

8. Links & related


🗓 State of the Field (updated Aug 2026)

Verdict: the vision encoder is the proven VLA bottleneck; the one intervention that reliably helps is action-supervised fine-tuning of the ViT — not more VQA.

📈 Trend

VLM4VLA (ICLR 2026) settled the diagnosis: general VLM scores don't predict control (r ≈ −0.36 on SimplerEnv); freezing the ViT is catastrophic (−21…−42 pp); action-supervised vision fine-tuning gives +18.1 pp. 2026 responses fork into geometry add-ons (Robo3R, StereoVLA, PointACT, RobotManip's CaPE) and dynamics-aligned encoders (DINO-latent in LDA-1B, video backbones in mimic-video).

⚖️ Approaches & trade-offs

Approach Pros Cons
Action-align the VLM ViT (end-to-end grads or FAST-token FT) Cheapest, best-evidenced; implicitly what flagships now do Recipe not standardized
Bolt-on metric 3D / stereo / camera-geometry encodings Real spatial-OOD gains Calibration + pipeline complexity
Swap toward dynamics encoders (DINO-latent, video) Physics priors without pixel modeling Overlaps the VAM bet; loses VLM semantics
ICML 2026 additions — 3D/motion-structured inputs Any3D-VLA mixes sim/sensor/estimated point clouds for domain-agnostic 3D; GeoMoLa learns motion latents from point-cloud evolution (4D) rather than appearance; Fourier features beat spectral bias for high-precision IL; RS-CL aligns features to proprioception (45.0→58.3% real) Each targets one failure mode; no unified recipe

⚠️ Limitations & open problems

  • In-distribution benchmarks barely register the choice — effects appear only under spatial/OOD stress.
  • No standardized action-supervised ViT fine-tuning recipe exists.
  • Geometry add-ons assume calibration quality that in-the-wild deployments lack.

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️