ICLR 2026 PixelVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture — Perception / prompting Trend tag: Pixel-level grounding · visual prompting
flowchart LR
Img[Image] --> Enc[Multiscale pixel-aware encoder]
VP[Visual prompts<br/>clicks, masks] --> VPE[SAM visual prompting encoder]
T[Text prompt] --> LM[Frozen VLA backbone<br/>OpenVLA/Prismatic-7B + LoRA]
Enc --> LM
VPE --> LM
LM --> A[Continuous action decoder<br/>7-DoF action]
See Figure 1 of the original paper for the authors' own architecture diagram (link below).
Current VLAs (i) lack pixel-level scene understanding — they reason at coarse patch granularity and miss fine spatial detail — and (ii) depend almost entirely on text prompts, limiting flexibility when a user wants to specify "this object" by clicking or masking it. PixelVLA targets both.
A visuomotor instruction-tuning framework combining a multiscale pixel-aware encoder with a visual prompting encoder (a lightweight SAM prompt encoder), so the model accepts text and visual cues (clicks, regions, masks). Rather than training from scratch, PixelVLA is built on top of an existing VLA (here OpenVLA / Prismatic-7B with a Llama-2-7B backbone, DINOv2+SigLIP vision): the original weights are frozen and only the new encoders plus LoRA adapters (rank 32) in the LLM are trained — the source of the large pretraining-cost saving.
Training has two stages: (1) continuous action training — freeze everything except a continuous action decoder and learn 7-DoF actions on Fractal + Bridge v2; (2) pixel-level understanding enhancement — apply LoRA to the LLM and jointly train the visual-prompting and pixel-aware encoders on Pixel-160K.
To build the data, the authors use a two-stage automated annotation pipeline: a gripper-aware region proposal stage (extract gripper-close frames, detect the gripper with SAM 2) and a multimodal object segmentation stage (extract target-object descriptions with an LLM, then Grounding DINO + SAM to segment and sample point/line/box visual prompts). This produces Pixel-160K — ~160K episodes / ~6.5M image–text–action triplets derived from Fractal and Bridge v2, with pixel-level mask and visual-prompt annotations. PixelVLA is presented as the first VLA to support both pixel-level understanding and multimodal (text + visual) prompting.
Across three standard VLA benchmarks — SimplerEnv-Google Robot, SimplerEnv-WidowX, and LIBERO — and two VLA backbones, manipulation success rate improves by 10.1%–28.7% over OpenVLA, while requiring only 1.5% of OpenVLA's pretraining cost. Representative numbers: SimplerEnv-Google Robot 61.4 visual-matching / 50.1 variant-aggregation vs. OpenVLA 32.7 / 40.0 (the +28.7 / +10.1 endpoints), and LIBERO average 86.7% vs. OpenVLA 76.5% and TraceVLA 78.1%. Dataset and code planned for open release.
Reframes "what a VLA sees" from patch tokens to pixel-level structure and "how a user instructs it" from text-only to multimodal prompts. Connects to the wider 2026 trend of grounding-aware VLAs (Embodied R1, OmniSAT) and is positioned as an efficient alternative — orders of magnitude less pretraining than OpenVLA-style scaling.
- OpenReview: https://openreview.net/forum?id=7M6ryCABIc
- arXiv: https://arxiv.org/abs/2511.01571
← Back to ICLR-2026