ICML 2026 Bring My Cup Personalizing Vision Language Action - Heungwoo/research GitHub Wiki

Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting — Training-free instance-level personalization for frozen VLAs

Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Sangoh Lee, Sangwoo Mo, Wook-Shin Han Traction (2026-06): 3 citations (arXiv)

Manipulating personal objects with VLA: existing models collapse "my cup" to a generic category, while VAP grounds the user-specific instance from a small visual memory and steers the frozen VLA through visual prompting (Figure 1 from Lee et al., 2026)

Problem

VLA models generalize well to generic instructions like "pick up the cup," but they fail on personalized commands such as "bring my cup," where the robot must act on one specific instance among visually similar distractors. Because user-specific instances and personalized phrasings ("my cup") are rarely covered during training, standard VLAs collapse to category-level recognition and mis-ground the intended target. Language inherently abstracts away the fine-grained cues — a unique pattern, a chipped rim — that distinguish a possession from its look-alikes, so even detailed text descriptions remain brittle. Per-object fine-tuning is impractical at deployment, motivating a training-free approach that adapts a frozen VLA to a new personal object from only a few reference images (e.g., 4–5 views).

Method

Visual Attentive Prompting (VAP) is a training-free input-side perceptual adapter that equips a frozen VLA with top-down selective attention. It leaves all VLA parameters untouched and only preprocesses the visual and language inputs — modifying what the model sees and how the instruction is phrased so a generic policy behaves as if it understood personalized commands.

VAP builds a non-parametric visual memory from a few reference images, grounds the target with frozen open-vocabulary detection and segmentation, then highlights the object and rewrites the instruction for the frozen VLA, tracking it over subsequent frames (Figure 3 from Lee et al., 2026)

VAP reformulates personalization as a two-stage inference process:

  1. Grounding (g): The reference set is treated as a non-parametric visual memory. At each timestep the system compares this memory to the live images via open-vocabulary detection (Grounding DINO for category proposals) plus embedding-based matching, producing a mask M_t of the personal object in each camera view. Masks are propagated across frames with SAM2 tracking using history H_{t-1}.
  2. Visual Prompting (p): The grounded mask steers the VLA by overlaying a semi-transparent highlight on the masked region (leaving the rest of the scene unchanged) and rewriting the instruction into a generic referring expression matching the cue (e.g., "pick up my cup" → "pick up the red cup"). The frozen VLA then acts on the prompted observation as if it were a standard generic task.

Results

Two new simulation benchmarks (Personalized-SIMPLER, Personalized-VLABench) and a real-world SO-101 tabletop benchmark (8 object categories) evaluate Success Rate (SR) and correct-object interaction (CMR/Correct Movement Ratio).

  • Personalized-SIMPLER (Google Robot, visual matching): VAP lifts SR from 8.5% → 60.3% and CMR from 10.5% → 89.2%. WidowX Task 3 reaches CMR/SR of 82.9 / 71.3; on placement Task 5, VAP hits 100% CMR / 95% SR versus a soft-prompt baseline at 79.2% / 62.5%.
  • Real-world (π0.5 policy): Across four pointing tasks, SR rises from 30% → 80% (text prompts only reach 40–45%). Across four pick-and-place tasks, SR rises from 18.8% → 56.2% (hard/soft prompts stay in the 26–30% range).
  • Efficiency: One-time initialization costs 0.26 s (Grounding DINO 0.19 s + crop/embed/segment 0.07 s); per-step SAM2 tracking adds 0.02 s atop 0.20 s VLA inference — roughly 10% per-step overhead.

Error analysis (Table 6) attributes most simulation failures to correct-prompt-but-failed-manipulation (Case 3), while real-world and VLABench failures show more multi-view mask inconsistency (Case 2).

Significance

VAP shows that instance-level personalization — a key requirement for home robots — can be achieved without any policy fine-tuning, simply by reshaping inputs to a frozen VLA. By bridging semantic understanding and instance-level control through a non-parametric visual memory plus mask-based visual prompting, it offers a practical building block for personalized assistance across multiple robots and policies (π0, π0.5), at modest runtime cost. The authors note dependence on initial segmentation quality and a current focus on single-object, short-horizon tasks as limitations to extend.

Links

← Back to ICML-2026