ICML 2026 Bring My Cup Personalizing Vision Language Action - Heungwoo/research GitHub Wiki
Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting — Training-free instance-level personalization for frozen VLAs
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Sangoh Lee, Sangwoo Mo, Wook-Shin Han Traction (2026-06): 3 citations (arXiv)

Problem
VLA models generalize well to generic instructions like "pick up the cup," but they fail on personalized commands such as "bring my cup," where the robot must act on one specific instance among visually similar distractors. Because user-specific instances and personalized phrasings ("my cup") are rarely covered during training, standard VLAs collapse to category-level recognition and mis-ground the intended target. Language inherently abstracts away the fine-grained cues — a unique pattern, a chipped rim — that distinguish a possession from its look-alikes, so even detailed text descriptions remain brittle. Per-object fine-tuning is impractical at deployment, motivating a training-free approach that adapts a frozen VLA to a new personal object from only a few reference images (e.g., 4–5 views).
Method
Visual Attentive Prompting (VAP) is a training-free input-side perceptual adapter that equips a frozen VLA with top-down selective attention. It leaves all VLA parameters untouched and only preprocesses the visual and language inputs — modifying what the model sees and how the instruction is phrased so a generic policy behaves as if it understood personalized commands.

VAP reformulates personalization as a two-stage inference process:
- Grounding (
g): The reference set is treated as a non-parametric visual memory. At each timestep the system compares this memory to the live images via open-vocabulary detection (Grounding DINO for category proposals) plus embedding-based matching, producing a maskM_tof the personal object in each camera view. Masks are propagated across frames with SAM2 tracking using historyH_{t-1}. - Visual Prompting (
p): The grounded mask steers the VLA by overlaying a semi-transparent highlight on the masked region (leaving the rest of the scene unchanged) and rewriting the instruction into a generic referring expression matching the cue (e.g., "pick up my cup" → "pick up the red cup"). The frozen VLA then acts on the prompted observation as if it were a standard generic task.
Results
Two new simulation benchmarks (Personalized-SIMPLER, Personalized-VLABench) and a real-world SO-101 tabletop benchmark (8 object categories) evaluate Success Rate (SR) and correct-object interaction (CMR/Correct Movement Ratio).
- Personalized-SIMPLER (Google Robot, visual matching): VAP lifts SR from 8.5% → 60.3% and CMR from 10.5% → 89.2%. WidowX Task 3 reaches CMR/SR of 82.9 / 71.3; on placement Task 5, VAP hits 100% CMR / 95% SR versus a soft-prompt baseline at 79.2% / 62.5%.
- Real-world (π0.5 policy): Across four pointing tasks, SR rises from 30% → 80% (text prompts only reach 40–45%). Across four pick-and-place tasks, SR rises from 18.8% → 56.2% (hard/soft prompts stay in the 26–30% range).
- Efficiency: One-time initialization costs 0.26 s (Grounding DINO 0.19 s + crop/embed/segment 0.07 s); per-step SAM2 tracking adds 0.02 s atop 0.20 s VLA inference — roughly 10% per-step overhead.
Error analysis (Table 6) attributes most simulation failures to correct-prompt-but-failed-manipulation (Case 3), while real-world and VLABench failures show more multi-view mask inconsistency (Case 2).
Significance
VAP shows that instance-level personalization — a key requirement for home robots — can be achieved without any policy fine-tuning, simply by reshaping inputs to a frozen VLA. By bridging semantic understanding and instance-level control through a non-parametric visual memory plus mask-based visual prompting, it offers a practical building block for personalized assistance across multiple robots and policies (π0, π0.5), at modest runtime cost. The authors note dependence on initial segmentation quality and a current focus on single-object, short-horizon tasks as limitations to extend.
Links
- arXiv: 2512.20014
- ICML 2026: https://icml.cc/virtual/2026/poster/62528
← Back to ICML-2026