ICML 2026 Neural Implicit Action Fields - Heungwoo/research GitHub Wiki

Neural Implicit Action Fields — from discrete waypoints to continuous-time action functions

Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Haoyun Liu, Jianzhuang Zhao, Xinyuan Chang, Tong Lin, Mu Xu, SongLin Dong, Zhiheng Ma, Yihong Gong, Sheng Zhong et al. (2026) Traction (2026-06): 1 citation (arXiv)

From discrete waypoints to continuous functions: resolution independence and analytical differentiability (Figure 1 from Liu et al., 2026)

Problem

Modern VLAs predict actions as discrete waypoints — fixed-horizon sequences of coordinates (ACT, Diffusion Policy) or compressed tokens such as B-spline control points (BEAST) and DCT coefficients (FAST). The authors argue this discretization is fundamentally misaligned with the continuous nature of physical motion, citing three limitations: (1) rigid time discretization — predictions are bound to the training frequency and cannot be queried at sub-step resolution without interpolation artifacts; (2) lack of high-order dynamics supervision — spline tokenizers quantize control points into codebooks and never constrain higher-order derivatives, yielding discontinuous velocity profiles and motion jitter; (3) dynamic inconsistency — discrete paradigms lack analytical differentiability, so position and velocity cannot be jointly supervised, and numerical differentiation amplifies quantization noise, making the feedforward terms required for impedance control unattainable. Agents are thus confined to stiff position control.

Method

Neural Implicit Action Fields (NIAF) reformulates action decoding from discrete tokenization to continuous function regression: the action chunk is modeled as a parameterized continuous-time function 𝒜(τ) = Φ(τ; θ). The MLLM acts as a hypernetwork / hierarchical spectral modulator: learnable query embeddings attend to the multimodal context and are decoded (one-step, parallel) into modulation latents Z, which are projected into FiLM-style coefficients (γ, β). These coefficients dynamically reconfigure the shared meta-parameters of a SIREN (Sinusoidal Representation Network) implicit decoder. Because γ scales frequency (via weights) and β shifts phase (via biases), the model decouples a stable shared kinematic backbone from instance-specific deformations.

The NIAF architecture: the MLLM hypernetwork modulates a shared SIREN via parallel decoding (Figure 2 from Liu et al., 2026)

A grouped hyper-modulation scheme constrains the latent length to Q = L×(G+1), aligning latent blocks with each SIREN layer (G weight/frequency tokens + 1 bias/phase token per layer). Using sinusoidal activations guarantees C∞ smoothness by construction. The implicit representation's analytical differentiability enables physics-informed supervision that explicitly constrains velocity, acceleration, and jerk in addition to position.

Results

On CALVIN, NIAF (0.77B, no robot-data pretraining) reaches average chain lengths of 4.66 on ABCD→D and 4.47 on ABC→D, beating BEAST (4.61 / 4.42), FLOWER (4.62 / 4.44), and larger pretrained models like UniVLA (9B). On LIBERO it attains a 97.9% average success rate (98.2 Spatial, 100.0 Object, 98.0 Goal, 95.5 Long), surpassing π₀ (94.2), OpenVLA-OFT (95.5), FLOWER (95.7), and BEAST (92.5). Under an identical Florence-2 Large backbone, the SIREN representation reaches 88.6 on CALVIN ABC→D and 95.5 on LIBERO-Long, ahead of BEAST-CT/BEAST-F/FAST/OFT. Ablations on CALVIN ABC→D show peaks at chunk size H=10 (4.47), monotonic gains with more weight groups (G=64 → 4.47), and sine activations beating ReLU (4.47 vs 3.91). Real-world experiments on AgileX Piper (single-arm) and Cobot Magic (bimanual) demonstrate stable impedance control and reduced jitter on delicate dynamic tasks (cup stacking, shape insertion, towel folding).

Significance

NIAF reframes the action head as a function, not a sequence — buying resolution independence (query at any control frequency without interpolation) and, crucially, analytically exact velocity/acceleration/jerk for feedforward impedance control. It delivers SOTA on CALVIN and LIBERO with a sub-1B backbone and no large-scale robot pretraining, scaling across backbones from Florence-2 to Qwen3-VL.

Links

← Back to ICML-2026