ICLR 2026 Cortical Policy - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA architecture ยท View transformer Trend tag: Dual-stream perception Authors: Xuening Zhang, Qi Lv, Xiang Deng, Miao Zhang, Xingbo Liu, Liqiang Nie (Harbin Institute of Technology, Shenzhen)
flowchart LR
Views[Multi-view obs] --> S[Static-view stream<br/>3D-foundation keypoint alignment]
Views --> D[Dynamic-view stream<br/>egocentric gaze pretraining]
S --> Fuse[Fuse + language conditioning]
D --> Fuse
Fuse --> Act[Action]
Existing view-transformer policies extract features in a view-specific way, leading to weak spatial reasoning and poor adaptation when the manipulation context shifts. The paper attributes this to the lack of separate mechanisms for stable geometric understanding versus dynamic, gaze-like attention.
Built on an RVT-2 backbone (preserving its two-stage processing, intra-view self-attention, and vision-language co-attention), the method adds a biologically-inspired dual-stream view transformer:
- Static-view stream โ aligns features of geometrically consistent keypoints extracted from the pretrained VGGT (Visual Geometry Grounded Transformer) 3D foundation model, supervised by a cross-view geometric-consistency loss.
- Dynamic-view stream โ pretrained on egocentric gaze estimation (position-aware pretraining producing gaze heatmaps) for position-aware adaptive attention, mirroring the cortical dorsal pathway.
The two streams are fused and conditioned on language to predict actions for spatially complex tasks.
- RLBench (18 tasks, 249 variations): 81.0% avg success vs RVT-2 77.5, SAM-E 70.6, RVT 62.9, PerAct 49.4. +3.5 pts over prior SOTA (RVT-2); top-1/top-2 on 14 of 18 tasks. (Baseline table: Hiveformer, PerAct, RVT, ฮฃ-agent, SAM-E, VIHE, RVT-2, 3D-MVP.)
- COLOSSEUM (14 perturbation factors): 69.9% avg vs RVT-2 60.5, PerAct 7.7; top performance on 9 of 14 conditions.
- Real-world (4 tasks, 10 trials each): on a stack-with-displacement task Cortical Policy reaches 80% vs 0% for static-view baselines, and ~30% higher success than RVT/RVT-2 on a spatial-reasoning task โ highlighting the dynamic stream's value under context shift.
- Cross-view geometric-consistency loss: +2.6%.
- Position-aware pretraining vs end-to-end variant: +1.9%.
- Dynamic-view gaze heatmaps are critical for the displacement/adaptation gains.
A neuroscience-motivated decomposition of perception in robot policies: separate the "what/where" geometric channel from the dynamic gaze-like channel. Sits alongside other 2026 work that questions monolithic visual encoders for manipulation and provides a concrete fusion recipe.
- OpenReview: https://openreview.net/forum?id=eWe8zqGvs5
- arXiv: https://arxiv.org/abs/2603.21051
โ Back to ICLR-2026