ICLR 2026 PA3FF - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Representations ā 3D feature fields Trend tag: Spatial / 3D for VLA Affiliations: Peking University, Beijing Institute of Technology, NUS
flowchart LR
subgraph Stage1[Stage 1: 3D Geometric Prior]
Big[140k point clouds] --> Sonata[Sonata / PTv3<br/>self-distillation pretraining]
end
subgraph Stage2[Stage 2: Part-Aware Refinement]
Parts[PartNet-Mobility +<br/>3DCoMPaT + PartObjaverse-Tiny] --> Cont[Contrastive learning<br/>L_Geo + L_Sem]
Sonata --> Refine[PTv3 with most downsampling<br/>removed + extra transformer<br/>blocks + per-point MLP]
SigLIP[SigLIP text encoder<br/>part names] --> Cont
Cont --> Refine
Refine --> PA3FF[PA3FF features<br/>3D, dense, part-aware, semantic]
end
subgraph Stage3[Stage 3: PADP]
PC[Point cloud] --> PA3FF
PA3FF --> Enc[Frozen feature backbone]
Lang[Language instruction] --> SigLIP2[SigLIP] --> CLS[CLS = part-name embedding]
CLS --> XEnc[Trainable Transformer encoder]
Enc --> XEnc
XEnc --> MLP2[2-layer MLP] --> DiffH[Diffusion action head<br/>DDPM train, DDIM infer]
State[Robot state q_t] --> XEnc
DiffH --> Act["Action chunk a_t..a_{t+H-1}"]
end
Articulated-object manipulation (drawers, microwaves, lids, dispensers) requires localizing functional parts ā handles, knobs, hinges ā across novel object instances and categories. The dominant pre-trained feature options each fail in a specific way:
- 2D foundation features (CLIP, DINOv2, SigLIP) lifted to 3D via multi-view fusion or neural rendering: long inference (often minutes), inconsistent features across views, low spatial resolution (ViT patches give 14Ć downsampled feature maps in DINOv2), and they miss thin/small functional parts that fall below patch resolution.
- Affordance/grasp-primitive methods: limited to motion primitives, not full trajectory control.
- 2D + 3D hybrids (GenDP): cosine-similarity dense fields lifted from DINOv2, but inherit DINOv2's view inconsistency and lack functional-part granularity.
- Native 3D backbones (Sonata, NDF): 3D-dense but not part-aware ā they see geometry but not "this is a handle, that is a body."
PA3FF's bet: a natively 3D, dense, semantic, and part-aware feature representation built from 3D part-proposal data ā feature distance reflects functional-part proximity, not just geometric similarity.
- Base: Sonata (Wu et al., 2025b), a self-supervised pretrained Point Transformer V3, originally trained on 140k scene-level point clouds via self-distillation.
- Architectural surgery for object-level use: PTv3 was designed for large scenes with aggressive downsampling. PA3FF removes most downsampling layers and stacks additional transformer blocks to preserve detail at the smaller object scale. Ablation Table 6 confirms this is worth ~+4 to +7 absolute SR.
Given features {f_k} and part labels {a_k} for N points, two complementary losses:
Geometric loss (Supervised Contrastive, intra/inter-part):
L_Geo = Ī£_i (-1/(N_{a_i}-1)) Ī£_{j} 1_{iā j} 1_{a_i=a_j} log(exp(f_i Ā· f_j / Ļ) / Ī£_k 1_{iā k} exp(f_i Ā· f_k / Ļ))
Pulls together points of the same part; pushes apart points of different parts.
Semantic loss (InfoNCE against SigLIP text embeddings of part names):
x_k = SigLIP(s_k) for s_k ā {part name 1, ..., part name m}
L_Sem = Ī£_i -log(exp(f_i Ā· x_{a_i} / Ļ) / Ī£_k exp(f_i Ā· x_k / Ļ))
Aligns 3D point features with text embeddings of part names so PA3FF features are also retrievable by language.
Total: L_total = L_Geo + L_Sem. A lightweight per-point MLP refines the Sonata output before this loss.
Training data for the contrastive stage: PartNet-Mobility (Xiang et al., 2020; Mo et al., 2018), 3DCoMPaT (Slim et al., 2025), PartObjaverse-Tiny (Yang et al., 2024b).
-
Inputs: point clouds
{P_1^t, ..., P_n^t}from n cameras + proprioceptive state q_t. Action chunkA_t = [a_t, ..., a_{t+H-1}]. - PA3FF as frozen backbone: extracts per-point features.
- Trainable Transformer encoder: aggregates per-point features into a global representation. The CLS token is initialized as the SigLIP embedding of the task-critical part name ā this routes the transformer's attention to the relevant functional part by construction.
- 2-layer MLP + diffusion head: outputs a chunk of robot actions.
-
Training: DDPM (denoising-diffusion-probabilistic-model) MSE objective
L(Ļ) = MSE(a_t, D_Īø(o_t, Ć£_t, k))with sampled noiseĆ£_t = āβĢ^k a_t + ā(1-βĢ^k) ε. - Inference: DDIM sampling for speed.
| Method | Test 1 (OS) novel pose | Test 2 (OI) novel instances | Test 3 (TP) novel part combos | Test 4 (TC) novel task class | Test 5 (OC) novel obj. category | Average |
|---|---|---|---|---|---|---|
| Act3D | 6.25 ±1.8 | 5.68 ±1.7 | 4.55 ±1.6 | 0.0 | 2.08 ±2.1 | 3.88 ±1.8 |
| RVT2 | 4.55 ±2.0 | 4.55 ±2.0 | 6.36 ±2.3 | 0.91 ±0.9 | 3.33 ±3.3 | 4.04 ±2.1 |
| 3D-DA | 8.08 ±2.7 | 5.05 ±2.2 | 4.04 ±1.9 | 0.0 | 3.70 ±3.6 | 4.26 ±1.0 |
| DP | 7.27 ±1.8 | 8.64 ±1.9 | 8.18 ±1.8 | 3.75 ±2.1 | 6.67 ±3.2 | 5.96 ±2.2 |
| DP3 | 23.18 ±2.8 | 23.18 ±2.8 | 18.18 ±2.6 | 7.73 ±1.8 | 6.67 ±3.2 | 15.40 ±2.6 |
| GenDP | 24.34 ±2.1 | 23.36 ±2.3 | 24.53 ±1.9 | 10.00 ±2.0 | 14.61 ±2.1 | 19.36 ±2.7 |
| PADP | 36.76 ±2.3 | 34.33 ±3.6 | 32.45 ±1.6 | 13.75 ±2.0 | 26.67 ±3.2 | 28.79 ±2.5 |
PADP wins every column, with the biggest gap on novel object categories (Test 5: +12.06 over GenDP) ā the test most directly probing part-aware generalization.
Real world: 8 articulated-manipulation tasks (Franka + UMI fingers + 3Ć Intel RealSense D415, 30 demos per task)
| Task | DP (train/test) | DP3 | GenDP | PADP |
|---|---|---|---|---|
| Pulling lid of pot | 4/2 | 6/4 | 7/6 | 8/6 |
| Open drawer | 1/1 | 4/3 | 5/5 | 6/6 |
| Close box | 2/1 | 3/3 | 3/3 | 5/5 |
| Close lid of laptop | 4/3 | 5/5 | 6/4 | 7/7 |
| Open microwave | 1/0 | 3/1 | 4/3 | 6/5 |
| Open bottle | 5/2 | 4/3 | 5/4 | 8/6 |
| Put lid on kettle | 1/0 | 3/1 | 4/2 | 6/5 |
| Press dispenser | 0/0 | 1/1 | 2/1 | 4/3 |
| Mean test (unseen) | 1.1/10 | 2.6/10 | 3.5/10 | 5.4/10 |
PADP achieves the highest mean unseen-object SR vs 35% for the best baseline (GenDP). Note a paper-internal inconsistency: the text claims a 58.75% PADP mean (ā +23.75 over 35%), but the per-task Table 2 test column sums to 43/80 = 53.75%, matching the separately stated +18.75% increment over the best baseline. The 53.75% figure is the one consistent with the per-task table.
| Method | Original | Spatial Disturbance | Object Disturbance | Environment Disturbance |
|---|---|---|---|---|
| DP | 50 | 40 (ā10) | 10 (ā40) | 0 (ā50) |
| DP3 | 40 | 20 (ā20) | 20 (ā20) | 10 (ā30) |
| GenDP | 50 | 20 (ā30) | 30 (ā20) | 30 (ā20) |
| PADP | 80 | 70 (ā10) | 60 (ā20) | 60 (ā20) |
DP collapses with background changes (image-only encoder); DP3 resists color but fails with distractors; GenDP best baseline but still drops to 30 in combined disturbance; PADP holds 60+ across all axes ā direct evidence that part-aware features are robust to environmental confounds.
-
3D shape correspondences via PA3FF + Functional Maps (Ovsjanikov et al., 2012) + Smooth Discrete Optimization (Magnet et al., 2022) outperforms DINOv2 in topologically dissimilar pairs (qualitative; see Fig. 6).
-
3D part segmentation on PartNetE (Table 5, mAP50 %):
Method Bottle Chair Display Lamp Storage Table Avg PointGroup 8.0 77.2 16.7 9.8 0.0 0.0 18.6 SoftGroup 22.4 87.7 27.5 19.4 11.6 14.2 30.5 PartSlip 79.4 84.4 82.9 68.3 32.8 32.3 63.4 PartSlip++ 78.5 86.0 74.1 66.9 36.7 33.5 62.6 PA3FF + agglomerative clustering 94.6 90.0 86.5 69.5 49.6 33.4 70.6 Despite never training on segmentation, PA3FF + clustering tops PartSlip by +7.2 mAP50 average.
| Method | Put in Drawer (%) | Turn Tap | Open Box |
|---|---|---|---|
| PADP (Full) | 62 | 69 | 66 |
| ā w/o Stacked Transformers | 58 | 63 | 59 |
| ā w/o Geometric Loss | 54 | 57 | 55 |
| ā w/o Semantic Loss | 46 | 53 | 52 |
| Sonata + DP3 | 39 | 50 | 44 |
| DP3 (baseline) | 37 | 47 | 39 |
Reading: Sonata + DP3 alone is only +2 over DP3 (39 vs 37) ā direct combination doesn't work. Adding the stacked-transformer surgery (but no contrastive refinement) reaches 46%, +7 over Sonata+DP3. Of the two contrastive losses, removing the semantic loss hurts most (62 ā 46, ā16) versus removing the geometric loss (62 ā 54, ā8), confirming that grounding 3D features in part-name text is what gives PA3FF its part awareness. (The paper's prose frames the full 62 ā 46 drop as removing "feature refinement," i.e. the whole contrastive stage, rather than the semantic loss alone.)
| Feature | Part-Aware | 3D Native | Dense | Semantic |
|---|---|---|---|---|
| DINOv2 | ā | ā | ā | ā |
| SigLIP / CLIP | ā | ā | ā | ā |
| ULIP | ā | ā | ā | ā |
| NDF | ā | ā | ā | ā |
| D3Field | ā | ā | ā | ā |
| LERF | ā | ā | ā | ā |
| Sonata | ā | ā | ā | ā |
| PA3FF | ā | ā | ā | ā |
PA3FF claims the unique intersection of all four properties.
The paper does not have a dedicated Limitations section. Implicit and partially stated limitations:
- Coverage of part labels: PA3FF requires part-annotation training data (PartNet-Mobility, 3DCoMPaT, PartObjaverse-Tiny). Generalization to entirely unseen part categories not in these corpora is uncertain.
- Small-demo regime: real-world experiments use only 30 demos per task. The paper reports good sample efficiency but no scaling curves.
- Backbone frozen: PA3FF is used as a frozen backbone in PADP ā the policy cannot fine-tune perception for task-specific features.
- Sensor dependence: all real experiments use 3Ć RealSense D415 RGB-D cameras; degradation under fewer or less precise depth sensors is not studied.
- Inference cost vs the 2D baselines: PA3FF avoids minute-scale lift-up costs but adds PTv3 backbone forward time; not benchmarked against 2D-feature DP runtime explicitly.
- The Test 4 (TC, novel task category) score remains low (13.75%) ā generalizing to entirely new task structures is still hard even with part-aware features.
PA3FF takes a different stance from the "VLA + 3D" cluster (FALCON, Spatial Forcing, SP-VLA): instead of patching 3D into a 2D-pretrained VLA, PA3FF builds a representation that is natively 3D and part-aware by construction and pairs it with a non-VLA diffusion policy. The bet is that for articulated-object manipulation specifically ā where the bottleneck is identifying a specific functional part across object variations ā a focused representation outperforms a generalist VLA.
Versus the closest method, GenDP (which lifts DINOv2 features to 3D via cosine-similarity fields), PA3FF's contrastive part-name supervision is the qualitative difference: GenDP can match a handle to a similar-looking handle but cannot answer "this point is on a handle" or distinguish a knob from a button. The +9.4 absolute average win on PartInstruct and the ā16 ablation drop without the semantic loss are the empirical evidence.
Versus VLA-based articulated manipulation (Ļ0, OpenVLA, etc.): the paper does not directly compare to large VLAs on PartInstruct. The methodological contribution is orthogonal ā a foundation feature that any policy (DP, VLA action head, point-cloud transformer) can plug in. The downstream-applications results (segmentation, correspondence) on PartNetE without further training are the strongest argument that PA3FF is a true representation rather than a task-specific encoder.
The bigger-picture story: as VLAs accumulate spatial competence (FALCON, SP-VLA, Spatial Forcing), having a high-quality 3D-native foundation feature underneath them ā rather than relying on lifted 2D features ā looks increasingly attractive. PA3FF is the first ICLR 2026 paper to explicitly target functional-part awareness as the missing piece.
- OpenReview: https://openreview.net/forum?id=qXfRXfAHOK
- PDF: https://openreview.net/pdf?id=qXfRXfAHOK
- Project page: https://pa3ff.github.io/
ā Back to ICLR-2026