ICLR 2026 PA3FF - Heungwoo/research GitHub Wiki

PA3FF — Part-Aware Dense 3D Feature Field

Venue: ICLR 2026 Category: Representations — 3D feature fields Trend tag: Spatial / 3D for VLA Affiliations: Peking University, Beijing Institute of Technology, NUS

Approach diagram

flowchart LR
  subgraph Stage1[Stage 1: 3D Geometric Prior]
    Big[140k point clouds] --> Sonata[Sonata / PTv3<br/>self-distillation pretraining]
  end
  subgraph Stage2[Stage 2: Part-Aware Refinement]
    Parts[PartNet-Mobility +<br/>3DCoMPaT + PartObjaverse-Tiny] --> Cont[Contrastive learning<br/>L_Geo + L_Sem]
    Sonata --> Refine[PTv3 with most downsampling<br/>removed + extra transformer<br/>blocks + per-point MLP]
    SigLIP[SigLIP text encoder<br/>part names] --> Cont
    Cont --> Refine
    Refine --> PA3FF[PA3FF features<br/>3D, dense, part-aware, semantic]
  end
  subgraph Stage3[Stage 3: PADP]
    PC[Point cloud] --> PA3FF
    PA3FF --> Enc[Frozen feature backbone]
    Lang[Language instruction] --> SigLIP2[SigLIP] --> CLS[CLS = part-name embedding]
    CLS --> XEnc[Trainable Transformer encoder]
    Enc --> XEnc
    XEnc --> MLP2[2-layer MLP] --> DiffH[Diffusion action head<br/>DDPM train, DDIM infer]
    State[Robot state q_t] --> XEnc
    DiffH --> Act["Action chunk a_t..a_{t+H-1}"]
  end
Loading

Problem

Articulated-object manipulation (drawers, microwaves, lids, dispensers) requires localizing functional parts — handles, knobs, hinges — across novel object instances and categories. The dominant pre-trained feature options each fail in a specific way:

  • 2D foundation features (CLIP, DINOv2, SigLIP) lifted to 3D via multi-view fusion or neural rendering: long inference (often minutes), inconsistent features across views, low spatial resolution (ViT patches give 14Ɨ downsampled feature maps in DINOv2), and they miss thin/small functional parts that fall below patch resolution.
  • Affordance/grasp-primitive methods: limited to motion primitives, not full trajectory control.
  • 2D + 3D hybrids (GenDP): cosine-similarity dense fields lifted from DINOv2, but inherit DINOv2's view inconsistency and lack functional-part granularity.
  • Native 3D backbones (Sonata, NDF): 3D-dense but not part-aware — they see geometry but not "this is a handle, that is a body."

PA3FF's bet: a natively 3D, dense, semantic, and part-aware feature representation built from 3D part-proposal data — feature distance reflects functional-part proximity, not just geometric similarity.

Detailed Method

Backbone (Sonata + PTv3 modifications)

  • Base: Sonata (Wu et al., 2025b), a self-supervised pretrained Point Transformer V3, originally trained on 140k scene-level point clouds via self-distillation.
  • Architectural surgery for object-level use: PTv3 was designed for large scenes with aggressive downsampling. PA3FF removes most downsampling layers and stacks additional transformer blocks to preserve detail at the smaller object scale. Ablation Table 6 confirms this is worth ~+4 to +7 absolute SR.

Contrastive learning losses

Given features {f_k} and part labels {a_k} for N points, two complementary losses:

Geometric loss (Supervised Contrastive, intra/inter-part):

L_Geo = Ī£_i (-1/(N_{a_i}-1)) Ī£_{j} 1_{i≠j} 1_{a_i=a_j} log(exp(f_i Ā· f_j / Ļ„) / Ī£_k 1_{i≠k} exp(f_i Ā· f_k / Ļ„))

Pulls together points of the same part; pushes apart points of different parts.

Semantic loss (InfoNCE against SigLIP text embeddings of part names):

x_k = SigLIP(s_k)   for s_k ∈ {part name 1, ..., part name m}
L_Sem = Σ_i -log(exp(f_i · x_{a_i} / τ) / Σ_k exp(f_i · x_k / τ))

Aligns 3D point features with text embeddings of part names so PA3FF features are also retrievable by language.

Total: L_total = L_Geo + L_Sem. A lightweight per-point MLP refines the Sonata output before this loss.

Training data for the contrastive stage: PartNet-Mobility (Xiang et al., 2020; Mo et al., 2018), 3DCoMPaT (Slim et al., 2025), PartObjaverse-Tiny (Yang et al., 2024b).

PADP (Part-Aware Diffusion Policy)

  • Inputs: point clouds {P_1^t, ..., P_n^t} from n cameras + proprioceptive state q_t. Action chunk A_t = [a_t, ..., a_{t+H-1}].
  • PA3FF as frozen backbone: extracts per-point features.
  • Trainable Transformer encoder: aggregates per-point features into a global representation. The CLS token is initialized as the SigLIP embedding of the task-critical part name — this routes the transformer's attention to the relevant functional part by construction.
  • 2-layer MLP + diffusion head: outputs a chunk of robot actions.
  • Training: DDPM (denoising-diffusion-probabilistic-model) MSE objective L(Ļ•) = MSE(a_t, D_Īø(o_t, Ć£_t, k)) with sampled noise Ć£_t = āˆšĪ²Ģ„^k a_t + √(1-β̄^k) ε.
  • Inference: DDIM sampling for speed.

Comprehensive Results

Simulation: PartInstruct (16 task classes, 5-level generalization protocol)

Method Test 1 (OS) novel pose Test 2 (OI) novel instances Test 3 (TP) novel part combos Test 4 (TC) novel task class Test 5 (OC) novel obj. category Average
Act3D 6.25 ±1.8 5.68 ±1.7 4.55 ±1.6 0.0 2.08 ±2.1 3.88 ±1.8
RVT2 4.55 ±2.0 4.55 ±2.0 6.36 ±2.3 0.91 ±0.9 3.33 ±3.3 4.04 ±2.1
3D-DA 8.08 ±2.7 5.05 ±2.2 4.04 ±1.9 0.0 3.70 ±3.6 4.26 ±1.0
DP 7.27 ±1.8 8.64 ±1.9 8.18 ±1.8 3.75 ±2.1 6.67 ±3.2 5.96 ±2.2
DP3 23.18 ±2.8 23.18 ±2.8 18.18 ±2.6 7.73 ±1.8 6.67 ±3.2 15.40 ±2.6
GenDP 24.34 ±2.1 23.36 ±2.3 24.53 ±1.9 10.00 ±2.0 14.61 ±2.1 19.36 ±2.7
PADP 36.76 ±2.3 34.33 ±3.6 32.45 ±1.6 13.75 ±2.0 26.67 ±3.2 28.79 ±2.5

PADP wins every column, with the biggest gap on novel object categories (Test 5: +12.06 over GenDP) — the test most directly probing part-aware generalization.

Real world: 8 articulated-manipulation tasks (Franka + UMI fingers + 3Ɨ Intel RealSense D415, 30 demos per task)

Task DP (train/test) DP3 GenDP PADP
Pulling lid of pot 4/2 6/4 7/6 8/6
Open drawer 1/1 4/3 5/5 6/6
Close box 2/1 3/3 3/3 5/5
Close lid of laptop 4/3 5/5 6/4 7/7
Open microwave 1/0 3/1 4/3 6/5
Open bottle 5/2 4/3 5/4 8/6
Put lid on kettle 1/0 3/1 4/2 6/5
Press dispenser 0/0 1/1 2/1 4/3
Mean test (unseen) 1.1/10 2.6/10 3.5/10 5.4/10

PADP achieves the highest mean unseen-object SR vs 35% for the best baseline (GenDP). Note a paper-internal inconsistency: the text claims a 58.75% PADP mean (→ +23.75 over 35%), but the per-task Table 2 test column sums to 43/80 = 53.75%, matching the separately stated +18.75% increment over the best baseline. The 53.75% figure is the one consistent with the per-task table.

Generalization (Open Bottle task, 10 trials)

Method Original Spatial Disturbance Object Disturbance Environment Disturbance
DP 50 40 (āˆ’10) 10 (āˆ’40) 0 (āˆ’50)
DP3 40 20 (āˆ’20) 20 (āˆ’20) 10 (āˆ’30)
GenDP 50 20 (āˆ’30) 30 (āˆ’20) 30 (āˆ’20)
PADP 80 70 (āˆ’10) 60 (āˆ’20) 60 (āˆ’20)

DP collapses with background changes (image-only encoder); DP3 resists color but fails with distractors; GenDP best baseline but still drops to 30 in combined disturbance; PADP holds 60+ across all axes — direct evidence that part-aware features are robust to environmental confounds.

Downstream applications (no policy training)

  • 3D shape correspondences via PA3FF + Functional Maps (Ovsjanikov et al., 2012) + Smooth Discrete Optimization (Magnet et al., 2022) outperforms DINOv2 in topologically dissimilar pairs (qualitative; see Fig. 6).

  • 3D part segmentation on PartNetE (Table 5, mAP50 %):

    Method Bottle Chair Display Lamp Storage Table Avg
    PointGroup 8.0 77.2 16.7 9.8 0.0 0.0 18.6
    SoftGroup 22.4 87.7 27.5 19.4 11.6 14.2 30.5
    PartSlip 79.4 84.4 82.9 68.3 32.8 32.3 63.4
    PartSlip++ 78.5 86.0 74.1 66.9 36.7 33.5 62.6
    PA3FF + agglomerative clustering 94.6 90.0 86.5 69.5 49.6 33.4 70.6

    Despite never training on segmentation, PA3FF + clustering tops PartSlip by +7.2 mAP50 average.

Ablation Studies (Table 6, "Put in Drawer", "Turn Tap", "Open Box")

Method Put in Drawer (%) Turn Tap Open Box
PADP (Full) 62 69 66
– w/o Stacked Transformers 58 63 59
– w/o Geometric Loss 54 57 55
– w/o Semantic Loss 46 53 52
Sonata + DP3 39 50 44
DP3 (baseline) 37 47 39

Reading: Sonata + DP3 alone is only +2 over DP3 (39 vs 37) — direct combination doesn't work. Adding the stacked-transformer surgery (but no contrastive refinement) reaches 46%, +7 over Sonata+DP3. Of the two contrastive losses, removing the semantic loss hurts most (62 → 46, āˆ’16) versus removing the geometric loss (62 → 54, āˆ’8), confirming that grounding 3D features in part-name text is what gives PA3FF its part awareness. (The paper's prose frames the full 62 → 46 drop as removing "feature refinement," i.e. the whole contrastive stage, rather than the semantic loss alone.)

Comparison with other foundation features (Table 4)

Feature Part-Aware 3D Native Dense Semantic
DINOv2 āœ— āœ— āœ— āœ—
SigLIP / CLIP āœ— āœ— āœ— āœ“
ULIP āœ— āœ“ āœ— āœ“
NDF āœ— āœ“ āœ“ āœ—
D3Field āœ— āœ“ āœ— āœ“
LERF āœ— āœ“ āœ— āœ“
Sonata āœ— āœ“ āœ“ āœ—
PA3FF āœ“ āœ“ āœ“ āœ“

PA3FF claims the unique intersection of all four properties.

Limitations (as stated)

The paper does not have a dedicated Limitations section. Implicit and partially stated limitations:

  • Coverage of part labels: PA3FF requires part-annotation training data (PartNet-Mobility, 3DCoMPaT, PartObjaverse-Tiny). Generalization to entirely unseen part categories not in these corpora is uncertain.
  • Small-demo regime: real-world experiments use only 30 demos per task. The paper reports good sample efficiency but no scaling curves.
  • Backbone frozen: PA3FF is used as a frozen backbone in PADP — the policy cannot fine-tune perception for task-specific features.
  • Sensor dependence: all real experiments use 3Ɨ RealSense D415 RGB-D cameras; degradation under fewer or less precise depth sensors is not studied.
  • Inference cost vs the 2D baselines: PA3FF avoids minute-scale lift-up costs but adds PTv3 backbone forward time; not benchmarked against 2D-feature DP runtime explicitly.
  • The Test 4 (TC, novel task category) score remains low (13.75%) — generalizing to entirely new task structures is still hard even with part-aware features.

Significance & Positioning

PA3FF takes a different stance from the "VLA + 3D" cluster (FALCON, Spatial Forcing, SP-VLA): instead of patching 3D into a 2D-pretrained VLA, PA3FF builds a representation that is natively 3D and part-aware by construction and pairs it with a non-VLA diffusion policy. The bet is that for articulated-object manipulation specifically — where the bottleneck is identifying a specific functional part across object variations — a focused representation outperforms a generalist VLA.

Versus the closest method, GenDP (which lifts DINOv2 features to 3D via cosine-similarity fields), PA3FF's contrastive part-name supervision is the qualitative difference: GenDP can match a handle to a similar-looking handle but cannot answer "this point is on a handle" or distinguish a knob from a button. The +9.4 absolute average win on PartInstruct and the āˆ’16 ablation drop without the semantic loss are the empirical evidence.

Versus VLA-based articulated manipulation (Ļ€0, OpenVLA, etc.): the paper does not directly compare to large VLAs on PartInstruct. The methodological contribution is orthogonal — a foundation feature that any policy (DP, VLA action head, point-cloud transformer) can plug in. The downstream-applications results (segmentation, correspondence) on PartNetE without further training are the strongest argument that PA3FF is a true representation rather than a task-specific encoder.

The bigger-picture story: as VLAs accumulate spatial competence (FALCON, SP-VLA, Spatial Forcing), having a high-quality 3D-native foundation feature underneath them — rather than relying on lifted 2D features — looks increasingly attractive. PA3FF is the first ICLR 2026 paper to explicitly target functional-part awareness as the missing piece.

Links

Related pages

← Back to ICLR-2026

āš ļø **GitHub.com Fallback** āš ļø