Anonymous supplementary material

Beyond Point-Attached
Semantics

Stable Object-Centric Semantic Fields for Robust Manipulation

Watch the research overview 2:44
Explore the research
Grasp Mug Held-out object · 3×

01 / The idea

Different observations.
Consistent part semantics.

A robot needs to know not just where an object is, but which part to use. A hammer is grasped by its handle and makes contact with its head, even when its shape or appearance changes.

Yet features attached to observed points can change with viewpoint, sensor noise, and sampling. We learn an object-centric semantic field that uses object context to provide more consistent part-aware features for a manipulation policy.

  1. 01Observe the object

    A partial point cloud

  2. 02Build a semantic field

    Geometry with part supervision

  3. 03Query part features

    Resampled 3D locations

  4. 04Condition the policy

    Semantic clouds + scene + state

The field changes how semantics are read, not where the robot observes. Query locations are resampled from the current object cloud; they are not fixed canonical points.
Read the abstract

Robotic manipulation often requires identifying functional parts, such as a mug handle or a hammer head. However, features attached to observed 3D points can vary with viewpoint and sensor noise, giving a policy inconsistent representations of the same part. We propose an object-centric semantic field to provide more consistent part-aware features for manipulation. We use the observed object cloud to build a continuous field, then read features from this field at 3D locations independently resampled from the cloud. Each feature uses the sampled object support as context, rather than directly reusing an individual point descriptor. Part classification distinguishes functional regions, cross-instance alignment brings corresponding part features together, and perturbation consistency encourages similar features under observation changes. The queried coordinates and features form semantic point clouds that are supplied to a DP3-based policy. We evaluate the approach on four RoboTwin simulation tasks and four real-world bimanual tasks, achieving average success rates of 69.3% and 67.5%, respectively. These improve on Utonia Point-wise by 7.0 and 32.5 percentage points, respectively, with real-world tests on held-out objects. A point-wise control with matched part supervision scores 63.5% in simulation, compared with our 69.3%. These results highlight the value of stable, object-conditioned semantic fields for manipulation across object instances and varying observations.

02 / Research overview

The method, in motion.

From the representation to the robot.

2 min 44 secEnglish narration · Subtitles included

03 / Method

Object context.
Part-aware field readout.

A frozen 3D encoder supplies geometry. Part supervision gives that geometry meaning. A continuous field brings both together at the locations used by the policy.

Two training stages: learn the field from part-annotated 3D models, then freeze it while learning the manipulation policy from demonstrations.
01 / Encode

Build object context

A frozen Utonia backbone extracts support features. A trainable adapter maps them into an object feature cache with three spatial feature planes and a pooled global feature.

02 / Learn

Distinguish and align parts

The decoder reads features at resampled query locations. PartNext labels supervise part classification; cross-instance alignment and perturbation consistency act on the embeddings.

03 / Act

Learn from demonstrations

Separate PointNets encode scene and semantic object clouds. Their features and the robot state condition a DP3 policy. The semantic field stays frozen during policy training.

A different object representation. The same action representation and diffusion learning objective.

04 / Real-world manipulation

One task policy.
Different held-out objects.

Each task policy learns from 50 demonstrations with designated training objects. At test time, it is applied to unseen physical instances without object-specific retraining.

Selected executions · 3× playback
01 Held-out mug 1
02 Held-out mug 2
03 Held-out mug 3
04 Held-out mug 4

Success rates cover all 20 trials per task, using randomly selected held-out objects, not just the executions shown here. Category-level fields are trained on PartNext models, with no scans or fine-tuning of the physical test objects.

Robot setup and object split
Two calibrated ZED 2i cameras provide RGB-D observations for the bimanual robot.
One training instance per category. Held-out sets contain five hammers, six teapots, six mugs, and six spoons. Stirring and pouring involve object pairs.

05 / Evaluation

Better manipulation.
Across four tasks in each setting.

69.3%

Simulation success

+7.0 points over Utonia Point-wise

67.5%

Real-world success

+32.5 points on held-out objects

+5.8pp

Beyond matched supervision

Over the part-supervised point-wise control

Success on unseen physical objects

OursUtonia Point-wise

Grasp Mug

85%
35%

Beat Cube

85%
45%

Stir Mug

50%
35%

Pour Water

50%
25%

20 trials per task. Rates are percentages of successful trials; no real-world standard deviations are reported.

Compare all real-world baselines
Real-world success rates (%) on held-out objects
MethodGrasp MugBeat CubeStir MugPour WaterAverage
DP335.040.015.020.027.5
DINOv2 Lifting35.045.025.025.032.5
Utonia Point-wise35.045.035.025.035.0
Ours85.085.050.050.067.5

A controlled comparison

More than part labels alone.

The part-supervised point-wise control uses the same encoder, training data, and supervision, but reads features attached to observed points. Our field reaches 69.3%, versus 63.5% for this control.

This comparison tests the field-based representation against direct point-wise readout. It does not isolate query resampling alone.

Average success across four simulation tasks (%)
VariantMean ± SD
Part-supervised point-wise63.5 ±4.2
Without part classification61.3 ±1.7
Without cross-instance alignment64.7 ±2.1
Without perturbation consistency68.9 ±1.8
Full model69.3 ±1.0

Standard deviations are over five per-seed task averages. The 0.4-point consistency-loss gap is smaller than the observed variation and is not conclusive evidence of a simulation benefit.

06 / Inside the representation

Corresponding parts stay alike.
Different parts stay distinct.

Qualitative feature visualizations show how functional regions are represented across object instances. Our aim is consistency for corresponding parts without losing the differences between handles, bodies, and other regions.

Real-world observations. A shared PCA projection is used within each method and category. Colors are not directly comparable across methods or categories, and these visualizations are qualitative, not a controlled noise-robustness test.
Part-aware features in simulation
Corresponding regions have similar feature colors across objects within a category, while different functional parts remain distinguishable.

Beyond Point-Attached Semantics

Stable part semantics.
More reliable manipulation.

Object-conditioned semantic fields offer a way to connect geometric observations to functional parts, helping a policy use the right part across different objects.

Watch the overview
Figure