Build object context
A frozen Utonia backbone extracts support features. A trainable adapter maps them into an object feature cache with three spatial feature planes and a pooled global feature.
Anonymous supplementary material
Stable Object-Centric Semantic Fields for Robust Manipulation
Watch the research overview 2:4401 / The idea
A robot needs to know not just where an object is, but which part to use. A hammer is grasped by its handle and makes contact with its head, even when its shape or appearance changes.
Yet features attached to observed points can change with viewpoint, sensor noise, and sampling. We learn an object-centric semantic field that uses object context to provide more consistent part-aware features for a manipulation policy.
A partial point cloud
Geometry with part supervision
Resampled 3D locations
Semantic clouds + scene + state
Robotic manipulation often requires identifying functional parts, such as a mug handle or a hammer head. However, features attached to observed 3D points can vary with viewpoint and sensor noise, giving a policy inconsistent representations of the same part. We propose an object-centric semantic field to provide more consistent part-aware features for manipulation. We use the observed object cloud to build a continuous field, then read features from this field at 3D locations independently resampled from the cloud. Each feature uses the sampled object support as context, rather than directly reusing an individual point descriptor. Part classification distinguishes functional regions, cross-instance alignment brings corresponding part features together, and perturbation consistency encourages similar features under observation changes. The queried coordinates and features form semantic point clouds that are supplied to a DP3-based policy. We evaluate the approach on four RoboTwin simulation tasks and four real-world bimanual tasks, achieving average success rates of 69.3% and 67.5%, respectively. These improve on Utonia Point-wise by 7.0 and 32.5 percentage points, respectively, with real-world tests on held-out objects. A point-wise control with matched part supervision scores 63.5% in simulation, compared with our 69.3%. These results highlight the value of stable, object-conditioned semantic fields for manipulation across object instances and varying observations.
02 / Research overview
From the representation to the robot.
2 min 44 secEnglish narration · Subtitles included
03 / Method
A frozen 3D encoder supplies geometry. Part supervision gives that geometry meaning. A continuous field brings both together at the locations used by the policy.
A frozen Utonia backbone extracts support features. A trainable adapter maps them into an object feature cache with three spatial feature planes and a pooled global feature.
The decoder reads features at resampled query locations. PartNext labels supervise part classification; cross-instance alignment and perturbation consistency act on the embeddings.
Separate PointNets encode scene and semantic object clouds. Their features and the robot state condition a DP3 policy. The semantic field stays frozen during policy training.
A different object representation. The same action representation and diffusion learning objective.
04 / Real-world manipulation
Each task policy learns from 50 demonstrations with designated training objects. At test time, it is applied to unseen physical instances without object-specific retraining.
Success rates cover all 20 trials per task, using randomly selected held-out objects, not just the executions shown here. Category-level fields are trained on PartNext models, with no scans or fine-tuning of the physical test objects.
05 / Evaluation
+7.0 points over Utonia Point-wise
+32.5 points on held-out objects
Over the part-supervised point-wise control
20 trials per task. Rates are percentages of successful trials; no real-world standard deviations are reported.
| Method | Grasp Mug | Beat Cube | Stir Mug | Pour Water | Average |
|---|---|---|---|---|---|
| DP3 | 35.0 | 40.0 | 15.0 | 20.0 | 27.5 |
| DINOv2 Lifting | 35.0 | 45.0 | 25.0 | 25.0 | 32.5 |
| Utonia Point-wise | 35.0 | 45.0 | 35.0 | 25.0 | 35.0 |
| Ours | 85.0 | 85.0 | 50.0 | 50.0 | 67.5 |
Single-arm and dual-arm manipulation
| Method | Hang Mug | Beat Hammer | Open Microwave | Put Cabinet | Average |
|---|---|---|---|---|---|
| DP3 | 23.0 ±4.0 | 72.0 ±2.0 | 62.0 ±3.0 | 72.0 ±2.0 | 57.3 |
| DINOv2 Lifting | 23.0 ±7.0 | 78.0 ±5.0 | 58.0 ±3.0 | 68.0 ±4.0 | 56.8 |
| Utonia Point-wise | 31.0 ±4.0 | 77.0 ±2.0 | 65.0 ±1.0 | 76.0 ±1.0 | 62.3 |
| G3Flow | 25.0 ±5.0 | 78.0 ±5.0 | 64.0 ±2.0 | 78.0 ±3.0 | 61.3 |
| Ours | 37.0 ±3.0 | 84.0 ±1.0 | 73.0 ±1.0 | 83.0 ±1.0 | 69.3 |
50 demonstrations per task; 3,000 training epochs; 100 evaluation episodes per task and seed. Simulation uses the same task and object setup for training and evaluation, unlike the held-out-object real-world protocol. G3Flow is evaluated in simulation only.
A controlled comparison
The part-supervised point-wise control uses the same encoder, training data, and supervision, but reads features attached to observed points. Our field reaches 69.3%, versus 63.5% for this control.
This comparison tests the field-based representation against direct point-wise readout. It does not isolate query resampling alone.
| Variant | Mean ± SD |
|---|---|
| Part-supervised point-wise | 63.5 ±4.2 |
| Without part classification | 61.3 ±1.7 |
| Without cross-instance alignment | 64.7 ±2.1 |
| Without perturbation consistency | 68.9 ±1.8 |
| Full model | 69.3 ±1.0 |
Standard deviations are over five per-seed task averages. The 0.4-point consistency-loss gap is smaller than the observed variation and is not conclusive evidence of a simulation benefit.
06 / Inside the representation
Qualitative feature visualizations show how functional regions are represented across object instances. Our aim is consistency for corresponding parts without losing the differences between handles, bodies, and other regions.
Beyond Point-Attached Semantics
Object-conditioned semantic fields offer a way to connect geometric observations to functional parts, helping a policy use the right part across different objects.
Watch the overview