NeurIPS 2026

UniRAP: Towards Unified Part-level Physical Affordance Reasoning and Actionable Perception

From linguistic intent to pixel-level affordances and executable contact geometry—in one unified model.

Linfei Li1 Ruining Hu1 Lin Zhang1,* Zhong Wang1 Fengyi Zhang4 Ying Shen1 Binqiang Wang2 Xin Zhang3

1Tongji University   2IEIT SYSTEMS   3NAIS, Neusoft Corporation

4The University of Queensland   *Corresponding author

UniRAP task overview across segmentation, affordance perception, and actionable perception
UniRAP transfers generic segmentation knowledge to part-level affordance reasoning and spatiotemporally consistent actionable perception.
Video

UniRAP in action

The idea

Reason about where to interact—and how

Vision-language perception has made impressive progress in aligning language with visual observations, yet grounding high-level semantics into part-level physical interaction remains challenging. UniRAP is a unified model that maps language instructions to actionable geometric representations. Through a unified interface token mechanism, it integrates images or videos, textual instructions, and optional visual prompts into a shared spatiotemporal representation space.

A Unified Affordance Decoder jointly predicts object detections, part-level affordance masks, and 4-DoF interaction poses. A curriculum transfer strategy then adapts the model from general visual parsing to interaction-aware perception, improving data efficiency when high-quality interaction labels are scarce.

01

Unified interfaces

Text, points, and boxes share a token-level interface for reasoning over both images and videos.

02

Actionable geometry

One decoder produces bounding boxes, part-level masks, and precise 4-DoF interaction poses.

03

Transfer by curriculum

Four progressive stages bridge abundant segmentation data and scarce interaction annotations.

Method

A unified path from prompts to actions

Architecture overview of the UniRAP model
Visual input, language, and optional visual references are fused through <REF> and <SEG> interfaces, then decoded into masks, detections, and 4-DoF poses.
Unified Affordance Decoder

One representation, multiple granularities

The UAD reuses dense segmentation features for three coupled tasks. Mask evidence initializes region geometry; task-specific tokens condition shared visual features; consistency constraints let detection and contact prediction refine one another.

DetectionSegmentation4-DoF pose
Unified Affordance Decoder architecture
Curriculum transfer training

From seeing objects to understanding interaction

Four-stage curriculum transfer training pipeline
1General segmentation
2Affordance pre-training
3Visual referring pre-training
4Actionable affordance fine-tuning
Results

Strong generalization from a 3B model

UniRAP performs consistently across referring segmentation, affordance grounding, video reasoning, and fine-grained interaction prediction.

80.9RefCOCO val cIoU
64.1ReVOS overall J&F
63.3Novel-split grasp accuracy
3×Novel grasp gain over RealVLG-R1
Quantitative comparisons

Results from the paper

Gray column groups denote zero-shot benchmarks. Best results among the displayed methods are highlighted.

Referring affordance segmentation
MethodSizeHANDALGraspNet SeenGraspNet Novel3DOI
gIoUcIoUgIoUcIoUgIoUcIoUgIoUcIoU
LISA7B15.411.817.717.725.224.121.513.7
UniPixel7B26.016.445.012.548.011.541.524.9
RAGNet7B60.560.363.364.045.633.237.437.4
UniRAP3B63.965.770.664.151.642.352.137.5
Reasoning affordance segmentation
MethodSizeHANDAL EasyHANDAL Hard3DOI
gIoUcIoUgIoUcIoUgIoUcIoU
LISA7B15.511.912.38.112.38.1
UniPixel7B26.715.724.511.638.521.0
RAGNet7B58.358.158.257.838.139.4
UniRAP3B63.264.162.963.855.754.5
Actionable perception on RealVLG-Benchmark
Method (3B)Seen
Seg. Fβ / Grasp gAcc
Similar
Seg. Fβ / Grasp gAcc
Novel
Seg. Fβ / Grasp gAcc
Qwen2.5-VL + SFT76.2 / 1.775.4 / 2.146.2 / 1.5
RealVLG-R1 (GRPO)83.9 / 40.384.1 / 31.949.4 / 17.1
RealVLG-R1 (GSPO)88.7 / 33.685.0 / 30.649.1 / 9.1
UniRAP92.0 / 56.789.4 / 56.775.8 / 63.3

General & reasoning segmentation

UniRAP reaches state-of-the-art results on referring expression segmentation at the 3B scale and remains competitive with larger models, while maintaining strong temporal consistency for video queries.

Part-level affordance grounding

On explicit and reasoning-based affordance segmentation, UniRAP improves generalization to unseen objects. On zero-shot 3DOI, it raises gIoU and cIoU by 17.6 and 15.1 points over RAGNet.

Dynamic examples

Image understanding that stays consistent over time

Examples extracted from the UniRAP supplementary presentation. Videos load only when they approach the viewport.

Referring video object segmentation

Referring sequence 1
Referring sequence 2
Referring sequence 3
Referring sequence 4
Referring sequence 5
Referring sequence 6

Reasoning video object segmentation

Reasoning sequence 1
Reasoning sequence 2
Reasoning sequence 3
Reasoning sequence 4
Reasoning sequence 5
Reasoning sequence 6
Reasoning sequence 7

Spatiotemporal actionable perception

Actionable sequence 1
Actionable sequence 2
Actionable sequence 3
Actionable sequence 4
Actionable sequence 5
Actionable sequence 6
Real-world experiments

Perception that can be executed

A Franka FR3 with an eye-in-hand RealSense D435i directly executes UniRAP's predicted 4-DoF poses across 13 household objects and both referring and reasoning prompts.

Franka FR3 real-world robotic experiment setup
+26.2%

overall grasp success rate compared with RealVLG-R1 (3B), across 10 trials per object.

Referring grasping demonstrations

“Please give me the blue spoon.”
“Please grasp the orange cup.”
“Please pick up the grey and black razor.”
“Please give me the screwdriver.”
“Please give me the milk box.”
“Please grasp the green brush.”

Reasoning grasping demonstrations

Braise pork ribs → pressure cooker
Blue tool on the induction cooker → stapler
Stir-fry vegetables → shovel
Cut fruit → knife
Remove lint from clothes → rolling sticker
Make a cup of tea → teapot
Pour water from a lidded vessel → teapot
Writable and erasable on a whiteboard → marker pen
Citation

BibTeX

@inproceedings{li2026unirap,
  title     = {UniRAP: Towards Unified Part-level Physical Affordance Reasoning and Actionable Perception},
  author    = {Li, Linfei and Hu, Ruining and Zhang, Lin and Wang, Zhong and Zhang, Fengyi and Shen, Ying and Wang, Binqiang and Zhang, Xin},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}