BIND: Binding 3D Robot Actions to 2D Image Features

Cameron Smith1, Arsh Tangri2, Vitor Guizilini2, Yue Wang1, Zubair Irshad2, Sergey Zakharov2
1University of Southern California    2Toyota Research Institute
In submission · 2026
🔍 Investigation
🤔Question
If image features are so rich and spatially robust, why are robot policies we train on top of them so data-inefficient and spatially-fragile?
💡Intuition
Policies have to learn the relationship between their robot action targets (e.g. EEF XYZ position) and their projecting image features (i.e. where that EEF XYZ position would land on the image).
🔗Solution
We explicitly BIND candidate robot actions to their projecting 2D image features.
📈Result
Large gains in data-efficiency and spatial robustness to out-of-distribution object positions and camera viewpoints.
Method Overview

Experiment 1.1: Data Efficiency

We train a model for increasing dataset subets from 5 to 70 episodes. BIND achieves strong success even at 5 demos while ACT still struggles at 70 demos.

5ep Train Data efficiency — 5-demo training set (3×2 tile of representative frames)
70ep Train Data efficiency — 70-demo training set (10×7 tile of representative frames)
BIND
100
100
100
ACT
0
0
0
Data Efficiency results chart

Experiment 1.2: Out-of-Distribution Object Position

We train only with the cup on the left side of the table and test with the cup on the right side. ACT doesn't even reach towards the object while our policy degrades more gracefully to 60% success.

TRAIN OOD object position — training distribution (left half)
TEST OOD object position — held-out positions (right half)
BIND
100
100
100
ACT
0
0
0
Out-of-Distribution Object Position results chart

Experiment 1.3: Out-of-Distribution Viewpoint

We train at one viewpoint and test at three OOD viewpoints: BIND retains near 80% success whereas ACT breaks completely.

TRAIN OOD viewpoint — training viewpoint
TEST OOD viewpoint — held-out viewpoint (view 2)
BIND
100
100
100
ACT
0
0
0
Out-of-Distribution Viewpoint results chart

Experiment 2: More Tasks

We train BIND on more long-horizon and dextrous tasks across fine grasps and wiping motions.

Teapot40 demos
Fold Towel35 demos
Cup Spill30 demos
Cup Stacking40 demos
OursBIND
100
100
100
100
ACTbaseline
0
0
0
0
Molmo5B VLA
0
0
0
0
Teapot per-method success bars
Fold Towel per-method success bars (3-rollout sample)
Cup Spill per-method success bars
Cup Stacking per-method success bars

Experiment 3: Simulation OOD Object Position and Viewpoint Testing

We also verify our experiments in sim with precise object and viewpoint definitions with a simple mujoco pick+place environment.

In-DistributionLeft Side of Table
OOD Object PositionRight Side of Table
OOD ViewpointNew Viewpoint Grid
In-dist progress — 6 methods
OOD-obj progress — 6 methods
Viewpoint curve — 6 methods

Experiment 4: Backbones — Same Head, Different Encoders

BIND is a general action head. We swap in various backbones (vision-based, geometry-based, VLMs) and show policy success.

DINOv3
baseline
DAv3
geometric
DynaFlip
robotics ViT
PaliGemma
VLM
π0
VLA

Method: Forward Pass Visualization