We train a model for increasing dataset subets from 5 to 70 episodes. BIND achieves strong success even at 5 demos while ACT still struggles at 70 demos.
We train only with the cup on the left side of the table and test with the cup on the right side. ACT doesn't even reach towards the object while our policy degrades more gracefully to 60% success.
We train at one viewpoint and test at three OOD viewpoints: BIND retains near 80% success whereas ACT breaks completely.
We train BIND on more long-horizon and dextrous tasks across fine grasps and wiping motions.




We also verify our experiments in sim with precise object and viewpoint definitions with a simple mujoco pick+place environment.
MuJoCo cube pick-and-place, fixed-yaw, train on the LEFT half of the workspace, eval on the 110-position held-out RIGHT half. Each cell = one method's per-position outcome grid (green = success, red = fail). Right chart = interactive 6-way success bars. All 6 methods are FRESH phe retrains under identical recipe. Only Ours generalizes (21%); all 5 baselines collapse to 0%.






Camera-viewpoint OOD on the same task, trained on 1 / 5 / 9 canonical cameras and evaluated on the 49-camera grid (196 eps). Each cell = one method's per-viewpoint mean progress across the 3 training-cam configurations. Right chart = interactive 6-way progress-vs-cams curve. Ours (FRESH) climbs 22 → 39 → 44; MolmoAct2* is still the [P] import at 37 → 44 → 44 (FRESH retrain not yet run); ACT reaches 49; Diffusion-x0 27; DP3's 5/9-cam mv-render is deferred.





BIND is a general action head. We swap in various backbones (vision-based, geometry-based, VLMs) and show policy success.