TAVISA Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning

Giacomo Spigler

Department of Intelligent Systems, Tilburg University

Overview of the TAVIS benchmark: TAVIS-Head and TAVIS-Hands task suites, the GR1T2 and Reachy2 robots, teleoperated datasets on Hugging Face, and the evaluation protocol (headcam vs fixedcam, ID vs OOD, GALT).
TAVIS at a glance: two task suites on two humanoid torsos, first-person VR teleoperation datasets, and an evaluation protocol that pairs an active head camera with a fixed workspace camera, in- and out-of-distribution splits, and the GALT metric for anticipatory gaze. Download (PDF)

Abstract

Active vision – where a policy controls its own gaze during manipulation – has emerged as a key capability for imitation learning, with multiple independent systems demonstrating its benefits in the past year. Yet there is no shared benchmark to compare approaches or quantify what active vision contributes, on which task types, and under what conditions. We introduce TAVIS, evaluation infrastructure for active-vision imitation learning, with two complementary task suites – TAVIS-Head (5 tasks, global search via pan/tilt necks) and TAVIS-Hands (3 tasks, local occlusion via wrist cameras) – on two humanoid torso embodiments (GR1T2, Reachy2), built on IsaacLab. TAVIS provides three evaluation primitives: a paired headcam-vs-fixedcam protocol on identical demonstrations; GALT (Gaze-Action Lead Time), a novel metric grounded in cognitive science and HRI that quantifies anticipatory gaze in learned policies; and procedural ID/OOD splits. Baseline experiments with Diffusion Policy and π0 reveal that (i) active vision generally helps, consistently across training runs, but benefits are task-conditional rather than uniform; (ii) multi-task policies degrade sharply under controlled distribution shifts on both suites; and (iii) imitation alone yields anticipatory gaze that lands on the task-relevant object, with lead times comparable to those of the human demonstrations, although head motion is less smooth than in the demonstrations, an effect of action chunking that success rate does not reveal.

Videos

Rollouts of the multi-task π0 policy with a head camera on GR1T2, shown at half speed: a third-person view (left) next to the robot's own head camera (right). The captions mark the head fixation and the grasp as the GALT detector finds them: the head settles on the target well before the hand gets there.

conditional-pick: the color of the card says which object to pick.
wait-then-act: wait for the light to turn green, then pick.
clutter-pick-cube: find the red cube among distractors.
clutter-pick-lift: find and lift the object named in the prompt.

More videos → first-person demonstrations for all eight tasks and both robots.

Tasks

TAVIS has two suites on two humanoid torsos (Fourier GR1T2 and Pollen Robotics Reachy2), with about 2,200 teleoperated demonstrations in LeRobot v3.0 format. Every demonstration records the head camera, a fixed workspace camera and both wrist cameras at the same time, so policies can be trained on the same demonstrations with or without control of their own viewpoint. Each task isolates one function of active vision:

TAVIS-Head · head camera on a pan/tilt neck

  • conditional-pick · conditional information gathering
  • wait-then-act · temporal monitoring
  • clutter-pick-cube · clutter disambiguation, target defined by appearance
  • clutter-pick-lift · clutter disambiguation, target defined by language
  • multi-shelf-scan · vertical workspace search

TAVIS-Hands · wrist cameras, head view occluded

  • peeking-box · looking into a container
  • occluded-reach · reaching around an occluder
  • blocked-clutter-pick-cube · wrist-only search
One demonstration per task: eight frames per task from the head camera (TAVIS-Head) or the wrist camera of the grasping arm (TAVIS-Hands).
One demonstration per task (GR1T2): head camera for TAVIS-Head, wrist camera of the grasping arm for TAVIS-Hands. Columns are aligned to task events: blue marks the head arriving at its final fixation and orange the grasp, so the interval between them is the GALT lead.

Citation

@article{spigler2026tavis,
  title={TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning},
  author={Spigler, Giacomo},
  journal={arXiv preprint arXiv:2605.07943},
  year={2026}
}