DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos

¹ Ludwig-Maximilians-Universität München (LMU Munich)
² Munich Center for Machine Learning (MCML)
³ University of Freiburg
⁴ Technical University of Munich (TUM)
⁵ Huawei Heisenberg Research Center (Munich)

ICRA 2026

Video Introduction

Abstract

Video diffusion models provide powerful realworld simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectoryconditioned video generation address this gap but often rely on 2D trajectories or single modality conditioning, which restricts their ability to produce controllable and consistent robotic demonstrations. We present DRAW2ACT, a depthaware trajectory-conditioned video generation framework that extracts multiple orthogonal representations from the input trajectory, capturing depth, semantics, shape and motion, and injects them into the diffusion model. Moreover, we propose to jointly generate spatially aligned RGB and depth videos, leveraging cross-modality attention mechanisms and depth supervision to enhance the spatio-temporal consistency. Finally, we introduce a multimodal policy model conditioned on the generated RGB and depth sequences to regress the robot’s joint angles. Experiments on Bridge V2, Berkeley Autolab, and simulation benchmarks show that DRAW2ACT achieves superior visual fidelity and consistency while yielding higher manipulation success rates compared to existing baselines.

Poster

BibTeX

@article{DRAW2ACT2026,
  title={DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos.},
  author={Bai Y, Yang L, Eskandar G, et al.},
  journal={arXiv e-prints, 2025: arXiv: 2512.14217.},
  year={2025},
  url={https://arxiv.org/abs/2512.14217}
}