Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance

Published in The Fourteenth International Conference on Learning Representations (ICLR), 2026

Align-Then-stEer (ATE) is a data-efficient, plug-and-play adaptation framework for pretrained vision-language-action models. It aligns disparate action spaces into a unified latent representation and steers VLA generation through latent guidance.

Project page

Recommended citation: Y. Zhang†, C. Wang†, O. Lu†, Y. Zhao, Y. Ge, Z. Sun, X. Li, C. Zhang, C. Bai*, and X. Li*. (2026). "Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance." ICLR. † Equal contribution; * corresponding author.
Download Paper