Bo Jiang.
I am a 4th-year Ph.D. candidate at Huazhong University of Science and Technology, where I am fortunate to be advised by Prof. Xinggang Wang and Prof. Wenyu Liu. Currently, I am an intern at ByteDance Seed, working on VLA models for physical AI. I also interned at Horizon Robotics, where I was advised by Dr. Qian Zhang, and at Applied Intuition, where I worked under the guidance of Chief Scientist Dr. Wei Zhan.
My research focuses on physical intelligence: building systems that can connect perception, reasoning, and action in the physical world. I am especially interested in end-to-end driving, vision-language-action models, and video world models.
News
- 2026.08 Senna was accepted by IJCV 2026.
- 2026.05 We released LaMo, a self-supervised latent motion prior for physically realistic video generation.
- 2026.01 VADv2 was accepted by ICLR 2026.
- 2025.09 RAD was accepted by NeurIPS 2025.
- 2025.02 DiffusionDrive was accepted by CVPR 2025 Highlight.
- 2023.07 VAD was accepted by ICCV 2023.
Selected Publications

Video World Model
LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation
Bo Jiang, Depu Meng, Yihan Hu, Yichen Xie, Tianshuo Xu, Wei Zhan
arXiv preprint, 2026
Learning latent motion priors through self-supervision to improve physical realism of generated videos across diverse scenarios.

Probabilistic Planning
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
Bo Jiang*, Shaoyu Chen*, Hao Gao, Bencheng Liao, Qian Zhang, Wenyu Liu, Xinggang Wang
International Conference on Learning Representations (ICLR), 2026
Probabilistic planning captures driving possibilities, allowing an end-to-end policy to account for uncertainty when acting.

Dual-System Driving VLA
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
Bo Jiang, Shaoyu Chen, Bencheng Liao, Xingyu Zhang, Wei Yin, Qian Zhang, Chang Huang, Wenyu Liu, Xinggang Wang
International Journal of Computer Vision (IJCV), 2026
A dual-system driving VLA connects vision-language reasoning with end-to-end planning, bridging decisions and actions.

Vectorized Autonomous Driving
VAD: Vectorized Scene Representation for Efficient Autonomous Driving
Bo Jiang*, Shaoyu Chen*, Qing Xu, Bencheng Liao, Jiajie Chen, Hao Gao, Qian Zhang, Wenyu Liu, Chang Huang, Xinggang Wang
International Conference on Computer Vision (ICCV), 2023
Vectorized scene representation toward real-time end-to-end autonomous driving.
More Publications
-
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
-
RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
Advances in Neural Information Processing Systems (NeurIPS), 2025
-
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
arXiv preprint arXiv:2503.07608
-
MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
International Journal of Computer Vision (IJCV), 2024
-
Lane Graph as Path: Continuity-Preserving Path-Wise Modeling for Online Lane Graph Construction
European Conference on Computer Vision (ECCV), 2024
-
HOPE: Hierarchical Spatial-Temporal Network for Occupancy Flow Prediction
CVPR 2022 Workshop on Autonomous Driving
Education
-
Huazhong University of Science and Technology
M.S. & Ph.D. in Information and Communication Engineering
-
Central South University
B.Sc. in Data Science and Big Data Technology
Internships
-
ByteDance Seed
Multimodal Interaction and World Model Research Intern
-
Applied Intuition
Physical AI World Model Research Intern
-
Horizon Robotics
Autonomous Driving Algorithm Research Intern
Honors & Awards
-
First-Class Ph.D. Scholarship, Huazhong University of Science and Technology
-
1st Place Winner, Google Waymo Open Dataset Challenge, Occupancy Flow Prediction Track
-
Outstanding Graduate, Central South University
Invited Talks
-
Advancing E2E-AD via Multimodal Planning, Reinforced Fine-Tuning, and Language Modality Integration
-
Closed-Loop Reinforcement Learning for On-Policy Autonomous Driving
-
Rethinking Vision-Language-Action Models for End-to-end Autonomous Driving