Top AI Repos โ open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
๐ Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
| Date | Stars |
|---|---|
| 2026-07-24 | 448 |
| 2026-07-25 | 448 |
| 2026-07-28 | 453 |
| 2026-07-30 | 453 |
| 2026-08-06 | 453 |
Today
โ stars today
This week
โ stars this week
This month
โ stars this month
Momentum
0.0
growth rate 0.00%/day
[](https://github.com/sindresorhus/awesome)
[](https://arxiv.org/abs/2512.16760)

[](https://github.com/worldbench/awesome-vla4ad/pulls)
# :sunglasses: Awesome VLA for Autonomous Driving
Autonomous driving has long relied on modular "Perception-Decision-Action" pipelines, whose hand-crafted interfaces and rule-based components often struggle in complex, dynamic, or long-tailed scenarios. Their cascaded structure also amplifies upstream perception errors, undermining downstream planning and control.
This survey reviews **vision-action (VA)** models and **vision-language-action (VLA)** models for autonomous driving. We trace the evolution from early VA approaches to modern VLA frameworks, and organize existing methods into two principal paradigms:
- **End-to-End VLA**, which integrates perception, reasoning, and planning within a single model.
- **Dual-System VLA**, which separates slow deliberation (via VLMs) from fast, safety-critical execution (via planners).
| <img width="100%" src="docs/figures/teaser.png"> |
|:-:|
For more details, kindly refer to our :books: [Paper](https://arxiv.org/abs/2512.16760), :globe_with_meridians: [Project Page](https://worldbench.github.io/vla4ad), and :hugs: [HuggingFace Leaderboard](https://huggingface.co/spaces/worldbench/vla4ad).
### :books: Citation
If you find this work helpful for your research, please kindly consider citing our paper:
```bib
@article{survey_vla4ad,
title = {Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future},
author = {Tianshuai Hu and Xiaolu Liu and Song Wang and Yiyao Zhu and Ao Liang and Lingdong Kong and Guoyang Zhao and Zeying Gong and Jun Cen and Zhiyu Huang and Xiaoshuai Hao and Linfeng Li and Hang Song and Xiangtai Li and Jun Ma and Shaojie Shen and Jianke Zhu and Dacheng Tao and Ziwei Liu and Junwei Liang},
journal = {arXiv preprint arXiv:2512.16760},
year = {2025},
}
```
```bib
@article{survey_3d_4d_world_models,
title = {{3D} and {4D} World Modeling: A Survey},
author = {Lingdong Kong and Wesley Yang and Jianbiao Mei and Youquan Liu and Ao Liang and Dekai Zhu and Dongyue Lu and Wei Yin and Xiaotao Hu and Mingkai Jia and Junyuan Deng and Kaiwen Zhang and Yang Wu and Tianyi Yan and Shenyuan Gao and Song Wang and Linfeng Li and Liang Pan and Yong Liu and Jianke Zhu and Wei Tsang Ooi and Steven C. H. Hoi and Ziwei Liu},
journal = {arXiv preprint arXiv:2509.07996},
year = {2025}
}
```
# Table of Contents
- [**1. Vision-Action Models**](#1-vision-action-models)
- [Action-Only Models](#one-action-only-models)
- [Perception-Action Models](#two-perception-action-models)
- [Image-Based World Models](#three-image-based-world-models)
- [Occupancy-Based World Models](#four-occupancy-based-world-models)
- [Latent-Based World Models](#five-latent-based-world-models)
- [**2. Vision-Language-Action Models**](#2-vision-language-action-models)
- [Textual Action Generator](#one-textual-action-generator)
- [Numerical Action Generator](#two-numerical-action-generator)
- [Explicit Action Guidance](#three-explicit-action-guidance)
- [Implicit Representations Transfer](#four-implicit-representations-transfer)
- [**3. Datasets & Benchmarks**](#3-datasets--benchmarks)
- [Vision-Action Datasets](#one-vision-action-datasets)
- [Vision-Language-Action Datasets](#two-vision-language-action-datasets)
- [**4. Applications**](#4-applications)
- [**5. Other Resources**](#5-other-resources)
## 1. Vision-Action Models
### :one: Action-Only Models
> :timer_clock: In chronologiExcerpt of 96,932 characters
Read on GitHubWould you bet a product on this? Bounded 0โ100 and slow moving.
matched fp:a03ce24f32752417, topic:embodied-ai, topic:autonomous-driving, desc:autonomous driving
matched fp:a03ce24f32752417, topic:large-language-models, topic:llm
matched fp:a03ce24f32752417, topic:vlm, desc:vision-language, readme:vision-language