Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[WACV 2024 Survey Paper] Multimodal Large Language Models for Autonomous Driving
| Date | Stars |
|---|---|
| 2026-07-24 | 313 |
| 2026-07-25 | 313 |
| 2026-07-28 | 313 |
| 2026-07-30 | 313 |
| 2026-08-06 | 313 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# A Survey on Multimodal Large Language Models for Autonomous Driving ### We add new references from CVPR 2024 in our repo, some references are from [自动驾驶之心](https://github.com/autodriving-heart/CVPR-2024-Papers-Autonomous-Driving). ### :boom: News: MAPLM (Tencent, UIUC) and LaMPilot (Purdue University) from our team are accepted by CVPR 2024. ### News: [LLVM-AD Workshop](https://llvm-ad.github.io/schedule/) is successfully organized at WACV 2024.  [WACV 2024 Proceedings](https://openaccess.thecvf.com/WACV2024_workshops/LLVM-AD) | [Arxiv](https://arxiv.org/abs/2311.12320) | [Workshop](https://llvm-ad.github.io/) | [Report by 机器之心](https://www.jiqizhixin.com/articles/2023-12-18-4) ### Summary of the 1st WACV Workshop on Large Language and Vision Models for Autonomous Driving (LLVM-AD) ## Abstract With the emergence of Large Language Models (LLMs) and Vision Foundation Models (VFMs), multimodal AI systems benefiting from large models have the potential to equally perceive the real world, make decisions, and control tools as humans. In recent months, LLMs have shown widespread attention in autonomous driving and map systems. Despite its immense potential, there is still a lack of a comprehensive understanding of key challenges, opportunities, and future endeavors to apply in LLM driving systems. In this repo, we present a systematic investigation in this field. We first introduce the background of Multimodal Large Language Models (MLLMs), the multimodal models development using LLMs, and the history of autonomous driving. Then, we overview existing MLLM tools for driving, transportation, and map systems together with existing datasets and benchmarks. Moreover, we summarized the works in The 1st WACV Workshop on Large Language and Vision Models for Autonomous Driving (LLVM-AD), which is the first workshop of its kind regarding LLMs in autonomous driving. To further promote the development of this field, we also discuss several important problems regarding using MLLMs in autonomous driving systems that need to be solved by both academia and industry. ## Awesome Papers ### MLLM for Perception & Planning & Control for Autonomous Driving Please ping us if you find any interesting new papers in this area. We will update them into the Table. And all of them will be included in the next version of the survey paper. | Model | Year | Backbone | Task | Modality | Learning | Input | Output | |-----------------------------------------------------------------------------|------|-------------------------|-----------------------------------|--------------------------|---------------------|----------------------|--------------------| | [Driving with LLMs](https://arxiv.org/abs/2310.01957) | 2023 | LLaMA | Perception Control | Vision, Language | Finetuning | Vector Query | Response / Actions | | [Talk2BEV](https://arxiv.org/abs/2310.02251) | 2023 | Flan5XXL Vicuna-13b | Perception Planning | Vision, Language | In-context learning | Image Query | Response | | [GAIA-1](https://arxiv.org/abs/2309.17080) | 2023 | - | Planning | Vision, Language | Pretraining | Video Prompt | Video | | [Dilu](https://arxiv.org/abs/2309.16292) | 2023 | GPT-3.5 GPT-4 | Planning Control | Language | In-context learning | Text | Action | | [Drive as You Speak](https://arxiv.org/abs/2309.10228) | 2023 | GPT-4 | Planning
Excerpt of 19,560 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:06628bfa55847af6, topic:multimodal, topic:vision-language-model, name:multimodal
matched fp:06628bfa55847af6, topic:large-language-models, topic:foundation-models
matched fp:06628bfa55847af6, topic:autonomous-driving, name:autonomous driving, desc:autonomous driving