Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[TMLR 2024] repository for VLN with foundation models
| Date | Stars |
|---|---|
| 2026-07-31 | 296 |
| 2026-08-05 | 296 |
| 2026-08-06 | 296 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Survey of Vision-and-Language Navigation
This is the official repository of "**[Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models](https://arxiv.org/pdf/2407.07035)**", a comprehensive survey
of recent progress in VLN with foundation models.
## 👏 Our Survey has been officially accepted by TMLR!!!
## Introduction
Vision-and-Language Navigation (VLN) has gained increasing attention over recent years and many approaches have emerged to advance their development. The remarkable achievements of foundation models have shaped the challenges and proposed methods for VLN research. In this survey, we provide a top-down review that adopts a principled framework for embodied planning and reasoning and emphasizes the current methods and future opportunities leveraging foundation models to address VLN challenges. We hope our in-depth discussions could provide valuable resources and insights: on the one hand, to document the progress and explore opportunities and potential roles for foundation models in this field, and on the other, to organize different challenges and solutions in VLN to foundation model researchers.
## Citation
If you find our work useful in your research, please consider citing:
@article{zhang2024vision,
title={Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models},
author={Zhang, Yue and Ma, Ziqiao and Li, Jialu and Qiao, Yanyuan and Wang, Zun and Chai, Joyce and Wu, Qi and Bansal, Mohit and Kordjamshidi, Parisa},
journal={arXiv preprint arXiv:2407.07035},
year={2024}
}
🔔 We will update this page frequently. If you believe additional work should be included, please do not hesitate to email us ([email protected]) or raise an issue. Your suggestions and comments are invaluable to ensuring the completeness of our resources.
## Content
---
- [Relevant Surveys](#relevant-surveys)
- [World Model](#word-modal)
- [Human Model](#human-modal)
- [VLN Agent](#vln-agent-learning-an-embodied-agent-for-reasoning-and-planning)
- [Behavior Analysis of the VLN Agent](#behavir-analysis)
---
## Relevant Surveys
| Title | Venue | Date | Code |
|:--------|:--------:|:--------:|:--------:|
[**Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions**](https://arxiv.org/abs/2203.12667)| ACL | 2022| [Github](https://github.com/eric-ai-lab/awesome-vision-language-navigation) |
[**Visual language navigation: A survey and open challenges**](https://link.springer.com/article/10.1007/s10462-022-10174-9)| - | 2023| - |
[**Vision-Language Navigation: A Survey and Taxonomy**](https://arxiv.org/abs/2108.11544)| - | 2021| - |
---
## World Model
A world model helps the VLN agent to understand their surrounding environments, predict how their actions would change the world state, and align their perception and actions with language instructions.
| Title | Venue | Date | Code |
|:--------|:--------:|:--------:|:--------:|
[**MapNav: A Novel Memory Representation via Annotated Semantic Maps for VLM-based Vision-and-Language Navigation**](https://arxiv.org/pdf/2502.13451)| ACL | 2025| - |
[**VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language Navigation**](https://arxiv.org/abs/2402.03561)| AAAI | 2024| - |
[**Volumetric Environment Representation for Vision-Language Navigation**](https://arxiv.org/pdf/2403.14158)| CVPR | 2024| [Github](https://github.com/DefaultRui/VLN-VER) |
[**Vision Language Navigation with Knowledge-driven Environmental Dreamer**](https://www.ijcai.org/proceedings/2023/0204.pdf)| IJCAI | 2023| - |
[**Frequency-enhanced Data Augmentation for Vision-and-Language Navigation**](https://openreview.net/pdf?id=eKFrXWb0sT)| NeurIPS | 2023| [Github](https://github.com/hekj/FDA?tab=readme-ov-file) |
[**Frequency-Enhanced Data Augmentation for Vision-and-Language Navigation**](https://proceedings.neurips.cc/paper_files/paper/2023/file/0d9e08f247ca7fbbfd5e5Excerpt of 20,887 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:f66ff1a643109491, llm:Repository description: '[TMLR 2024] repository for VLN with foundation models' — VLN = Vision-and-Language Navigation (vision + language navigation research with foundation models).
matched fp:f66ff1a643109491, llm:Repository description: '[TMLR 2024] repository for VLN with foundation models' — VLN = Vision-and-Language Navigation (vision + language navigation research with foundation models).
matched fp:f66ff1a643109491, llm:Repository description: '[TMLR 2024] repository for VLN with foundation models' — VLN = Vision-and-Language Navigation (vision + language navigation research with foundation models).
matched fp:f66ff1a643109491, llm:Repository description: '[TMLR 2024] repository for VLN with foundation models' — VLN = Vision-and-Language Navigation (vision + language navigation research with foundation models).