Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
This is the official repository for Retrieval Augmented Visual Question Answering
| Date | Stars |
|---|---|
| 2026-07-31 | 252 |
| 2026-08-04 | 252 |
| 2026-08-06 | 252 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Retrieval-augmented Visual Question Answering with Fine-grained Late-interaction Multi-modal Retrieval [](https://paperswithcode.com/sota/retrieval-on-infoseek?p=preflmr-scaling-up-fine-grained-late) [](https://paperswithcode.com/sota/visual-question-answering-vqa-on-infoseek?p=preflmr-scaling-up-fine-grained-late) [](https://paperswithcode.com/sota/retrieval-on-ok-vqa?p=fine-grained-late-interaction-multi-modal-1) [](https://paperswithcode.com/sota/visual-question-answering-on-ok-vqa?p=fine-grained-late-interaction-multi-modal-1) This is the official repository of the Retrieval Augmented Visual Question Answering (RAVQA) project. The project covers RAVQA and RAVQA-v2 (equipped with Fine-grained Late-interaction Multi-modal Retrieval). # 🔥🔥News - [19/12/2024] 🔥🔥🔥 We released the multilingual version(Chinese + English ) of PreFLMR, you can download PreFLMR ENCN model [here](https://huggingface.co/LinWeizheDragon/PreFLMR_ViT-L_ENCN). - [03/09/2024] We have uploaded the images used in the M2KR benchmark [here](https://huggingface.co/datasets/BByrneLab/M2KR_Images) . - [10/08/2024] We received many requests regarding adding multilingual abilities to PreFLMR. We announce that **we are now training the Chinese version of PreFLMR and will release it very soon**. Stay tuned! - [05/06/2024] 🔥🔥🔥The PreFLMR paper has been accepted to appear at ACL 2024! The camera-ready version of the paper has been updated [here](https://arxiv.org/abs/2402.08327) to include more details and analyses. Along with the acceptance, we have made some important updates to help you use the model and extend your research easier: - Added an evaluation script that reproduces the results in the PreFLMR paper [here](https://github.com/LinWeizheDragon/FLMR?tab=readme-ov-file#new-evaluate-the-preflmr-models-on-all-m2kr-benchmarks) - Added the updated benchmark results with the transformer implementation [here](#benchmark-results-for-preflmr-in-the-dedicated-flmr-codebase) - Added an example script to fine-tune PreFLMR on a custom retrieval dataset [here](https://github.com/LinWeizheDragon/FLMR?tab=readme-ov-file#new-finetune-the-preflmr-model-on-downstream-datasets) - **IMPORTANT**: fixed the OVEN data splits in the M2KR benchmark, and updated each entry with a fixed instruction to ensure the evaluation result is not affected by random sampling of instructions. Please delete your local cache and download the dataset again. - [13/04/2024] 🔥 We highlight another valuable and concurrent research on training instruction-following, universal, multi-task multi-modal retrievers: [UniIR: Training and Benchmarking Universal Multimodal Information Retrievers](https://tiger-ai-lab.github.io/UniIR/), which was done by the researchers of the University of Waterloo. They also shared the M-Beir benchmark which can be used to train and evaluate multi-modal universal information retrievers. In the near future, we may collaborate to combine the two benchmarks together to facilitate the advance of this field. - [06/03/2024] 🔥🔥🔥The implementation based on huggingface-transformers is now available [here](https://github.com/linweizhedragon/FLMR)! - [20/02/2024] 🔥🔥🔥 The [PreFLMR project page](https://preflmr.github.io/) has been launched! Explore a captivating demo showcasing PreFLMR_ViT-G, our largest model yet. Additionally, access pre-trained checkpoints and the M2KR
Excerpt of 42,834 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:0e638ad7938bf8ee, name:visual question answering, desc:visual question answering
matched fp:0e638ad7938bf8ee, name:retrieval augmented, desc:retrieval augmented