jolibrain/colette
quality grade B, 66 out of 100Multimodal RAG to search and interact locally with technical documents of any kind
- stars
- 301
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
340 results
Multimodal RAG to search and interact locally with technical documents of any kind
Visualizing the attention of vision-language models
✨✨Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy
[ICCV 2025] Official Implementation for "Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition"
[WACV 2024 Survey Paper] Multimodal Large Language Models for Autonomous Driving
[NeurIPS'24] This repository is the implementation of "SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models"
[CVPR 2024] Official PyTorch Code for "PromptKD: Unsupervised Prompt Distillation for Vision-Language Models"
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Recipes for learning, fine-tuning, and adapting ColPali to your multimodal RAG use cases. 👨🏻🍳
[ICLR 2026 Oral (top 1.2%)] Official implementation of DepthLM
[CVPR 2024 Highlight] OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
[ECCV 2026] Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
[CVPR 2025 Highlight] Official code for "Olympus: A Universal Task Router for Computer Vision Tasks"
An open-source implementation for training LLaVA-NeXT.
Official repository for VisionZip (CVPR 2025)
[IJCV 2026] Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
Evaluating text-to-image/video/3D models with VQAScore
A curated list of awesome prompt/adapter learning methods for vision-language models like CLIP.
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.
Awesome Unified Multimodal Models
The implementation of "Prismer: A Vision-Language Model with Multi-Task Experts".
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
24,538 repositories in the index in total.