DirtyHarryLYL/LLM-in-Vision
quality grade D, 35 out of 100Recent LLM-based CV and related works. Welcome to comment/contribute!
- stars
- 871
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
344 results
Recent LLM-based CV and related works. Welcome to comment/contribute!
[CVPR 2024] OneLLM: One Framework to Align All Modalities with Language
LLM2CLIP significantly improves already state-of-the-art CLIP models.
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Make text LLMs listen and speak
[CVPR2023] The implementation for "DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation"
[CVPR2024] MMA-Diffusion: MultiModal Attack on Diffusion Models
Translate EPUB books using Large Language Models while preserving the original text. The translated content is displayed side-by-side with the original, creating bilingual books perfect for language learning and cross-reference reading.
Tarsier -- a family of large-scale video-language models, which is designed to generate high-quality video descriptions , together with good capability of general video understanding.
✨✨[NeurIPS 2025] VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.
R1-onevision, a visual language model capable of deep CoT reasoning.
open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
Official code implementation of Vary-toy (Small Language Model Meets with Reinforced Vision Vocabulary)
📖 A curated list of resources dedicated to hallucination of multimodal large language models (MLLM).
LaVIT: Empower the Large Language Model to Understand and Generate Visual Content
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)
OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data
VideoLLM-online: Online Video Large Language Model for Streaming Video (CVPR 2024)
Lumina-DiMOO - An Open-Sourced Multi-Modal Large Diffusion Language Model
Code for 3D-LLM: Injecting the 3D World into Large Language Models
[ICCV 2025] LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning
Official implementation for "Multimodal Chain-of-Thought Reasoning in Language Models" (stay tuned and more will be updated)
PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
24,523 repositories in the index in total.