sstzal/DiffTalk
quality grade F, 18 out of 100[CVPR2023] The implementation for "DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation"
- stars
- 472
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
340 results
[CVPR2023] The implementation for "DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation"
[CVPR2024] MMA-Diffusion: MultiModal Attack on Diffusion Models
Tarsier -- a family of large-scale video-language models, which is designed to generate high-quality video descriptions , together with good capability of general video understanding.
✨✨[NeurIPS 2025] VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.
R1-onevision, a visual language model capable of deep CoT reasoning.
open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
Official code implementation of Vary-toy (Small Language Model Meets with Reinforced Vision Vocabulary)
📖 A curated list of resources dedicated to hallucination of multimodal large language models (MLLM).
LaVIT: Empower the Large Language Model to Understand and Generate Visual Content
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)
OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data
VideoLLM-online: Online Video Large Language Model for Streaming Video (CVPR 2024)
Lumina-DiMOO - An Open-Sourced Multi-Modal Large Diffusion Language Model
Code for 3D-LLM: Injecting the 3D World into Large Language Models
[ICCV 2025] LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning
Official implementation for "Multimodal Chain-of-Thought Reasoning in Language Models" (stay tuned and more will be updated)
PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
SEED-Story: Multimodal Long Story Generation with Large Language Model
State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
Cell2Sentence: Teaching Large Language Models the Language of Biology
A flexible and efficient codebase for training visually-conditioned language models (VLMs)
[NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling
24,538 repositories in the index in total.