TencentARC/SEED-Story
quality grade D, 42 out of 100SEED-Story: Multimodal Long Story Generation with Large Language Model
- stars
- 885
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
344 results
SEED-Story: Multimodal Long Story Generation with Large Language Model
State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
Cell2Sentence: Teaching Large Language Models the Language of Biology
A flexible and efficient codebase for training visually-conditioned language models (VLMs)
[NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.
GPT4Tools is an intelligent system that can automatically decide, control, and utilize different visual foundation models, allowing the user to interact with images during a conversation.
🛰️ Official repository of paper "RemoteCLIP: A Vision Language Foundation Model for Remote Sensing" (IEEE TGRS)
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
Vision-Language Pathology Foundation Model - Nature Medicine
Frontier Multimodal Foundation Models for Image and Video Understanding
A Scientific Multimodal Foundation Model
A feed-forward 3D foundation model for reconstructing scenes from streaming data
Transformers 3rd Edition
Fine-tune Gemma 4 and 3n with audio, images and text on Apple Silicon, using PyTorch and Metal Performance Shaders.
[T-IV] This repository collects research papers of large Vision Language Models in Autonomous driving and Intelligent Transportation System. The repository will be continuously updated to track the latest update.
Official repo for "More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models" (ICLR 2026)
[NeurIPS 2025]《SD-VLM: Spatial Measuring and Understanding with Depth-encoded Vision Language Models》
This repository collects and organises state‑of‑the‑art papers on spatial reasoning for Multimodal Vision–Language Models (MVLMs).
[ECCV2024] 🐙Octopus, an embodied vision-language model trained with RLEF, emerging superior in embodied visual planning and programming.
Harnessing 1.4M GPT4V-synthesized Data for A Lite Vision-Language Model
[CVPR 2024 Highlight] Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
24,523 repositories in the index in total.