SkalskiP/awesome-foundation-and-multimodal-models
quality grade D, 40 out of 100👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]
- stars
- 638
- stars gained this week
- —this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
344 results
👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]
This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.
Camera monitoring with VLM
Meta-Transformer for Unified Multimodal Learning
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
Show, Attend, and Tell | a PyTorch Tutorial to Image Captioning
Collection of AWESOME vision-language models for vision tasks
24,523 repositories in the index in total.