penghao-wu/vstar
quality grade D, 43 out of 100PyTorch Implementation of "V* : Guided Visual Search as a Core Mechanism in Multimodal LLMs"
- stars
- 714
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
340 results
PyTorch Implementation of "V* : Guided Visual Search as a Core Mechanism in Multimodal LLMs"
Detects reggaeton genre with Machine Learning and sends packets to disable BT speakers (hopefully)
A python library and collection of notebooks for making art with machine learning.
Magenta.js: Music and Art Generation with Machine Learning in the browser
🔮 Libraries & tools for enabling Machine Learning driven user-experiences on the web
DeepFuze is a state-of-the-art deep learning tool that seamlessly integrates with ComfyUI to revolutionize facial transformations, lipsyncing, Face Swapping, Lipsync Translation, video generation, and voice cloning.
Use deep learning to generate and harmonize music in the style of Bach
BLOCK (AAAI 2019), with a multimodal fusion library for deep learning models
Create automatic playlists by using Deep Learning to *listen* to the music.
Fusing Histology and Genomics via Deep Learning - IEEE TMI
Using deep learning to generate music in MIDI format.
This repository contains various models targetting multimodal representation learning, multimodal fusion for downstream tasks such as multimodal sentiment analysis.
FaceChain is a deep-learning toolchain for generating your Digital-Twin.
NeuralTalk is a Python+numpy project for learning Multimodal Recurrent Neural Networks that describe images with sentences.
MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
The simplest, fastest repository for training/finetuning small-sized VLMs.
[CVPR 2025] Magma: A Foundation Model for Multimodal AI Agents
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
[ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) | 基于CPM基础模型的中英双语多模态大模型系列
🩺 首个会看胸部X光片的中文多模态医学大模型 | The first Chinese Medical Multimodal Model that Chest Radiographs Summarization.
Codebase for Aria - an Open Multimodal Native MoE
Implementation of CoCa, Contrastive Captioners are Image-Text Foundation Models, in Pytorch
A curated list of Multimodal Related Research.
24,538 repositories in the index in total.