richard-peng-xia/MMed-RAG
quality grade D, 45 out of 100[ICLR'25] MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
- stars
- 337
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
344 results
[ICLR'25] MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
The official repo for "SpatialBot: Precise Spatial Understanding with Vision Language Models.
Benchmark environment for evaluating vision-language models (VLMs) on popular video games!
VisualGPT, CVPR 2022 Proceeding, GPT as a decoder for vision-language models
Real-time Vision Language Model interaction via webcam - WebRTC-based web interface
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
[NeurIPS 2025] Efficient Reasoning Vision Language Models
Experiments and data for the paper "When and why vision-language models behave like bags-of-words, and what to do about it?" Oral @ ICLR 2023
[ICLR'25] Official code for the paper 'MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs'
The code for "TokenPacker: Efficient Visual Projector for Multimodal LLM", IJCV2025
Convert files (PDF, image, Word, PPT, Excel, notebooks, code snippets) to markdown using powerful multimodal LLM
(CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.
HPT - Open Multimodal LLMs from HyperGAI
Medical Multimodal LLMs
[ICLR 2025] Official implementation of "DiffSplat: Repurposing Image Diffusion Models for Scalable 3D Gaussian Splat Generation".
Reading notes about Multimodal Large Language Models, Large Language Models, and Diffusion Models
Connecting openFrameworks to Google MediaPipe Machine Learning Framework over UDP
[ECCV 2024 Oral] Code for paper: An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
[NeurIPS 2025] OmniSVG is the first family of end-to-end multimodal SVG generators that leverage pre-trained Vision-Language Models (VLMs), capable of generating complex and detailed SVGs, from simple icons to intricate anime characters.
😎 up-to-date & curated list of awesome Attacks on Large-Vision-Language-Models papers, methods & resources.
The official repository for ArGue: Attribute-Guided Prompt Tuning For Vision-Language Models (CVPR 2024)
Official implementation for "CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text Labels" (AAAI 2023)
Coding a Multimodal (Vision) Language Model from scratch in PyTorch with full explanation: https://www.youtube.com/watch?v=vAmKB7iPkWw
【ICML 2025 Spotlight】 Official Repo for Paper ‘’HealthGPT : A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation‘’
24,523 repositories in the index in total.