Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Recent LLM-based CV and related works. Welcome to comment/contribute!
| Date | Stars |
|---|---|
| 2026-07-31 | 871 |
| 2026-08-06 | 871 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# LLM-in-Vision Recent LLM (Large Language Models)-based CV and multi-modal works. Welcome to comment/contribute! ### 2025.2 - (CVPR 2025.2) Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection, [[Paper]](https://arxiv.org/abs/2412.04455), [[Project]](https://zhoues.github.io/Code-as-Monitor/) ### 2024.6 <!-- - (arXiv 2024.6) , [[Paper]](), [[Project]]()--> - (arXiv 2024.6) Bootstrap3D: Improving **3D Content Creation** with Synthetic Data, [[Paper]](https://arxiv.org/pdf/2406.00093), [[Project]](https://sunzey.github.io/Bootstrap3D/) ### 2024.5 <!-- - (arXiv 2024.5) , [[Paper]](), [[Project]]()--> - (arXiv 2024.5) VideoOFA: Two-Stage Pre-Training for **Video-to-Text** Generation, [[Paper]](https://arxiv.org/pdf/2305.03204) - (arXiv 2024.5) Grounded 3D-LLM with Referent Tokens, [[Paper]](https://arxiv.org/pdf/2405.10370), [[Project]](https://groundedscenellm.github.io/grounded_3d-llm.github.io) - (arXiv 2024.5) **Self-supervised Pre-training** for Transferable Multi-modal Perception, [[Paper]](https://arxiv.org/pdf/2405.17942v1) - (arXiv 2024.5) Multi-modal Generation via Cross-Modal **In-Context Learning**, [[Paper]](https://arxiv.org/pdf/2405.18304v1) - (arXiv 2024.5) RoboCasa: Large-Scale **Simulation** of Everyday Tasks for Generalist **Robots**, [[Paper]](https://robocasa.ai/assets/robocasa_rss24.pdf), [[Project]](https://robocasa.ai/) - (arXiv 2024.5) Unveiling the Tapestry of **Consistency** in Large Vision-Language Models, [[Paper]](https://arxiv.org/pdf/2405.14156) - (arXiv 2024.5) Dense **Connector** for MLLMs, [[Paper]](https://arxiv.org/pdf/2405.13800), [[Project]](https://github.com/HJYao00/DenseConnector) - (arXiv 2024.5) Adapting Multi-modal Large Language Model to Concept Drift in the **Long-tailed** Open World, [[Paper]](https://arxiv.org/pdf/2405.13459), [[Project]](https://github.com/Anonymous0Knight/ConceptDriftMLLMs) - (arXiv 2024.5) VTG-LLM: INTEGRATING TIMESTAMP KNOWLEDGE INTO VIDEO LLMS FOR ENHANCED **VIDEO** TEMPORAL **GROUNDING**, [[Paper]](https://arxiv.org/pdf/2405.13382), [[Project]](https://github.com/gyxxyg/VTG-LLM) - (arXiv 2024.5) **Calibrated** Self-Rewarding Vision Language Models, [[Paper]](https://arxiv.org/pdf/2405.14622), [[Project]](https://github.com/YiyangZhou/CSR) - (arXiv 2024.5) From Text to Pixel: Advancing **Long-Context** Understanding in MLLMs, [[Paper]](https://arxiv.org/pdf/2405.14213), [[Project]](https://github.com/YujieLu10/Seeker) - (arXiv 2024.5) **Explaining** Multi-modal Large Language Models by Analyzing their Vision Perception, [[Paper]](https://arxiv.org/pdf/2405.14612) - (arXiv 2024.5) Octopi: Object Property Reasoning with Large **Tactile-Language** Models, [[Paper]](https://arxiv.org/pdf/2405.02794), [[Project]](https://github.com/clear-nus/octopi) - (arXiv 2024.5) Auto-Encoding Morph-**Tokens** for **Multimodal** LLM, [[Paper]](https://arxiv.org/pdf/2405.01926), [[Project]](https://github.com/DCDmllm/MorphTokens) - (arXiv 2024.5) What matters when building **vision-language** models? [[Paper]](https://arxiv.org/pdf/2405.02246) ### 2024.4 <!-- - (arXiv 2024.4) , [[Paper]](), [[Project]]()--> - (arXiv 2024.4) VITRON: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing, [[Paper]](http://haofei.vip/downloads/papers/Skywork_Vitron_2024.pdf), [[Project]](https://vitron-llm.github.io/) - (arXiv 2024.4) GROUNDHOG: Grounding Large Language Models to Holistic **Segmentation**, [[Paper]](https://arxiv.org/pdf/2402.16846), [[Project]](https://groundhog-mllm.github.io/) - (arXiv 2024.4) **Hallucination** of Multimodal Large Language Models: A **Survey**, [[Paper]](https://arxiv.org/pdf/2404.18930), [[Project]](https://github.com/showlab/Awesome-MLLM-Hallucination) - (arXiv 2024.4) PLLaVA: Parameter-free LLaVA Extension from Images to Videos for **Video Dense Captioning**, [[Paper]](https://arxiv.org/pdf/2404.16994), [[Project]](https://github.com/magic-research/PLLaV
Excerpt of 137,244 characters
Read on GitHub304
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:79e7a6b4e3272bfb, llm:Repository description: 'Recent LLM-based CV and related works. Welcome to comment/contribute!' — indicates a collection of LLM-based computer vision works (vision + LLMs).
matched fp:79e7a6b4e3272bfb, llm:Repository description: 'Recent LLM-based CV and related works. Welcome to comment/contribute!' — indicates a collection of LLM-based computer vision works (vision + LLMs).
matched fp:79e7a6b4e3272bfb, llm:Repository description: 'Recent LLM-based CV and related works. Welcome to comment/contribute!' — indicates a collection of LLM-based computer vision works (vision + LLMs).