Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.
| Date | Stars |
|---|---|
| 2026-07-31 | 760 |
| 2026-08-04 | 760 |
| 2026-08-06 | 760 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Multimodal & Large Language Models **Note:** This paper list is only used to record papers I read in the daily arxiv for personal needs. I only subscribe to and cover the following subjects: Artificial Intelligence (cs.AI), Computation and Language (cs.CL), Computer Vision and Pattern Recognition (cs.CV), and Machine Learning (cs.LG). If you find I missed some important and exciting work, it would be super helpful to let me know. Thanks! **Update:** Starting from June 2024, I'm focusing on reading and recording papers that I believe offer unique insights and substantial contributions to the field. ## Table of Contents - [Survey](#survey) - [Position Paper](#position-paper) - [Structure](#structure) - [Planning](#planning) - [Reasoning](#reasoning) - [Generation](#generation) - [Representation Learning](#representation-learning) - [LLM Analysis](#llm-analysis) - [LLM Safety](#llm-safety) - [LLM Evaluation](#llm-evaluation) - [LLM Reasoning](#llm-reasoning) - [LLM Application](#llm-application) - [LLM with Memory](#llm-with-memory) - [LLM with Human](#llm-with-human) - [Inference-time Scaling (via RL)](#inference-time-scaling) - [Long-Context LLM](#long-context-llm) - [LLM Foundation](#llm-foundation) - [Scaling Law](#scaling-law) - [LLM Data Engineering](#llm-data-engineering) - [VLM Data Engineering](#vlm-data-engineering) - [Alignment](#alignment) - [Scalable Oversight&SuperAlignment](#scalable-oversight-&-superalignment) - [RL Foundation](#rl-foundation) - [Beyond Bandit](#beyond-bandit) - [Agent](#agent) - [DeepResearch](#deepresearch) - [SWE-Agent](#swe-agent) - [Evolution](#evolution) - [Interaction](#interaction) - [Critique Modeling](#critic-modeling) - [MoE/Specialized](#moe/specialized) - [Vision-Language Foundation Model](#vision-language-foundation-model) - [Vision-Language Model Analysis & Evaluation](#vision-language-model-analysis&evaluation) - [Vision-Language Model Application](#vision-language-model-application) - [Multimodal Foundation Model](#multimodal-foundation-model) - [Image Generation](#image-generation) - [Diffusion](#diffusion) - [Document Understanding](#document-understanding) - [Tool Learning](#tool-learning) - [Instruction Tuning](#instruction-tuning) - [Incontext Learning](#incontext-learning) - [Learning from Feedback](#learning-from-feedback) - [Reward Modeling](#reward-modeling) - [Video Foundation Model](#video-foundation-model) - [Key Frame Detection](#key-frame-detection) - [Pretraining](#pretraining) - [Vision Model](#vision-model) - [Adaptation of Foundation Model](#adaptation-of-foundation-model) - [Prompting](#prompting) - [Efficiency](#efficiency) - [Analysis](#analysis) - [Grounding](#grounding) - [VQA Task](#vqa-task) - [VQA Dataset](#vqa-dataset) - [Social Good](#social-good) - [Application](#application) - [Benchmark & Evaluation](#benchmark-&-evaluation) - [Dataset](#dataset) - [Robustness](#robustness) - [Hallucination&Factuality](#hallucination&factuality) - [Cognitive NeuronScience & Machine Learning](#cognitive-neuronscience-&-machine-learning) - [Theory of Mind](#theory-of-mind) - [Cognitive NeuronScience](#cognitive-neuronscience) - [World Model](#world-model) - [Resource](#resource) ## Survey - **Multimodal Learning with Transformers: A Survey;** Peng Xu, Xiatian Zhu, David A. Clifton - **Multimodal Machine Learning: A Survey and Taxonomy;** Tadas Baltrusaitis, Chaitanya Ahuja, Louis-Philippe Morency; Introduce 4 challenges for multi-modal learning, including representation, translation, alignment, fusion, and co-learning. - **FOUNDATIONS & RECENT TRENDS IN MULTIMODAL MACHINE LEARNING: PRINCIPLES, CHALLENGES, & OPEN QUESTIONS;** Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency - **Multimodal research in vision and language: A review of current and emerging trends;** Shagun Uppal et al; - **Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods;** Aditya Mogadala et al - **Challenges and Prospects in Vis
Excerpt of 224,661 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:160afbd1c89ee7c1, topic:multimodal, name:multimodal, desc:multimodal
matched fp:160afbd1c89ee7c1, topic:large-language-models