Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A curated list of Multimodal Related Research.
| Date | Stars |
|---|---|
| 2026-07-24 | 1393 |
| 2026-07-25 | 1393 |
| 2026-07-28 | 1393 |
| 2026-07-30 | 1393 |
| 2026-08-06 | 1393 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome Multimodal Research [](https://github.com/sindresorhus/awesome)    > This repo is reorganized from [Paul Liang's repo: Reading List for Topics in Multimodal Machine Learning](https://github.com/pliang279/awesome-multimodal-ml), feel free to raise pull requests! ## News **[03/2023] OpenAI:** *[ChatGPT plugins](https://platform.openai.com/docs/plugins/introduction) are tools designed specifically for language models with safety as a core principle, and help ChatGPT access up-to-date information, run computations, or use third-party services. https://openai.com/blog/chatgpt-plugins* > *"We’re also hosting two plugins ourselves, a [web browser](https://openai.com/blog/chatgpt-plugins#browsing) and [code interpreter](https://openai.com/blog/chatgpt-plugins#code-interpreter). We’ve also open-sourced the code for a knowledge base [retrieval plugin](https://github.com/openai/chatgpt-retrieval-plugin), to be self-hosted by any developer with information with which they’d like to augment ChatGPT."* **[03/2023] Google Research:** *[Bard](https://blog.google/technology/ai/try-bard/) is an early experiment that lets you collaborate with generative AI, powered by a research large language model (LLM), specifically a lightweight and optimized version of LaMDA. https://bard.google.com/* **[03/2023] OpenAI:** *[GPT-4](https://cdn.openai.com/papers/gpt-4.pdf) is a large multimodal model (accepting image and text inputs, emitting text outputs) that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks. https://openai.com/research/gpt-4* **[03/2023] Google Research:** *[PaLM-E](https://palm-e.github.io/assets/palm-e.pdf) is a new generalist robotics model that overcomes these issues by transferring knowledge from varied visual and language domains to a robotics system. https://ai.googleblog.com/2023/03/palm-e-embodied-multimodal-language.html* **[03/2023] OpenAI:** *[ChatGPT and Whisper APIs](https://openai.com/blog/introducing-chatgpt-and-whisper-apis), developers can now integrate ChatGPT and Whisper models into their apps and products through API. https://openai.com/blog/introducing-chatgpt-and-whisper-apis* **[02/2023] MSR:** *[Kosmos-1](https://arxiv.org/pdf/2302.14045.pdf) is a multimodal large language model (MLLM) that is capable of perceiving multimodal input, following instructions, and performing in-context learning for not only language tasks but also multimodal tasks. https://github.com/microsoft/unilm#llm--mllm-multimodal-llm* **[01/2023] Google Research:** *[2022 & beyond: Language, vision and generative models](https://ai.googleblog.com/2023/01/google-research-2022-beyond-language.html), a post of a series in which researchers across Google will highlight some exciting progress in 2022 and present the vision for 2023 and beyond. https://ai.googleblog.com/2023/01/google-research-2022-beyond-language.html* **[11/2022] OpenAI:** *[ChatGPT](https://chat.openai.com/chat) is a sibling model to [InstructGPT](https://openai.com/research/instruction-following), which is trained to follow an instruction in a prompt and provide a detailed response. https://openai.com/blog/chatgpt* **[08/2022] MSR:** *[Multimodal Pretraining](https://github.com/microsoft/unilm#multimodal-x--language): [BEiT-3](https://arxiv.org/abs/2208.10442) is a general-purpose multimodal foundation model, which achieves state-of-the-art transfer performance on both vision and vision-language tasks. https://github.com/microsoft/unilm/tree/master/beit* **[04/2022] OpenAI:** *[DALL·E 2](https://arxiv.org/pdf/2204.06125.pdf) is a new AI system that can creat
Excerpt of 12,918 characters
Read on GitHubFeiyang(Vance) Chen · Creatify AI · United States
202
Mario · Universität Stuttgart · Germany
2
Jun Wang · Salesforce Research · United States
1
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:9f0dad8103c3e250, topic:multimodal, name:multimodal, desc:multimodal
matched fp:9f0dad8103c3e250, topic:awesome, desc:curated list, readme:reading list