Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
| Date | Stars |
|---|---|
| 2026-07-24 | 674 |
| 2026-07-25 | 674 |
| 2026-07-28 | 675 |
| 2026-07-30 | 675 |
| 2026-08-06 | 675 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<p align="center" width="100%"> <img src="assets/logo.png" alt="Stanford-Alpaca" style="width: 70%; min-width: 300px; display: block; margin: auto;"> </p> # <span style="background: linear-gradient(45deg, #667eea 0%, #764ba2 25%, #f093fb 50%, #f5576c 75%, #4facfe 100%); -webkit-background-clip: text; -webkit-text-fill-color: transparent; background-clip: text; font-weight: bold; font-size: 1.1em;">**OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM (ICLR 2026)**</span> <br /> [](https://arxiv.org/abs/2510.15870) [](https://github.com/NVlabs/OmniVinci) [](https://huggingface.co/nvidia/omnivinci) [](https://nvlabs.github.io/OmniVinci) [](https://youtu.be/w84pPuGFH4o?si=OUFhhiXeQbzil7gN) <div align="center"> </div> [Hanrong Ye*†](https://sites.google.com/site/yhrspace/home), [Chao-Han Huck Yang†](https://huckiyang.github.io/), [Arushi Goel†](https://scholar.google.com/citations?user=tj08PZcAAAAJ&hl=en), [Wei Huang†](https://aaron-weihuang.com/), [Ligeng Zhu†](https://lzhu.me/), [Yuanhang Su†](https://scholar.google.com/citations?user=n335GwUAAAAJ&hl=en), [Sean Lin†](https://www.nvidia.com/en-us/), [An-Chieh Cheng†](https://www.anjiecheng.me/), [Zhen Wan†](https://scholar.google.com/citations?user=OH_1qwMAAAAJ&hl=en), [Jinchuan Tian†](https://jctian98.github.io/), [Yuming Lou†](https://github.com/Louym), [Dong Yang†](https://scholar.google.com/citations?user=PHvliUgAAAAJ&hl=en), [Zhijian Liu](https://zhijianliu.com/), [Yukang Chen](https://yukangchen.com/), [Ambrish Dantrey](https://www.nvidia.com/en-us/), [Ehsan Jahangiri](https://www.nvidia.com/en-us/), [Sreyan Ghosh](https://sreyan88.github.io/), [Daguang Xu](https://scholar.google.com/citations?user=r_VHYHAAAAAJ&hl=en), [Ehsan Hosseini Asl](https://scholar.google.com/citations?user=I9w3ON4AAAAJ&hl=en), [Danial Mohseni Taheri](https://danialtaheri.github.io/), [Vidya Murali](https://www.linkedin.com/in/vidya-n-murali/), [Sifei Liu](https://sifeiliu.net/), [Yao Lu](https://www.linkedin.com/in/yao-jason-lu-a0291938/), [Oluwatobi Olabiyi](https://www.linkedin.com/in/oluwatobi-olabiyi-08955123/), [Yu-Chiang Frank Wang](https://scholar.google.com/citations?user=HSGvdtoAAAAJ&hl=en), [Rafael Valle](https://rafaelvalle.github.io/), [Bryan Catanzaro](https://www.linkedin.com/in/bryancatanzaro/), [Andrew Tao](https://scholar.google.com/citations?user=Wel9l1wAAAAJ&hl=en), [Song Han](https://hanlab.mit.edu/songhan), [Jan Kautz](https://jankautz.com/), [Hongxu Yin*^†](https://hongxu-yin.github.io/), [Pavlo Molchanov^](https://www.pmolchanov.com/) <span style="color: rgb(133, 184, 55);">**NVIDIA**</span> *Corresponding Author | †Core Contribution | ^Equal Advisory <p align="center" width="100%"> <img src="assets/performance.png" alt="Stanford-Alpaca" style="width: 100%; min-width: 300px; display: block; margin: auto;"> </p> Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study the design choices across model architecture and data curation. For model architecture, we present three key innovations: **(i)** OmniAlignNet for strengthening alignment between vision and audio embeddings in a shared omni-modal latent space; **(ii)** Temporal Embedding Grouping for capturing relative temporal alignment between vision and audio signals; and **(iii)** Constrained Rotary Time Embedding for encoding absolute temporal information in omni-modal embeddings. We introduce a curation and synthesis pipeline that generates 24M single-modal and omni-modal conversations. We find that modalities reinforce one anoth
Excerpt of 11,152 characters
Read on GitHub6
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:12d038beb39e98e9, topic:deep-learning
matched fp:12d038beb39e98e9, topic:large-language-models
matched fp:12d038beb39e98e9, topic:vision-language-model