Top AI Repos β open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Baichuan-Omni: Towards Capable Open-source Omni-modal LLM π
| Date | Stars |
|---|---|
| 2026-07-31 | 273 |
| 2026-08-06 | 273 |
Today
β stars today
This week
β stars this week
This month
β stars this month
Momentum
0.0
growth rate 0.00%/day
<div align=center><img src="assets/logo.jpg" height="100%" width="90%"/></div>
<h2 align="center"> <a href="https://arxiv.org/abs/2410.08565">Baichuan-Omni Technical Report</a></h2>
<h5 align="center"> If our project helps you, please give us a star β and cite our <a href="#citation">paper</a>!</h2>
<h5 align="center">
[](https://huggingface.co/papers/2410.08565)
[](https://github.com/westlake-baichuan-mllm/bc-omni)
[](https://arxiv.org/abs/2410.08565)
[](https://hits.seeyoufarm.com)
<!-- [](https://github.com/DAMO-NLP-SG/VideoLLaMA2/blob/main/LICENSE) -->
[](https://github.com/westlake-baichuan-mllm/bc-omni/issues?q=is%3Aopen+is%3Aissue)
[](https://github.com/westlake-baichuan-mllm/bc-omni/issues?q=is%3Aissue+is%3Aclosed) <br>
<div style="display: flex; justify-content: center;">
<img src="assets/lidar.png" style="width: 48%; height: auto;"/>
<img src="assets/lidar_avg.png" style="width: 48%; height: auto;"/>
</div>
## News
- **[2025/01/26]** π₯ The code and model weights for both the base and alignment versions of **Baichuan-Omni-1.5** are now publicly available. Long wait! See [here](https://github.com/baichuan-inc/Baichuan-Omni-1.5)!
- **[2025/01/25]** π₯ We are excited to present **Baichuan-Omni-1.5**, an enhanced version that is capable of processing four modalitiesβtext, image, audio, and videoβwhile supporting both text and audio outputs in an end-to-end manner.
- **[2024/10/11]** π We have released the technical report of **Baichuan-Omni**. See [here](https://arxiv.org/abs/2410.08565)!
## Introduction
The salient multimodal capabilities and interactive experience of GPT-4o highlight its critical role in practical applications, yet it lacks a high-performing open-source counterpart.
In this paper, we introduce Baichuan-Omni, the first high-performing open-source Multimodal Large Language Model (MLLM) adept at concurrently processing and analyzing modalities of image, video, audio, and text, while delivering an advanced multimodal interactive experience. We propose an effective multimodal training schema starting with 7B model and proceeding through two stages of multimodal alignment and multitask fine-tuning across audio, image, video, and text modal.
This approach equips the language model with the ability to handle visual and audio data effectively. Demonstrating strong performance across various omni-modal and multimodal benchmarks, we aim for this contribution to serve as a competitive baseline for the open-source community in advancing multimodal understanding and real-time interaction.
## Architecture & Training Schema
The training pipeline is divided into two phases:
<div align=center><img src="assets/pipeline.svg" height="100%" width="90%"/></div>
**Phase 1: Multimodal Alignment Pretraining.** The pre-training and alignment processes consist of Image-Language, Video-Language, and Audio-Language branches.
- The Image-Language branch utilizes a visual encoder to process images and undergoes training in three stages, focusing on image captioning, visual question answering tasks, and further enhancing alignment with the large language model (LLM).
- The Video-Language branch builds on the Image-Language branch, using the same visual encoder and video pExcerpt of 7,331 characters
Read on GitHubWould you bet a product on this? Bounded 0β100 and slow moving.
matched fp:ad6c0b16ea8e2168, llm:Repository title and description: 'Baichuan-Omni: Towards Capable Open-source Omni-modal LLM' β indicates an open-source multimodal large language model (LLM).
matched fp:ad6c0b16ea8e2168, llm:Repository title and description: 'Baichuan-Omni: Towards Capable Open-source Omni-modal LLM' β indicates an open-source multimodal large language model (LLM).
matched fp:ad6c0b16ea8e2168, llm:Repository title and description: 'Baichuan-Omni: Towards Capable Open-source Omni-modal LLM' β indicates an open-source multimodal large language model (LLM).