Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
| Date | Stars |
|---|---|
| 2026-07-31 | 420 |
| 2026-08-03 | 421 |
| 2026-08-06 | 421 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align="right">
<a href="README_CN.md">简体中文</a> | <b>English</b>
</div>
<h1 align="center">Qianfan-VL Series</h1>
<p align="center">
<strong>Domain-Enhanced Vision-Language Models for Enterprise</strong><br>
<strong>3B to 70B Parameters | OCR Specialist</strong><br>
<strong>Document Understanding & OCR Enhancement</strong><br>
<strong>Chain-of-Thought Support</strong>
</p>
<div align="center">
📚 **[Cookbook](https://github.com/baidubce/qianfan-models-cookbook)** |
</div>
---
## Introduction
**Qianfan-VL Series** is a family of domain-enhanced vision-language models built for enterprise users who need production-grade visual understanding. The series delivers deep optimization for high-frequency industrial deployment scenarios — from document parsing and OCR to complex visual reasoning — while retaining strong general-purpose multimodal capabilities. Whether deployed at the edge or in the cloud, Qianfan-VL Series provides the visual intelligence backbone that enterprise applications demand.
## 🔥 News
- **[2026/03/18] Qianfan-OCR Released!** — Introducing **Qianfan-OCR**, a 4B end-to-end model that unifies document parsing, layout analysis, table extraction, formula recognition, chart understanding & key information extraction in one single model. No more chaining detection → recognition → LLM. Key innovation: **Layout-as-Thought** — an optional ⟨think⟩ phase where the model reasons about bounding boxes, element types & reading order before generating output. Think "CoT for document layout."
- 🏆 **OmniDocBench v1.5:** 93.12 (#1 end-to-end)
- 🏆 **OCRBench:** 880 (#1 overall, all models)
- 🏆 **KIE avg:** 87.9 (outperforms Gemini-3.1-Pro & Qwen3-VL-235B)
- ⚡ **1.024 pages/sec** on a single A100 (W8A8)
- 🌍 Supports **192 languages** across Latin, Cyrillic, Arabic, South Asian, Southeast Asian & CJK scripts
- Trained on 1,024 Kunlun P800 chips processing 2.85T tokens across 4 stages
- 📄 [Paper](https://arxiv.org/abs/2603.13398) | 🤗 [HuggingFace Collection](https://huggingface.co/collections/baidu/qianfan-vl) | 💻 [GitHub](https://github.com/baidubce/Qianfan-OCR) | [ModelScope](https://modelscope.cn/models/baidu-qianfan/Qianfan-OCR)
- **[2025/09/22] Qianfan-VL 1.0 Released!** — Launching the Qianfan-VL model series (3B / 8B / 70B) with multi-size coverage from edge to cloud, enhanced OCR & document understanding, and Chain-of-Thought reasoning for 8B and 70B variants. 📄 [Tech Report](https://arxiv.org/abs/2603.13398)
## Qianfan-VL 1.0 Key Features
### 🚀 Multi-Size Models
Provides 3B, 8B, and 70B model variants to meet different scenario requirements from edge to cloud
### 📝 OCR & Document Understanding Enhancement
- **Full-scenario OCR recognition**: Supports handwriting, printed text, scene text, formulas, and more
- **Complex layout understanding**: Table parsing, chart understanding, document structuring capabilities
- **Multi-language support**: Chinese, English, and multilingual document processing
### 🧠 Chain-of-Thought Capability
8B and 70B models support Chain-of-Thought capability, demonstrating excellent performance in complex scenarios like mathematics and reasoning computation, applicable to teaching assistance, photo problem-solving, automatic grading, and more
## Model Specifications
| Model Name | Parameters | Context Length | CoT Support | Application Scenarios | Model Download |
|---------|--------|-----------|---------|----------|---------|
| **Qianfan-VL-3B** | 3B | 32k | ❌ | Edge real-time scenarios, OCR text recognition | 🤗 **[HuggingFace](https://huggingface.co/baidu/Qianfan-VL-3B)** / 🤖 **[ModelScope](https://modelscope.cn/models/baidu-qianfan/Qianfan-VL-3B)** |
| **Qianfan-VL-8B** | 8B | 32k | ✅ | Server-side general scenarios, fine-tuning optimization | 🤗 **[HuggingFace](https://huggingface.co/baidu/Qianfan-VL-8B)** / 🤖 **[ModelScope](https://modelscope.cn/models/baidu-qianfan/Qianfan-VL-8B)** |
| **Qianfan-VL-70B** | 70B | 32k | ✅ | OffliExcerpt of 14,838 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:f7dc8ff1627dd432, desc:vision-language