Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[CVPR2024] The code for "Osprey: Pixel Understanding with Visual Instruction Tuning"
| Date | Stars |
|---|---|
| 2026-07-31 | 843 |
| 2026-08-06 | 843 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<p align="center" width="100%"> <img src="assets/osprey.png" width="90%"> </p> <div align=center>  [](https://arxiv.org/pdf/2312.10032.pdf) [](https://huggingface.co/datasets/AntGroup-MI/Osprey-724K) [](https://youtu.be/YsxqHBBnDfk) [](http://111.0.123.204:8000/) </div> <div align=center> Demo username & password: <b>osprey</b> </div> --- <div align=center> <img src="./assets/qmsht.gif" /> <br> A part of <i>Along the River During the Qingming Festival</i> (清明上河图) <br> <img src="./assets/qyqx.gif" /> <br> <i>Spirited Away</i> (千与千寻) <br> </div> <details open><summary>💡 Some of our other multimodal-LLM projects may interest you ✨. </summary><p> <!-- may --> > [**VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM**](https://arxiv.org/abs/2501.00599) <br> > Yuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng, Boqiang Zhang, Long Li, Xin Li, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, Lidong Bing <br> [](https://github.com/DAMO-NLP-SG/VideoRefer) [](https://github.com/DAMO-NLP-SG/VideoRefer) [](https://arxiv.org/abs/2501.00599) <br> > [**TokenPacker: Efficient Visual Projector for Multimodal LLM**](https://arxiv.org/abs/2407.02392) <br> > Wentong Li*, Yuqian Yuan*, Jian Liu, Dongqi Tang, Song Wang, Jianke Zhu, Lei Zhang <br> [](https://github.com/CircleRadon/TokenPacker) [](https://github.com/CircleRadon/TokenPacker) [](https://arxiv.org/abs/2407.02392) <br> ## Updates 📌 [2025/4/22]🔥 Our defined metrics (Sem. Sim. & Sem. IoU) on Referring Object Classification have been adopted in [Describe Anything Model](https://arxiv.org/pdf/2504.16072) (NVIDIA & UC Berkeley). [2025/2/27]🔥 Our new work, [VideoRefer Suite](https://github.com/DAMO-NLP-SG/VideoRefer), has been accept to CVPR2025! This project focuses on video referring. [2024/11/27]🔥 Our defined metrics (Sem. Sim. & Sem. IoU) on Referring Object Classification have been adopted in [ChatRex](https://arxiv.org/abs/2411.18363) (IDEA). [2024/3/29]🔥 We released [Osprey-Chat](https://huggingface.co/sunshine-lwt/Osprey-Chat-7b/tree/main) model, which exhibits better conversation and image-level understanding&reasoning capabilities. [2024/2/27]🔥 Osprey has been accepted to CVPR2024! [2024/1/15]🔥 We released the [evaluation](./osprey/eval/README.md) code. [2023/12/29]🔥 We released the training code and [Osprey-724K](https://huggingface.co/datasets/AntGroup-MI/Osprey-724K) dataset. [2023/12/18]🔥 We released the code, [osprey-7b model](https://huggingface.co/sunshine-lwt/Osprey-7b/tree/main) and [online demo](http://111.0.123.204:8000/) for Osprey. ## What is Osprey 👀 Osprey is a mask-text instruction tuning approach that extends MLLMs by incorporating pixel-wise mask regions into language instructions, enabling **fine-grained visual understanding**. Based on input mask region, Osprey generate the semantic descriptions including **short description** and **detailed description**. Our Osprey can seamlessly integrate with [SAM](https://github.com/facebookresearch/segment-anything) in point-prompt, box-prompt and segmentation everything modes to generate the semantics asso
Excerpt of 9,243 characters
Read on GitHubYuqian Yuan · Zhejiang University · China
20
15
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:41bac70fa3914698, desc:instruction tuning