Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
YuLan: An Open-Source Large Language Model
| Date | Stars |
|---|---|
| 2026-07-31 | 634 |
| 2026-08-04 | 634 |
| 2026-08-06 | 634 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align=center> <img src="https://github.com/RUC-GSAI/YuLan-Chat/blob/main/assets/YuLan-logo.jpg" width="400px"> <h1>YuLan: An Open-Source Large Language Model</h1> <a href="https://github.com/RUC-GSAI/YuLan-Chat/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-MIT-blue" alt="license"></a> <a href="https://arxiv.org/abs/2406.19853" target="_blank"><img src=https://img.shields.io/badge/arXiv-b5212f.svg?logo=arxiv></a> <a href="https://huggingface.co/yulan-team"><img alt="Static Badge" src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-blue?color=8A2BE2"></a> <a><img src="https://img.shields.io/github/stars/RUC-GSAI/YuLan-Chat"></a> </div> YuLan-Chat models are chat-based large language models, which are developed by the researchers in GSAI, Renmin University of China (YuLan, which represents Yulan Magnolia, is the campus flower of Renmin University of China). The newest version is developed by pre-training from scratch, and supervised fine-tuning via curriculum learning with high-quality English and Chinese instructions and human preference data. The model has the following technical characteristics: - Owing to large-scale pre-training on high-quality English, Chinese, and multilingual data, the language ability of the model has been improved. - Owing to the curriculum learning strategy for human alignment, the helpfulness, honesty, and harmlessness of our model have been enhanced. - To well support Chinese longer inputs and outputs, we expand the vocabulary with Chinese words and the maximum input length. It can support 4k context now. > YuLan-Chat系列模型是中国人民大学高瓴人工智能学院师生共同开发的支持聊天的大语言模型(名字"玉兰"取自中国人民大学校花)。最新版本从头完成了整个预训练过程,并采用课程学习技术基于中英文双语数据进行有监督微调,包括高质量指令和人类偏好数据。该版模型具有如下技术特点: > - 由于在大规模中英双语数据上进行了继续预训练,模型的语言能力得到提高; > - 由于采用了课程学习方法进行人类对齐训练,模型在真实场景下的有用性、诚实性与无害性得到了增强; > - 为了更好的支持中文和更长的输入输出,模型的词表及长度得到了扩充,目前可支持4k上下文。 ## News * **\[Dec. 25, 2024\]** We release **YuLan-Mini**, a highly capable 2.4B lightweight LLM using only 1T pre-training data. See more [details](https://github.com/RUC-GSAI/YuLan-Mini). * **\[July. 1, 2024\]** We release **YuLan-Base-12B**, an LLM trained from scratch, and its chat-based version **YuLan-Chat-3-12B**. We pre-train the base model on over 1.6TB tokens of English, Chinese, and multilingual data, and then perform supervised fine-tuning via curriculum learning with high-quality English and Chinese instructions and human preference data to obtain the chat model. * **\[Aug. 18, 2023\]** Our **YuLan-Chat-2-13B** achieves the 5th position of [OpenCompass](https://opencompass.org.cn/leaderboard-llm) benchmark! * **\[Aug. 02, 2023\]** We release **YuLan-LLaMA-2-13B** and **YuLan-Chat-2-13B**. Both models have been continually pre-trained on English and Chinese corpus based on LLaMA-2, and YuLan-Chat-2-13B is the chat-based LLM based on YuLan-LLaMA-2-13B, with high-quality English and Chinese instructions. * **\[Aug. 02, 2023\]** We release **YuLan-Chat-1-65B-v2**, a chat-based LLM based on LLaMA. It has been continually pre-trained on English and Chinese corpus, and then instruction-tuned with high-quality English and Chinese instructions. * **\[Jun. 08, 2023\]** We release **YuLan-Chat-1-13B-v1** and **YuLan-Chat-1-65B-v1**, and the corresponding INT-8 quantization scripts. > * **\[2024年7月1日\]** 我们发布了**YuLan-Base-12B**,一个完全从头训练的Base模型,以及其Chat化版本**YuLan-Chat-3-12B**。我们在超过1.6TB词元的中、英文和多语数据上进行了大规模预训练,得到了Base模型,然后基于高质量双语指令和人类偏好数据,使用课程学习方法进行有监督微调,最终得到了Chat化的版本。 > * **\[2023年8月2日\]** 我们发布了**YuLan-LLaMA-2-13B**和**YuLan-Chat-2-13B**两个模型,其都在LLaMA-2的基础上进行了双语继续预训练,YuLan-Chat-2-13B在YuLan-LLaMA-2-13B基础上进行了双语高质量对话指令微调。 > * **\[2023年8月2日\]** 我们发布了**YuLan-Chat-1-65B-v2**模型,其在LLaMA-65B的基础上进行了双语继续预训练,然后用高质量双语指令进行了微调。 > * **\[2023年6月8日\]** 我们发布了**YuLan-Chat-1-13B-v1**和**YuLan-Chat-1-65B-v1**两个模型,以及对应的int8量化脚本。 ## Model Zoo Due to the license limitation, for models based on LLaMA, we only provide the weight difference with the original checkpoints; for mod
Excerpt of 18,254 characters
Read on GitHubYutao ZHU · Renmin University of China · China
25
22
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:3072447c502d14e8, topic:large-language-models