Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Polyglot: Large Language Models of Well-balanced Competence in Multi-languages
| Date | Stars |
|---|---|
| 2026-07-31 | 487 |
| 2026-08-02 | 487 |
| 2026-08-06 | 487 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Polyglot: Large Language Models of Well-balanced Competence in Multi-languages ## 1. Introduction ### Why another multilingual model? Various multilingual models such as [mBERT](https://huggingface.co/bert-base-multilingual-cased), [BLOOM](https://huggingface.co/bigscience/bloom), and [XGLM](https://arxiv.org/abs/2112.10668) have been released. Therefore, someone might ask, "why do we need to make multilingual models again?" Before answering the question, we would like to ask, "Why do people around the world make monolingual models in their language even though there are already many multilingual models?" We would like to point out there is a dissatisfaction with the non-English language performance of the current multilingual models as one of the most significant reason. So we want to make multilingual models with higher non-English language performance. This is the reason we need to make multilingual models again and why we name them ['Polyglot'](https://www.spanish.academy/blog/what-is-the-difference-between-a-polyglot-and-a-multilingual-person/). ## 2. Projects ### 1) Polyglot-Ko When we started our research, we have already had 1.2TB of Korean data collected by [TUNiB](https://tunib.ai/). Before we collected a large amount of multilingual data, we decided to try Korean modeling with the dataset we already had. This Korean model can be used for performance comparison with the multilingual models, and this model itself would help many Korean companies and researchers. | Size | Training Status | Model Card | Model Checkpoints | |:----:|:------------------------------------------------------------------------------------------:|:---------------------------------------------------------------:|:-------------------------------------------------------------------------:| | 1.3B | [Finished](https://wandb.ai/eleutherai/polyglot-ko/groups/polyglot-ko-1.3B) | [Available](https://huggingface.co/EleutherAI/polyglot-ko-1.3b) | [Available](https://huggingface.co/EleutherAI/polyglot-ko-1.3b/tree/main) | | 3.8B | [Finished](https://wandb.ai/eleutherai/polyglot-ko/groups/polyglot-ko-3.8B) | [Available](https://huggingface.co/EleutherAI/polyglot-ko-3.8b) | [Available](https://huggingface.co/EleutherAI/polyglot-ko-3.8b/tree/main) | | 5.8B | [Finished](https://wandb.ai/eleutherai/polyglot-ko/groups/polyglot-ko-5.8B) | [Available](https://huggingface.co/EleutherAI/polyglot-ko-5.8b) | [Available](https://huggingface.co/EleutherAI/polyglot-ko-5.8b/tree/main) | |12.8B | [Finished](https://wandb.ai/eleutherai-oslo/polyglot-ko-12_8b) | [Available](https://huggingface.co/EleutherAI/polyglot-ko-12.8b) | [Available](https://huggingface.co/EleutherAI/polyglot-ko-12.8b/tree/main) - Technical report: https://arxiv.org/abs/2306.02254 - Contributors: [Hyunwoong Ko](https://github.com/hyunwoongko), [Kichang Yang](https://github.com/jason9693), [Minho Ryu](https://github.com/bzantium), [Taekyoon Choi](https://github.com/Taekyoon), [Seungmu Yang](https://github.com/Ronalmoo), [jiwung Hyun](https://github.com/kabbi159), [Sungho Park](https://github.com/naem1023) 💡 We are collaborating with KoAlpaca team which is creating a series of Korean instruct fine-tuned models. As a result, we were able to release the Koalapca-Polyglot models. Please refer to [here](https://github.com/Beomi/KoAlpaca) to see more details. ### 2) Japanese StableLM We co-worked with StabilityAI Japan to create open source Japanese language models. We've mainly contributed to dataset collection part for this project. | Size | Training Status
Excerpt of 8,184 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:000394fa2167634d, llm:Repository description: 'Polyglot: Large Language Models of Well-balanced Competence in Multi-languages' (EleutherAI project), indicates it's a foundation multilingual LLM.
matched fp:000394fa2167634d, llm:Repository description: 'Polyglot: Large Language Models of Well-balanced Competence in Multi-languages' (EleutherAI project), indicates it's a foundation multilingual LLM.
matched fp:000394fa2167634d, llm:Repository description: 'Polyglot: Large Language Models of Well-balanced Competence in Multi-languages' (EleutherAI project), indicates it's a foundation multilingual LLM.