Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
[Nature Communications] The official codes for "Towards Building Multilingual Language Model for Medicine"
| Date | Stars |
|---|---|
| 2026-07-31 | 284 |
| 2026-08-05 | 284 |
| 2026-08-06 | 284 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# MMedLM <div align="center"> <img src="./assets/logo.png" width="200"/> <div align="center"></div> </div> [Nature Communications] The official codes for "Towards Building Multilingual Language Model for Medicine". [Paper (Arxiv version)](https://arxiv.org/abs/2402.13963) [Paper (Nature Communications)](https://www.nature.com/articles/s41467-024-52417-z) [Leaderboard](https://henrychur.github.io/MultilingualMedQA/) Models: [MMedLM-7B](https://huggingface.co/Henrychur/MMedLM), [MMedLM 2-7B](https://huggingface.co/Henrychur/MMedLM2), [MMedLM 2-1.8B](https://huggingface.co/Henrychur/MMedLM2-1.8B), [MMed-Llama 3-8B](https://huggingface.co/Henrychur/MMed-Llama-3-8B), [MMed-Llama3-8B-EnIns](https://huggingface.co/Henrychur/MMed-Llama-3-8B-EnIns), [MMed-Llama3.1-70B](https://huggingface.co/Henrychur/MMed-Llama3.1-70B) Datasets: [MMedC](https://huggingface.co/datasets/Henrychur/MMedC),[MMedBench](https://huggingface.co/datasets/Henrychur/MMedBench) ## Introduction In this paper, we aim to develop an open-source, multilingual language model for medicine. In general, we present the contribution from the following aspects: 1. **Corpus dataset.** For multilingual medical-specific adaptation, we construct a new multilingual medical corpus, that contains approximately 25.5B tokens encompassing 6 main languages, termed as MMedC, that enables auto-regressive training for existing general LLMs. 2. **Benchmark.** To monitor the development of multilingual LLMs in medicine, we propose a new multilingual medical multi-choice question-answering benchmark with rationale, termed as MMedBench. 3. **Model Evaluation.** We have assessed a number of popular LLMs on our benchmark, along with those further auto-regressive trained on MMedC, as a result, our final model, termed as MMedLM 2, with only 7B parameters, achieves superior performance compared to all other open-source models, even rivaling GPT-4 on MMedBench.  ## News [2025.3.7]  We release [MMed-Llama3.1-70B](https://huggingface.co/Henrychur/MMed-Llama3.1-70B), a new 70B **medical multilingual medical foundation model** based on LLama3.1-70B. With an aoto-regressive continues training on the newest MMedC, this model(80.51) achieves superior performance compared to GPT-4(74.27) and supports 8 main languages. [2024.9.27] Our paper has been accepted by Nature Communications! [2024.5.24] We release [MMed-Llama 3-8B](https://huggingface.co/Henrychur/MMed-Llama-3-8B) and [MMed-Llama3-8B-EnIns](https://huggingface.co/Henrychur/MMed-Llama-3-8B-EnIns). MMed-Llama 3 is based on Llama 3 and futher pretrained on MMedC, and MMed-Llama 3 EnIns is a fine-tuned version with additional English instructions (from PMC-LLaMA). [2024.3.1] We release [MMedLM 2-1.8B](https://huggingface.co/Henrychur/MMedLM2-1.8B), a 1.8B light-weight model based on InternLM 2-1.8B. With an auto-regressive continues training on MMedC, MMedLM 2-1.8B can exceed the performance of most 7B models, including InternLM and LLaMA 2. [2024.2.21] Our leaderboard web can be found [here](https://henrychur.github.io/MultilingualMedQA/). We look forward to more superior efforts in multilingual medical LLMs!. [2024.2.21] Our pre-print paper is released ArXiv. Dive into our findings [here](https://arxiv.org/abs/2402.13963). [2024.2.20] We release [MMedLM](https://huggingface.co/Henrychur/MMedLM) and [MMedLM 2](https://huggingface.co/Henrychur/MMedLM2). With an auto-regressive continues training on MMedC, these models achieves superior performance compared to all other open-source models, even rivaling GPT-4 on MMedBench. [2023.2.20] We release [MMedC](https://huggingface.co/datasets/Henrychur/MMedC), a multilingual medical corpus containing 25.5B tokens. [2023.2.20] We release [MMedBench](https://huggingface.co/datasets/Henrychur/MMedBench), a new multilingual medical multi-choice question-answering benchmark with rationale. ## Usage ### Environment In our experiments, we used
Excerpt of 14,301 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:ae988c3ca7abe290, llm:Repository description: 'The official codes for "Towards Building Multilingual Language Model for Medicine"' (Nature Communications). Language: Python.
matched fp:ae988c3ca7abe290, llm:Repository description: 'The official codes for "Towards Building Multilingual Language Model for Medicine"' (Nature Communications). Language: Python.
matched fp:ae988c3ca7abe290, llm:Repository description: 'The official codes for "Towards Building Multilingual Language Model for Medicine"' (Nature Communications). Language: Python.