facebookresearch/lingua
quality grade C, 58 out of 100Meta Lingua: a lean, efficient, and easy-to-hack codebase to research LLMs.
- stars
- 4.8k
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Core deep-learning frameworks and libraries for pretraining and distributed training.
Signals: deep-learning, neural-network, pytorch, tensorflow, jax, distributed-training, training, deepspeed
2,694 results
Meta Lingua: a lean, efficient, and easy-to-hack codebase to research LLMs.
Seamlessly integrate LLMs into scikit-learn.
LLM training in simple, raw C/CUDA
IDDM (Industrial, landscape, animate, latent diffusion), support LDM, DDPM, DDIM, PLMS, webui and distributed training. Pytorch实现扩散模型,生成模型,分布式训练
Mistral: A strong, northwesterly wind: Framework for transparent and accessible large-scale language model training, built with Hugging Face 🤗 Transformers.
The official implementation of “Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training”
From scratch implementation of a sparse mixture of experts language model inspired by Andrej Karpathy's makemore :)
Implementation of the specific Transformer architecture from PaLM - Scaling Language Modeling with Pathways
Implementation of the training framework proposed in Self-Rewarding Language Model, from MetaAI
Pretraining and inference code for a large-scale depth-recurrent language model
An implementation of local windowed attention for language modeling
Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch
Saprot: Protein Language Model with Structural Alphabet (AA+3Di)
MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining
PhoBERT: Pre-trained language models for Vietnamese (EMNLP-2020 Findings)
Implementation of Toolformer, Language Models That Can Use Tools, by MetaAI
Ongoing research training transformer language models at scale, including: BERT & GPT-2
Guide to using pre-trained large language models of source code
Empower Sequence Labeling with Task-Aware Language Model
博客配套视频链接: https://space.bilibili.com/383551518?spm_id_from=333.1007.0.0 b 站直接看 配套 github 链接:https://github.com/nickchen121/Pre-training-language-model 配套博客链接:https://www.cnblogs.com/nickchen121/p/15105048.html
Implementation of π₀, the robotic foundation model architecture proposed by Physical Intelligence
[ICLR 2025 Spotlight] Official implementation of "Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts"
Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
[COLM 2025] Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
24,524 repositories in the index in total.