facebookresearch/lingua
quality grade C, 58 out of 100Meta Lingua: a lean, efficient, and easy-to-hack codebase to research LLMs.
- stars
- 4.8k
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Core deep-learning frameworks and libraries for pretraining and distributed training.
Signals: deep-learning, neural-network, pytorch, tensorflow, jax, distributed-training, training, deepspeed
2,672 results
Meta Lingua: a lean, efficient, and easy-to-hack codebase to research LLMs.
LLM training in simple, raw C/CUDA
IDDM (Industrial, landscape, animate, latent diffusion), support LDM, DDPM, DDIM, PLMS, webui and distributed training. Pytorch实现扩散模型,生成模型,分布式训练
Mistral: A strong, northwesterly wind: Framework for transparent and accessible large-scale language model training, built with Hugging Face 🤗 Transformers.
The official implementation of “Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training”
From scratch implementation of a sparse mixture of experts language model inspired by Andrej Karpathy's makemore :)
Implementation of the specific Transformer architecture from PaLM - Scaling Language Modeling with Pathways
Implementation of the training framework proposed in Self-Rewarding Language Model, from MetaAI
Pretraining and inference code for a large-scale depth-recurrent language model
An implementation of local windowed attention for language modeling
Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch
Saprot: Protein Language Model with Structural Alphabet (AA+3Di)
MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining
PhoBERT: Pre-trained language models for Vietnamese (EMNLP-2020 Findings)
Implementation of Toolformer, Language Models That Can Use Tools, by MetaAI
Ongoing research training transformer language models at scale, including: BERT & GPT-2
Guide to using pre-trained large language models of source code
Empower Sequence Labeling with Task-Aware Language Model
博客配套视频链接: https://space.bilibili.com/383551518?spm_id_from=333.1007.0.0 b 站直接看 配套 github 链接:https://github.com/nickchen121/Pre-training-language-model 配套博客链接:https://www.cnblogs.com/nickchen121/p/15105048.html
[ICLR 2025 Spotlight] Official implementation of "Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts"
Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
[COLM 2025] Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
Implementation of Classifier Free Guidance in Pytorch, with emphasis on text conditioning, and flexibility to include multiple text embedding models
Transformer models from BERT to GPT-4, environments from Hugging Face to OpenAI. Fine-tuning, training, and prompt engineering examples. A bonus section with ChatGPT, GPT-3.5-turbo, GPT-4, and DALL-E including jump starting GPT-4, speech-to-text, text-to-speech, text to image generation with DALL-E, Google Cloud AI,HuggingGPT, and more
24,540 repositories in the index in total.