kyegomez/Sophia
quality grade D, 41 out of 100Effortless plugin and play Optimizer to cut model training costs by 50%. New optimizer that is 2x faster than Adam on LLMs.
- stars
- 382
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Core deep-learning frameworks and libraries for pretraining and distributed training.
Signals: deep-learning, neural-network, pytorch, tensorflow, jax, distributed-training, training, deepspeed
2,694 results
Effortless plugin and play Optimizer to cut model training costs by 50%. New optimizer that is 2x faster than Adam on LLMs.
Repo for Rho-1: Token-level Data Selection & Selective Pretraining of LLMs.
Pytorch Library for Relational Table Learning with LLMs.
该系列的目的是让读者可以在基础的pytorch上,不依赖任何其他现成的外部库,从零开始理解并实现一个大语言模型的所有组成部分,以及训练微调代码,因此读者仅需python,pytorch和最基础深度学习背景知识即可。
Awesome LLM pre-training resources, including data, frameworks, and methods.
Accelerating your LLM training to full speed! Made with ❤️ by ServiceNow Research
LLM-And-More is a professional, plug-and-play, llm trainer and application builder that guides you through the complete LLM workflow from data to evaluation, from training to deployment, from idea to sevice. / LLM-And-More 是一个专业、开箱即用的大模型训练及应用构建一站式解决方案,包含从数据到评估、从训练到部署、从想法到服务的全流程最佳实践。
No description
Build a ChatGPT like LLM from scratch in PyTorch, explained step by step.
Latency and Memory Analysis of Transformer Models for Training and Inference
CINO: Pre-trained Language Models for Chinese Minority (少数民族语言预训练模型)
Physics of Language Models: Part 4.2, Canon Layers at Scale where Synthetic Pretraining Resonates in Reality
A minimalistic framework for transparently training language models and storing comprehensive checkpoints for in-depth learning dynamics research.
Pytorch implementation of R-BERT: "Enriching Pre-trained Language Model with Entity Information for Relation Classification"
Empower Sequence Labeling with Task-Aware Neural Language Model | a PyTorch Tutorial to Sequence Labeling
Homepage for ProLong (Princeton long-context language models) and paper "How to Train Long-Context Language Models (Effectively)"
Collection of training data management explorations for large language models
Train a language model to answer Slack messages as you.
ICML'2022: Black-Box Tuning for Language-Model-as-a-Service & EMNLP'2022: BBTv2: Towards a Gradient-Free Future with Large Language Models
Collaborative Training of Large Language Models in an Efficient Way
Fully Open Language Models with Stellar Performance
Large-scale language modeling tutorials with PyTorch
[T-PAMI 2025] PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
The official repo for [JSTARS'24] "MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining"
24,524 repositories in the index in total.