Xirider/finetune-gpt2xl
quality grade D, 42 out of 100Guide: Finetune GPT2-XL (1.5 Billion Parameters) and finetune GPT-NEO (2.7 B) on a single GPU with Huggingface Transformers using DeepSpeed
- stars
- 436
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Adapting pretrained models: PEFT/LoRA, instruction tuning, RLHF, DPO and preference alignment.
Signals: fine-tuning, finetuning, lora, peft, qlora, rlhf, dpo, instruction-tuning
283 results
Guide: Finetune GPT2-XL (1.5 Billion Parameters) and finetune GPT-NEO (2.7 B) on a single GPU with Huggingface Transformers using DeepSpeed
No description
Inference, Fine Tuning and many more recipes with Gemma family of models
Single image to Lora Model for Flux in ComfyUI using Llm and Flux Kontext
Finetuning of Falcon-7B LLM using QLoRA on Mental Health Conversational Dataset
[EMNLP 2024] LongAlign: A Recipe for Long Context Alignment of LLMs
Banishing LLM Hallucinations Requires Rethinking Generalization
输入现代汉语句子,生成古汉语风格的句子。基于荀子基座大模型,采用“文言文(古文)- 现代文平行语料”中的部分数据进行LoRA微调训练而得。
Deepspeed、LLM、Medical_Dialogue、医疗大模型、预训练、微调
Easy, fast, and cheap pretrain,finetune, serving for everyone
Implementation for "Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs"
AutoAudit—— the LLM for Cyber Security 网络安全大语言模型
Finetune ALL LLMs with ALL Adapeters on ALL Platforms!
HuatuoGPT2, One-stage Training for Medical Adaption of LLMs. (An Open Medical GPT)
Library for industrial alignment.
该仓库主要记录 LLMs 算法工程师相关的顶会论文研读笔记(多模态、PEFT、小样本QA问答、RAG、LMMs可解释性、Agents、CoT)
LLM Tuning with PEFT (SFT+RM+PPO+DPO with LoRA)
Dự án bao gồm: 1. Xây dựng bộ dữ Instructions Vietnamese (chất lượng, nhiều, và đa dạng). 2.LLM Training, Finetuning, Evaluating & Testing trên Open-source mô hình ngôn ngữ: Bloomz,T5, UL2, LLaMA (1&2), OpenLLaMA, GPT-J pythia etc. 3. Ứng dụng và Giao diện Người dùng (UI)
Multi-agent Social Simulation + Efficient, Effective, and Stable alternative of RLHF. Code for the paper "Training Socially Aligned Language Models in Simulated Human Society".
Train a Language Model with GRPO to create a schedule from a list of events and priorities
[ICML'24] Data and code for our paper "Training-Free Long-Context Scaling of Large Language Models"
ArcticTraining is a framework designed to simplify and accelerate the post-training process for large language models (LLMs)
PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models(NeurIPS 2024 Spotlight)
Official PyTorch implementation of DistiLLM: Towards Streamlined Distillation for Large Language Models (ICML 2024)
24,535 repositories in the index in total.