UCSC-VLAA/MedReason
quality grade D, 35 out of 100MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
- stars
- 283
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Adapting pretrained models: PEFT/LoRA, instruction tuning, RLHF, DPO and preference alignment.
Signals: fine-tuning, finetuning, lora, peft, qlora, rlhf, dpo, instruction-tuning
283 results
MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署.
A proof-of-concept project that showcases the potential for using small, locally trainable LLMs to create next-generation documentation tools.
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search (NeurIPS 2024)
Simple Python library/structure to ablate features in LLMs which are supported by TransformerLens
🚀 2026届大模型算法岗实习面经 | 包含 DeepSeek/Qwen 技术报告解析、手撕 PPO/RoPE/Transformer、RLHF 核心与八股文 | 持续更新中...
A library for easily merging multiple LLM experts, and efficiently train the merged LLM.
LLM finetuned for medical question answering
An Open Source Toolkit For LLM Distillation
[ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA TraDo series.
🚀WebUI integrated platform for latest LLMs | 各大语言模型的全流程工具 WebUI 整合包。支持主流大模型API接口和开源模型。支持知识库,数据库,角色扮演,mj文生图,LoRA和全参数微调,数据集制作,live2d等全流程应用工具
大语言模型微调,Qwen2VL、Qwen2、GLM4指令微调
轻量级 LLM Post-training 框架,支持 SFT、RLVR、On-Policy KD、Guide KD 及混合训练;实现单轮/多轮 Guide 蒸馏、多教师蒸馏、Reward 混合训练与自动化数据分流👩🎓👨🎓
Tuning LLMs with no tears💦; Sample Design Engineering (SDE) for more efficient downstream-tuning.
Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning
GenEval: An object-focused framework for evaluating text-to-image alignment
Hypernetworks that adapt LLMs for specific benchmark tasks using only textual task description as the input
A very simple GRPO implement for reproducing r1-like LLM thinking.
[ICLR 2025] LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
An Open-source Toolkit for LLM Development
Training LLMs with QLoRA + FSDP
Create Custom LLMs
Video Diffusion Alignment via Reward Gradients. We improve a variety of video diffusion models such as VideoCrafter, OpenSora, ModelScope and StableVideoDiffusion by finetuning them using various reward models such as HPS, PickScore, VideoMAE, VJEPA, YOLO, Aesthetics etc.
24,535 repositories in the index in total.