chrisliu298/awesome-on-policy-distillation
quality grade B, 72 out of 100A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
- stars
- 591
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
RL algorithms, environments, simulators and decision-making systems.
Signals: reinforcement-learning, deep-reinforcement-learning, rl, gymnasium, openai-gym, multi-agent-reinforcement-learning, imitation-learning
653 results
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
Official Repository for "Eureka: Human-Level Reward Design via Coding Large Language Models" (ICLR 2024)
Fine-tuned MARL algorithms on SMAC (100% win rates on most scenarios)
Summary of key papers and blogs about diffusion models to learn about the topic. Detailed list of all published diffusion robotics papers.
CleanDiffuser: An Easy-to-use Modularized Library for Diffusion Models in Decision Making
A General Automated Machine Learning framework to simplify the development of End-to-end AutoML toolkits in specific domains.
The Machine Learning project including ML/DL projects, notebooks, cheat codes of ML/DL, useful information on AI/AGI and codes or snippets/scripts/tasks with tips.
Multi-Agent Connected Autonomous Driving (MACAD) Gym environments for Deep RL. Code for the paper presented in the Machine Learning for Autonomous Driving Workshop at NeurIPS 2019:
A sandbox game/simulator that demonstrates machine learning with evolutionary algorithms.
Notebooks from the Machine Learning Specialization
A small Roguelike game that uses Machine Learning to power its entities. Originally used in talks by Ciro & Alessia.
Program learning to play Flappy Bird by machine learning (Neuroevolution)
Use deep learning, genetic programming and other methods to predict stock and market movements
Distributed Deep Learning-based Offloading for Mobile Edge Computing Networks
Deep Pink is a chess AI that learns to play chess using deep learning.
深度学习、强化学习、模仿学习与机器人
No description
Code for "Social-STGCNN: A Social Spatio-Temporal Graph Convolutional Neural Network for Human Trajectory Prediction" CVPR 2020
Phase-Functioned Neural Networks for Character Control
Lagrangian Neural Networks
N-BEATS is a neural-network based model for univariate timeseries forecasting. N-BEATS is a ServiceNow Research project that was started at Element AI.
Using the genetic algorithm and neural networks I trained up 5 snakes who will then fuse to become the ultimate snake, this is how I did it
An index of recommendation algorithms that are based on Graph Neural Networks. (TORS)
High Frequency Trading Price Prediction using LSTM Recursive Neural Networks
24,523 repositories in the index in total.