Tencent-Hunyuan/MixGRPO
quality grade C, 57 out of 100[ECCV 2026] MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
- stars
- 1.2k
- stars gained this week
- —
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
RL algorithms, environments, simulators and decision-making systems.
Signals: reinforcement-learning, deep-reinforcement-learning, rl, gymnasium, openai-gym, multi-agent-reinforcement-learning, imitation-learning
653 results
[ECCV 2026] MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
Generating sets of formulaic alpha (predictive) stock factors via reinforcement learning.
Stock Trading Bot using Deep Q-Learning
Computational framework for reinforcement learning in traffic control
Training a humanoid robot for locomotion using Reinforcement Learning
ChainerRL is a deep reinforcement learning library built on top of Chainer.
Implement AlphaZero/AlphaGo Zero methods on Chinese chess.
A simple and well styled PPO implementation. Based on my Medium series: https://medium.com/@eyyu/coding-ppo-from-scratch-with-pytorch-part-1-4-613dfc1b14c8.
PyTorch implementation of Deep Reinforcement Learning: Policy Gradient methods (TRPO, PPO, A2C) and Generative Adversarial Imitation Learning (GAIL). Fast Fisher vector product TRPO.
Minimal Deep Q Learning (DQN & DDQN) implementations in Keras
PyTorch implementation of Asynchronous Advantage Actor Critic (A3C) from "Asynchronous Methods for Deep Reinforcement Learning".
PyTorch implementation of Soft Actor-Critic (SAC), Twin Delayed DDPG (TD3), Actor-Critic (AC/A2C), Proximal Policy Optimization (PPO), QT-Opt, PointNet..
SMAC: The StarCraft Multi-Agent Challenge
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning & ReCall: Learning to Reason with Tool Call for LLMs via Reinforcement Learning
Concise and beautiful algorithms written in Julia
TextWorld is a sandbox learning environment for the training and evaluation of reinforcement learning (RL) agents on text-based games.
Softlearning is a reinforcement learning framework for training maximum entropy policies in continuous domains. Includes the official implementation of the Soft Actor-Critic algorithm.
A collection of reference Jupyter notebooks and demo AI/ML applications for enterprise use cases: marketing, pricing, supply chain, smart manufacturing, and more.
[NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards
🚀🀄️ A fast and strong AI for riichi mahjong, powered by Rust and deep reinforcement learning.
For deep RL and the future of AI.
Free course that takes you from zero to Reinforcement Learning PRO 🦸🏻🦸🏽
[NeurIPS 2023 Spotlight] LightZero: A Unified Benchmark for Monte Carlo Tree Search in General Sequential Decision Scenarios (awesome MCTS)
Deep Neuroevolution
24,523 repositories in the index in total.