PRIME-RL/Entropy-Mechanism-of-RL
quality grade D, 35 out of 100The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.
- stars
- 447
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
RL algorithms, environments, simulators and decision-making systems.
Signals: reinforcement-learning, deep-reinforcement-learning, rl, gymnasium, openai-gym, multi-agent-reinforcement-learning, imitation-learning
653 results
The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.
A curated list of reinforcement learning with verifiable rewards (continually updated)
Code to reproduce the experiments in Sample Efficient Reinforcement Learning via Model-Ensemble Exploration and Exploitation (MEEE).
Spectrum sharing in vehicular networks based on multi-agent reinforcement learning, IEEE Journal on Selected Areas in Communications
FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control
Official Implementation for the paper "d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning"
Collection of Python code that solves the Gymnasium Reinforcement Learning environments, along with YouTube tutorials.
Reproduce results of the research article "Deep Reinforcement Learning Based Resource Allocation for V2V Communications"
Tactics2D: A Reinforcement Learning Environment Library with Generative Scenarios for Driving Decision-making
A Collection of Multi-Agent Reinforcement Learning (MARL) Resources
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
Code for the paper "FinRL-DeepSeek: LLM-Infused Risk-Sensitive Reinforcement Learning for Trading Agents" arXiv:2502.07393
JAX-accelerated Meta-Reinforcement Learning Environments Inspired by XLand and MiniGrid 🏎️
Lyapunov-guided Deep Reinforcement Learning for Stable Online Computation Offloading in Mobile-Edge Computing Networks
Reinforcement learning using Markov Decision Processes. For JS, written in C++.
A simple and highly efficient RTS-game-inspired environment for reinforcement learning (formerly Gym-MicroRTS)
[NeurIPS 2025] Reinforcement Learning for Reasoning in Large Language Models with One Training Example
A wrapper framework for Reinforcement Learning in the Webots robot simulator using Python 3.
Code for reco-gym: A Reinforcement Learning Environment for the problem of Product Recommendation in Online Advertising
Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Mirror of Stable-Baselines: a fork of OpenAI Baselines, implementations of reinforcement learning algorithms
DIAMBRA Arena: a New Reinforcement Learning Platform for Research and Experimentation
Use Reinforcement Learning to train an autonomous driving agent in CARLA Simulator
24,523 repositories in the index in total.