# RLlib benchmarking

**URL:** <https://discuss.ray.io/t/rllib-benchmarking/23378>\
**Category:** Debugging and performance tuning\
**Created:** [December 17, 2025, 11:52am UTC](https://discuss.ray.io/t/rllib-benchmarking/23378 "2025-12-17T11:52:32Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Diplomat](https://avatars.discourse-cdn.com/v4/letter/d/aca169/32.png) [@Diplomat](https://discuss.ray.io/u/Diplomat)\
**Post date:** [December 17, 2025, 11:52am UTC](https://discuss.ray.io/t/rllib-benchmarking/23378/1 "2025-12-17T11:52:32Z")

</div>

I’m searching for RLlib benchmark results.  
(I’d simply like to validate e.g. PPO on Mujoco’s Walker2d where I experience significantly inferior RLlib performance compared to e.g. SB3)

- [GitHub - ray-project/rl-experiments: Keeping track of RL experiments](https://github.com/ray-project/rl-experiments) seems a bit outdated
- the examples folder e.g. [ray/rllib/examples/algorithms/ppo/benchmark\_ppo\_mujoco.py at master · ray-project/ray · GitHub](https://github.com/ray-project/ray/blob/master/rllib/examples/algorithms/ppo/benchmark_ppo_mujoco.py) doesn’t seem to contain exact benchmark results

---

<div class="post-metadata">

**Author:** ![RunLLM](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/runllm/32/8249_2.png) [@RunLLM](https://discuss.ray.io/u/RunLLM)\
**Post date:** [December 17, 2025, 11:52am UTC](https://discuss.ray.io/t/rllib-benchmarking/23378/2 "2025-12-17T11:52:57Z")

</div>

It seems likely that RLlib benchmark results are not directly included in the example scripts like `benchmark_ppo_mujoco.py`; these scripts provide configurations and stopping criteria for PPO on MuJoCo environments (including Walker2d), but do not contain actual benchmark numbers or logs within the file itself. The [rl-experiments repo](https://github.com/ray-project/rl-experiments) is referenced as a place for benchmark results, but it may not always be up to date or contain the latest results for every environment or algorithm combination you are interested in.

My understanding is that for the most current and detailed RLlib benchmark results, you should check both the [rl-experiments repo](https://github.com/ray-project/rl-experiments) and the [tuned\_examples directory](https://github.com/ray-project/ray/tree/master/rllib/tuned_examples), as suggested in the RLlib developer documentation. These locations are where the community is encouraged to contribute and update benchmark results, but there is no guarantee of completeness or recency for every environment. Would you like more detail on how to run your own benchmarks or interpret the provided example scripts?

Sources:

- [benchmark\_ppo\_mujoco.py](https://github.com/ray-project/ray/blob/master/rllib/examples/algorithms/ppo/benchmark_ppo_mujoco.py)
- [rl-experiments repo](https://github.com/ray-project/rl-experiments)

Hint: Mention @RunLLM in the post for followups.
