# \[RLlib, Tune, PPO\] episode\_reward\_mean based on new episodes for each iteration

**URL:** <https://discuss.ray.io/t/rllib-tune-ppo-episode-reward-mean-based-on-new-episodes-for-each-iteration/20753>\
**Category:** Configure Algorithm, Training, Evaluation, Scaling\
**Created:** [November 25, 2024, 8:42pm UTC](https://discuss.ray.io/t/rllib-tune-ppo-episode-reward-mean-based-on-new-episodes-for-each-iteration/20753 "2024-11-25T20:42:30Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![PhilippWillms](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/philippwillms/32/3885_2.png) [@PhilippWillms](https://discuss.ray.io/u/PhilippWillms)\
**Post date:** [November 25, 2024, 8:42pm UTC](https://discuss.ray.io/t/rllib-tune-ppo-episode-reward-mean-based-on-new-episodes-for-each-iteration/20753/1 "2024-11-25T20:42:30Z")

</div>

Hi community,

I would like to discuss with you an interesting observation which I recently regarding the `episode_reward_mean` which is logged for each training iteration.

The `episode_reward_mean` visible in the `progress.csv` file is always calculated taking into account ALL episodes which have been executed in the trial. Data reference can be obtained from `hist_stats/episode_reward`.

However, what do you think about an `episode_reward_mean_per_iteration` ? In that case, the metric would calculate a mean on the NEW episodes occurring in the current iteration ONLY.

I see following benefits:

1. Measurement on convergence to a (local) optimum in the reward function
2. Better judgement on solution quality in the sense of “will any more iterations make sense”

Happy to hear your thoughts

BR Philipp

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [November 25, 2024, 10:18pm UTC](https://discuss.ray.io/t/rllib-tune-ppo-episode-reward-mean-based-on-new-episodes-for-each-iteration/20753/2 "2024-11-25T22:18:34Z")

</div>

Hi @PhilippWillms,

As far as I am aware it does not use all episodes in the trial it only uses a configurable number of the most recent n episodes. This can be configured in the `reporting` options. Default is 100.

[https://docs.ray.io/en/latest/rllib/rllib-training.html#specifying-reporting-options](https://docs.ray.io/en/latest/rllib/rllib-training.html#specifying-reporting-options)

metrics\_num\_episodes\_for\_smoothing – Smooth rollout metrics over this many episodes, if possible. In case rollouts (sample collection) just started, there may be fewer than this many episodes in the buffer and we’ll compute metrics over this smaller number of available episodes. In case there are more than this many episodes collected in a single training iteration, use all of these episodes for metrics computation, meaning don’t ever cut any “excess” episodes. Set this to 1 to disable smoothing and to always report only the most recently collected episode’s return.
