# Tuning entropy in PPO

**URL:** <https://discuss.ray.io/t/tuning-entropy-in-ppo/1754>\
**Category:** RLlib\
**Created:** [April 16, 2021, 8:11am UTC](https://discuss.ray.io/t/tuning-entropy-in-ppo/1754 "2021-04-16T08:11:29Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ulrikah](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ulrikah/32/847_2.png) [@ulrikah](https://discuss.ray.io/u/ulrikah)\
**Post date:** [April 16, 2021, 8:11am UTC](https://discuss.ray.io/t/tuning-entropy-in-ppo/1754/1 "2021-04-16T08:11:29Z")

</div>

Hi,

I’m trying to tune exploration settings in PPO. In the default config, the entropy related values are `entropy_coeff = 0.0` and `entropy_schedule = None`. It doesn’t make sense to me, as the way I would interpret those default settings is that the agent has no incentive to explore. However, when I’ve experimented with increasing the entropy coefficient or by scheduling decaying entropy values, the models generally tend to perform worse than with the default settings. Few of the [tuned examples](https://github.com/ray-project/ray/tree/master/rllib/tuned_examples/ppo) modify these entropy related settings, which just seem odd to me. I’m pretty sure I don’t understand the full picture here, so does anyone care to explain?

Here is the loss calculation in PPO for reference:

```python
total_loss = reduce_mean_valid(
    -surrogate_loss
    + policy.kl_coeff * action_kl
    + policy.config["vf_loss_coeff"] * vf_loss
    - policy.entropy_coeff * curr_entropy

```

PS: The background for asking this is that I’m comparing SAC and PPO for a custom made environment. In SAC, there’s a sense that the higher the `maximum_entropy` is set, the more exploration happens. I know that the entropy in SAC and PPO signify different things, but what I’m trying to do is to compare exploration rates in the two algorithms.

---

<div class="post-metadata">

**Author:** ![sven1977](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sven1977/32/53_2.png) [@sven1977](https://discuss.ray.io/u/sven1977)\
**Post date:** [April 16, 2021, 1:35pm UTC](https://discuss.ray.io/t/tuning-entropy-in-ppo/1754/2 "2021-04-16T13:35:05Z")

</div>

Hey @ulrikah , thanks for the question. It’s true, by default (entropy\_coeff=0.0), we don’t incentivize the algo to produce high-entropy actions. I think it depends on the task you want to learn. For example for CartPole or Atari, you may not want to have too much focus on high entropy while you learn (I usually chose small initial\_alphas for SAC testing on CartPole, for lowering the entropy in general), but for HalfCheetah, this is completely different.

---

<div class="post-metadata">

**Author:** ![ulrikah](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ulrikah/32/847_2.png) [@ulrikah](https://discuss.ray.io/u/ulrikah)\
**Post date:** [April 16, 2021, 1:41pm UTC](https://discuss.ray.io/t/tuning-entropy-in-ppo/1754/3 "2021-04-16T13:41:22Z")

</div>

Thanks for the response!

Is this heuristic something you have developed empirically or is it related to the size of the action space or some other aspect of the environment?
