# PPO entropy not decreasing in Ray=1.11.0 as Ray=1.2.0?

**URL:** <https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842>\
**Category:** RLlib\
**Created:** [April 17, 2022, 5:48am UTC](https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842 "2022-04-17T05:48:29Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![pengzh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/pengzh/32/813_2.png) [@pengzh](https://discuss.ray.io/u/pengzh)\
**Post date:** [April 17, 2022, 5:48am UTC](https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842/1 "2022-04-17T05:48:29Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- Medium: It contributes to significant difficulty to complete my task, but I can work around it.

Hi, I am training PPO agents in [MetaDrive environment](https://github.com/metadriverse/metadrive) and find that the training dynamics significantly diverges between ray==1.2.0 and ray==1.11.0

 ![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/2X/b/b1e2fbcc3fe1e105aad46417675bd7e9e719a957.jpeg)

You can find that the entropy goes to 0 in ray=1.2.0 but is stuck around 2 in ray=1.11.0

The hyper parameters of PPO are strictly identical in both trials. Environment is identical in both experiments. Summary:

- `sgd_minibatch_size`: 512
- `train_batch_size`: 1600 (since in MARL env the number of agents vary from 20 to 40, so the ACTUAL batch size might range from 20K to 50K)
- `rollout_segment_length`: 200
- `entropy_coeff`: 0
- `lr`: 3e-4
- `num_sgd_iters`: 5
- `num_workers`: 4

I have identified some differences but they are not the major causes according to my experiments:

- In `MultiGPUTrainOneStep` the batch is not shuffled as in ray=1.2.0. But my experiment using `simple_optimizer` yields same result. So the batch shuffling is not the cause.
- The auto-adjust `rollout_segment_length` does not affect my result since I can assure ` train_batch_size 1600` is divisible by `num_workers 4 * envs_per_worker 1 * rollout_segment_length 200 = 800`

Please note that this is not a strict comparison. The figures above only tell that the entropy of action distribution has different behavior under different version of ray. I want to get insight on what particular part during the update might affect the entropy dynamics. Thanks!

---

<div class="post-metadata">

**Author:** ![pengzh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/pengzh/32/813_2.png) [@pengzh](https://discuss.ray.io/u/pengzh)\
**Post date:** [April 17, 2022, 6:44am UTC](https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842/2 "2022-04-17T06:44:55Z")

</div>

Oops the figure is bad

This is ray=1.11.0 entropy dynamics:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/2X/6/6c71ddaddb4835bed1404b49ae0a3746c8ca7c80.jpeg)

And this is ray=1.2.0 entropy dynamics:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/2X/0/0b11aa084f0ace85d23b048f58d7815d38502099.jpeg)

---

<div class="post-metadata">

**Author:** ![hossein836](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hossein836/32/2559_2.png) [@hossein836](https://discuss.ray.io/u/hossein836)\
**Post date:** [April 17, 2022, 9:44am UTC](https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842/3 "2022-04-17T09:44:46Z")

</div>

how many times did you repeat the experiment? maybe it is because of bad initialization. that’s very common on policy gradient methods. as I know they didn’t change anything about ppo or default settings of trainer which means you are running exact same codes

---

<div class="post-metadata">

**Author:** ![pengzh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/pengzh/32/813_2.png) [@pengzh](https://discuss.ray.io/u/pengzh)\
**Post date:** [April 18, 2022, 2:54am UTC](https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842/4 "2022-04-18T02:54:02Z")

</div>

Experiment repeat 4 times. I don’t think this is due to the randomness in weight initialization. Maybe some implicit changes in training workflow might result this.

I am now grid-searching different versions of ray using the same config and environment and hope figure out which version of ray I can trust  
(though this task is far from my research…)

---

<div class="post-metadata">

**Author:** ![pengzh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/pengzh/32/813_2.png) [@pengzh](https://discuss.ray.io/u/pengzh)\
**Post date:** [April 18, 2022, 3:15am UTC](https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842/5 "2022-04-18T03:15:05Z")

</div>

In ray=1.10.0, the entropy is also decreasing so slowly

 ![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/2X/0/0b139d72970f5134571199fc36cc58cd00719ece.png)

---

<div class="post-metadata">

**Author:** ![pengzh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/pengzh/32/813_2.png) [@pengzh](https://discuss.ray.io/u/pengzh)\
**Post date:** [April 18, 2022, 7:04am UTC](https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842/6 "2022-04-18T07:04:29Z")

</div>

Super interesting! I have tried many ray versions and find that ray=1.4.0 already yields strange entropy. Here is the plots:

ray=1.4.0:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/2X/f/f036536e0a8c836321fd281503c7497d7f13a497.png)

ray=1.3.0:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/2X/3/32d93de348c28aa9886ca9f450301d9b3cee514b.png)

ray=1.2.0:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/2X/c/c6318245eaf108c0719a2e0487f8b8e798079b6a.png)

The conclusion is that the change between ray=1.3.0 and ray=1.4.0 causes the difference. I will dive into to see why

---

<div class="post-metadata">

**Author:** ![pengzh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/pengzh/32/813_2.png) [@pengzh](https://discuss.ray.io/u/pengzh)\
**Post date:** [April 18, 2022, 1:32pm UTC](https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842/7 "2022-04-18T13:32:28Z")

</div>

I have identified the strange behavior of PPO entropy emerged in ray=1.4.0

Could anyone help to identify the possible cause to this? Thank much in advance!

---

<div class="post-metadata">

**Author:** ![pengzh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/pengzh/32/813_2.png) [@pengzh](https://discuss.ray.io/u/pengzh)\
**Post date:** [April 18, 2022, 3:13pm UTC](https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842/8 "2022-04-18T15:13:35Z")

</div>

I can report more details here:

In ray=1.3.0:

- the value loss increases quickly (in first 15 iterations) to 40
- the episode length is relative short since agents explore and die quickly and the number of episodes per iteration is 1.5

In ray=1.4.0:

- the value loss increases slowly to 20
- the episode length is large since agents are almost not moving (taking random actions) and the episode per iteration is 1

We all know that “entropy is too high” means the exploration is too strong. This might because the value function is not learned rapidly in ray=1.4.0.

I am not sure which component in the system might affect this.

---

<div class="post-metadata">

**Author:** ![pengzh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/pengzh/32/813_2.png) [@pengzh](https://discuss.ray.io/u/pengzh)\
**Post date:** [January 9, 2023, 12:49am UTC](https://discuss.ray.io/t/ppo-entropy-not-decreasing-in-ray-1-11-0-as-ray-1-2-0/5842/9 "2023-01-09T00:49:01Z")

</div>

![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/2X/f/f2262baaa3ca65bef93b9ab33f9e624482f92eb6.png)

The issue is still here even with ray=2.2.0

Independent PPO trainer in a dense multi-agent environment achieves too high entropy.
