# PPO nan in actor logits

**URL:** <https://discuss.ray.io/t/ppo-nan-in-actor-logits/15140>\
**Category:** RLlib\
**Created:** [July 5, 2024, 11:22pm UTC](https://discuss.ray.io/t/ppo-nan-in-actor-logits/15140 "2024-07-05T23:22:59Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [July 6, 2024, 4:41pm UTC](https://discuss.ray.io/t/ppo-nan-in-actor-logits/15140/2 "2024-07-06T16:41:10Z")

</div>

Hi @tlaurie99,

Welcome to the forum.

If I had to venture a guess as to where the Nan’s originate it would be here:

` 250 self.dist = torch.distributions.normal.Normal(mean, torch.exp(log_std))`

It has been my experience that when using a continuous action space sometime during training the std logits that parameterize the action distribution can become very negative. Which leads to an std close to zero which causes a Nan when it divides by ~zero on the backward calculation of the normal distribution.

RLLIB is unique the popular frameworks in that it uses the nn policy to generate the log\_std values.

> <https://github.com/ray-project/ray/blob/72bdb2ac7fe3df753390d6a78eba35b845f38ce5/rllib/models/torch/torch_action_dist.py#L248>

If you look at cleanrl or sb3s implementation you will see that they register the log\_std as a parameter of the model so they can be learned but not as part of the nn layers.

Since you are already using a custom model you might try implementing this alternative to see if it helps.

cleanrl:

> <https://github.com/vwxyzjn/cleanrl/blob/65789babaae033433078504b4ff0b925d5e27b99/cleanrl/ppo_continuous_action.py#L129>

sb3:

> <https://github.com/DLR-RM/stable-baselines3/blob/d8148deeaad3dbd1fb2b601e6f21d71f210366b1/stable_baselines3/common/distributions.py#L150>

---

_[View the full topic](https://discuss.ray.io/t/ppo-nan-in-actor-logits/15140)._
