# Lack of convergence when increasing the number of workers

**URL:** <https://discuss.ray.io/t/lack-of-convergence-when-increasing-the-number-of-workers/11082>\
**Category:** RLlib\
**Created:** [June 19, 2023, 8:30am UTC](https://discuss.ray.io/t/lack-of-convergence-when-increasing-the-number-of-workers/11082 "2023-06-19T08:30:23Z")\
**Posts on this page:** 1\
**Showing post:** 10

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [November 24, 2024, 4:17pm UTC](https://discuss.ray.io/t/lack-of-convergence-when-increasing-the-number-of-workers/11082/10 "2024-11-24T16:17:28Z")

</div>

@Elena,

If this was PPO I would suggest that it was the std logits for the variance approaching zero. This problem has been reported many times. I have not seen it arise as an issue with SAC though. And I do not know enough of the RLlib implementation details of SAC off the top of my head to know if it is likely the issue.

There is an easy test and fix if that is the problem. You can set this in the config.  
`model_config_dict` `{"free_log_std": True}`

You can find a more detailed discussion here:

> [@PPO nan in actor logits](https://discuss.ray.io/t/ppo-nan-in-actor-logits/15140):
>
> How severe does this issue affect your experience of using Ray? -High: custom models and custom policies I am testing for work do not run effectively. I am using custom models and custom policies within the PyFlyt environment. Specifically using MAFixedwingDogfightEnv which I create a custom MultiAgentEnv with a reward wrapper as per the PyFlyt example on rllib. I am using two models 1.) mixture of gaussian critic with a TorchFC actor 2.) a critic and actor that are both TorchFC, but I will be…

---

_[View the full topic](https://discuss.ray.io/t/lack-of-convergence-when-increasing-the-number-of-workers/11082)._
