# Jump-Start Reinforcement Learning

**URL:** <https://discuss.ray.io/t/jump-start-reinforcement-learning/15991>\
**Category:** RLlib\
**Created:** [October 12, 2024, 2:06am UTC](https://discuss.ray.io/t/jump-start-reinforcement-learning/15991 "2024-10-12T02:06:19Z")\
**Posts on this page:** 1\
**Showing post:** 11

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [December 11, 2024, 2:34am UTC](https://discuss.ray.io/t/jump-start-reinforcement-learning/15991/11 "2024-12-11T02:34:36Z")

</div>

It looks like the jump-start has helped in that you have a higher win rate and the variance of the win rate appears much lower. Any idea why you the performance degrades over training in both cases?

I feel like a broken record because I point a lot of people in this direction. But there can be some issues training PPO with continuous activations. If you have not seen it you might want to give this post a read to see if you are experiencing this issue.

> [@PPO nan in actor logits](https://discuss.ray.io/t/ppo-nan-in-actor-logits/15140):
>
> How severe does this issue affect your experience of using Ray? -High: custom models and custom policies I am testing for work do not run effectively. I am using custom models and custom policies within the PyFlyt environment. Specifically using MAFixedwingDogfightEnv which I create a custom MultiAgentEnv with a reward wrapper as per the PyFlyt example on rllib. I am using two models 1.) mixture of gaussian critic with a TorchFC actor 2.) a critic and actor that are both TorchFC, but I will be…

---

_[View the full topic](https://discuss.ray.io/t/jump-start-reinforcement-learning/15991)._
