# Decay of StochasticSampling

**URL:** <https://discuss.ray.io/t/decay-of-stochasticsampling/5856>\
**Category:** RLlib\
**Created:** [April 18, 2022, 7:03pm UTC](https://discuss.ray.io/t/decay-of-stochasticsampling/5856 "2022-04-18T19:03:41Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![carlorop](https://avatars.discourse-cdn.com/v4/letter/c/b19c9b/32.png) [@carlorop](https://discuss.ray.io/u/carlorop)\
**Post date:** [April 18, 2022, 7:03pm UTC](https://discuss.ray.io/t/decay-of-stochasticsampling/5856/1 "2022-04-18T19:03:41Z")

</div>

According to the documentation, StochasticSampling is “An exploration that simply samples from a distribution”, I am still wondering what StochasticSampling does.If I am not wrong it adds random noise to the actions.

Assume that I am training a continuous agent such as the continuous PPO. Is this noise constant during the training? If not, how can I chose the decay of the StochasticSampling noise?

---

<div class="post-metadata">

**Author:** ![sven1977](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sven1977/32/53_2.png) [@sven1977](https://discuss.ray.io/u/sven1977)\
**Post date:** [April 26, 2022, 8:53am UTC](https://discuss.ray.io/t/decay-of-stochasticsampling/5856/2 "2022-04-26T08:53:24Z")

</div>

Hey @carlorop , thanks for posting this question!

The `StochasticSampling` exploration component does not have any decay mechanisms built-in. It really simply samples from the distribution given by the model’s outputs (e.g. n logits → n probabilities (add to 1.0) → sample an action from the thus parameterized categorical distribution).  
It’s used by algos such as PPO/IMPALA/APPO/PG/etc… (basically most on-policy algos) by default.

Other exploration components (e.g. `EpsilonGreedy`) do have a decay mechanism.

---

<div class="post-metadata">

**Author:** ![carlorop](https://avatars.discourse-cdn.com/v4/letter/c/b19c9b/32.png) [@carlorop](https://discuss.ray.io/u/carlorop)\
**Post date:** [June 9, 2022, 3:33pm UTC](https://discuss.ray.io/t/decay-of-stochasticsampling/5856/3 "2022-06-09T15:33:43Z")

</div>

Thank you very much for your reply, however, I am still trying to get my head around it for the case of multidimensional continuous outputs. From your answer, I assume that in the case of discrete algorithms it just adds noise to the logits. Would it just add noise to the actions in the case of continuous distributions?
