# How does StochasticSampling work?

**URL:** <https://discuss.ray.io/t/how-does-stochasticsampling-work/6506>\
**Category:** RLlib\
**Created:** [June 14, 2022, 9:14pm UTC](https://discuss.ray.io/t/how-does-stochasticsampling-work/6506 "2022-06-14T21:14:44Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![2dm](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/2dm/32/885_2.png) [@2dm](https://discuss.ray.io/u/2dm)\
**Post date:** [June 14, 2022, 9:14pm UTC](https://discuss.ray.io/t/how-does-stochasticsampling-work/6506/1 "2022-06-14T21:14:44Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- Medium: It contributes to significant difficulty to complete my task, but I can work around it.

Hi,

I am trying to understand how StochasticSampling works during training and evaluation.

From [this](https://discuss.ray.io/t/meaning-of-stochasticsampling-for-exploration/5056/6) and [this](https://discuss.ray.io/t/decay-of-stochasticsampling/5856) posts I understand that it supposed to take the model output (I use random\_timesteps=0), add some kind of noise , and then sample an action from it. However, I am very confused about what is actually happens.  
My questions:

1. Where in the code is this noise added?
2. Are there any properties to this noise?
3. In StochasticSampling class documentation there is: `Also allows for scheduled parameters for the distributions, such as lowering stddev, temperature, etc.. over time.` The only example I found was SoftQ that overrides `get_exploration_action` to add the temperature. Is there another example where stddev (is that the std of the noise?) is used?
4. In evaluation - if I use `"explore": True` in order to keep the policy stochastic, are the actions generated from the policy output alone (without argmax), or does StochasticSampling also affects their generation in this mode?
5. In PPO - does the entropy added to the loss “adds on top of” the exploration mechanism (particularly StochasticSampling) that is used? i.e. are there more randomly generated actions in this case because the agent has 2 sources that contribute to the exploration?

Thank you!

---

<div class="post-metadata">

**Author:** ![christy](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/christy/32/2934_2.png) [@christy](https://discuss.ray.io/u/christy)\
**Post date:** [June 16, 2022, 5:40am UTC](https://discuss.ray.io/t/how-does-stochasticsampling-work/6506/2 "2022-06-16T05:40:36Z")

</div>

Hi there! 👋🏼

Would you like to ask your question in RLlib Office Hours? It sounds like a good topic!

✍🏼 Just add discuss link to your question to this doc: [RLlib Office Hours - Google Docs](https://docs.google.com/document/d/1AG7lgtzHu8mpQT2UnHHhCnWuiHOMIM_0zSoRJZNrtt8/edit)

Thanks! Hope to see you there!

---

<div class="post-metadata">

**Author:** ![carlorop](https://avatars.discourse-cdn.com/v4/letter/c/b19c9b/32.png) [@carlorop](https://discuss.ray.io/u/carlorop)\
**Post date:** [June 22, 2022, 1:20pm UTC](https://discuss.ray.io/t/how-does-stochasticsampling-work/6506/3 "2022-06-22T13:20:05Z")

</div>

Are these Office Hours recorded? I would be very interested in knowing the answers for these questions

---

<div class="post-metadata">

**Author:** ![2dm](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/2dm/32/885_2.png) [@2dm](https://discuss.ray.io/u/2dm)\
**Post date:** [June 22, 2022, 4:32pm UTC](https://discuss.ray.io/t/how-does-stochasticsampling-work/6506/4 "2022-06-22T16:32:15Z")

</div>

I couldn’t make it to the last office hours to ask there, so you won’t find the answers in the recordings (they do record it, link is in the google doc @christy shared).

I still hope to get some help here.

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [June 27, 2022, 3:07pm UTC](https://discuss.ray.io/t/how-does-stochasticsampling-work/6506/5 "2022-06-27T15:07:50Z")

</div>

Hi @carlorop ,

1. Also in reference to your [other post](https://discuss.ray.io/t/decay-of-stochasticsampling/5856/3): [StochasticSampling](https://github.com/ray-project/ray/blob/master/rllib/utils/exploration/stochastic_sampling.py) will, if used in an algorithm, be called in the policies, like [here](https://github.com/ray-project/ray/blob/master/rllib/algorithms/simple_q/simple_q_tf_policy.py#L116-L118).
2. The distribution of the actions (and therefore of the noise if you will) is parameterized by the outputs of your model. For example, for a guassian diagonal of size _l_, 2\*_l_ outputs of the model will be needed to parameterize this distribution.
3. The parameters of the an exploratory action sampling step depend on the distribution used. Stddev is one of these.
4. A policy includes the stochastic sampling step in it’s _compute\_action_ methods. Therefore, choosing `explore=True` will be lead to output of an exploratory action.
5. The entropy is calculated based on the distribution parameters output by your model. This way, the entropy loss can efficiently control the variance.
