# Fetch action probability distribution from trained policy

**URL:** <https://discuss.ray.io/t/fetch-action-probability-distribution-from-trained-policy/7722>\
**Category:** RLlib\
**Created:** [September 29, 2022, 5:02pm UTC](https://discuss.ray.io/t/fetch-action-probability-distribution-from-trained-policy/7722 "2022-09-29T17:02:16Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![steff](https://avatars.discourse-cdn.com/v4/letter/s/e36b37/32.png) [@steff](https://discuss.ray.io/u/steff)\
**Post date:** [September 29, 2022, 5:02pm UTC](https://discuss.ray.io/t/fetch-action-probability-distribution-from-trained-policy/7722/1 "2022-09-29T17:02:16Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- High: It blocks me to complete my task.

How can I get the action probability distribution from a trained policy for a particular state using Ray/RLLib version 2?

Tried using policy.compute\_single\_action(state, full\_fetch=True), but that only fetches additional information for the selected action.

Thanks,  
Stefan

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [September 30, 2022, 1:22pm UTC](https://discuss.ray.io/t/fetch-action-probability-distribution-from-trained-policy/7722/2 "2022-09-30T13:22:18Z")

</div>

Hi @steff,

Others may have a better way but the best way I know of is to use the action\_dist\_inputs key in the extra fetches dictionary to create your own action distribution and compute the probabilities from there.

@arturn This is the third request for this I have seen this month in the forumns. Perhaps it makes sense to store all the action probabilities in addition to the selected actions.

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [September 30, 2022, 2:01pm UTC](https://discuss.ray.io/t/fetch-action-probability-distribution-from-trained-policy/7722/3 "2022-09-30T14:01:20Z")

</div>

Thanks! I’ll open a feature request!

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [September 30, 2022, 2:02pm UTC](https://discuss.ray.io/t/fetch-action-probability-distribution-from-trained-policy/7722/4 "2022-09-30T14:02:39Z")

</div>

@mannyv and yes, that’s how we do it ourselves. Fetch a fresh action distribution and use the provided inputs!

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [September 30, 2022, 2:09pm UTC](https://discuss.ray.io/t/fetch-action-probability-distribution-from-trained-policy/7722/5 "2022-09-30T14:09:55Z")

</div>

@steff In the meantime, please have a look at our Policy classes! [For example in SAC policies](https://github.com/ray-project/ray/blob/master/rllib/algorithms/sac/sac_torch_policy.py#L245) we turn action distribution inputs into distributions that you can sample from.

---

<div class="post-metadata">

**Author:** ![ihopethiswillfi](https://avatars.discourse-cdn.com/v4/letter/i/9f8e36/32.png) [@ihopethiswillfi](https://discuss.ray.io/u/ihopethiswillfi)\
**Post date:** [March 17, 2023, 7:13pm UTC](https://discuss.ray.io/t/fetch-action-probability-distribution-from-trained-policy/7722/6 "2023-03-17T19:13:55Z")

</div>

@arturn I’m another one watching that [feature request](https://github.com/ray-project/ray/issues/28923). It’s a blocker for me. I’ve looked at the provided example from arturn but unfortunately that code is way beyond my understanding.

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [March 17, 2023, 10:27pm UTC](https://discuss.ray.io/t/fetch-action-probability-distribution-from-trained-policy/7722/7 "2023-03-17T22:27:46Z")

</div>

@ihopethiswillfi Do you know what action distribution the algorithm you are looking at is using?  
Sample Batches passed around RLlib can contain a field `SampleBatch.ACTION_DIST_INPUTS`.  
That contains what is needed to reconstruct the action distribution object together with the class.

For example, a `DiagGaussian` distribution will take two vectors as inputs - both should be found in said field of a SampleBatch. So if you are running a PPO training, the batches passed around in your `PPO.training_step()` method will contain these inputs.

We are currently working on replacing the concept of `Policy` and will likely include the action distribution itself in such batches in the future.

---

<div class="post-metadata">

**Author:** ![ihopethiswillfi](https://avatars.discourse-cdn.com/v4/letter/i/9f8e36/32.png) [@ihopethiswillfi](https://discuss.ray.io/u/ihopethiswillfi)\
**Post date:** [March 18, 2023, 8:00am UTC](https://discuss.ray.io/t/fetch-action-probability-distribution-from-trained-policy/7722/8 "2023-03-18T08:00:13Z")

</div>

@arturn Perfect. Yea I really miss the `prediction_proba()` that I use pretty much every time I train a traditional ML model 🙂

I’m using PPO and my action space is discrete with 3 possible actions. However I think that I should dive more into the inner workings of RLlib and RL in general, before I will be able to solve this. Thanks a lot for your help.
