# Custom action space

**URL:** <https://discuss.ray.io/t/custom-action-space/11496>\
**Category:** Configure Algorithm, Training, Evaluation, Scaling\
**Created:** [July 20, 2023, 12:03pm UTC](https://discuss.ray.io/t/custom-action-space/11496 "2023-07-20T12:03:17Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Username1](https://avatars.discourse-cdn.com/v4/letter/u/ee59a6/32.png) [@Username1](https://discuss.ray.io/u/Username1)\
**Post date:** [July 20, 2023, 12:03pm UTC](https://discuss.ray.io/t/custom-action-space/11496/1 "2023-07-20T12:03:17Z")

</div>

Hello, I am in need to use a [Multinomial Distribution](https://en.wikipedia.org/wiki/Multinomial_distribution) as my action and observation space. This is not even included on Gym’s spaces.

One option that I have been working on, is to create a custom Gym space and map it to a Multinomial distribution. This involved doing surgery on RLLIB’s source code, but it has been working so far, however, it is very laborious and I only implemented it for TF.

I wonder if something similar (i.e. new action and observation space) can be achieved with [this functionality](https://docs.ray.io/en/latest/rllib/rllib-models.html#custom-action-distributions) called “custom action distributions”.

The reason that is not clear, is that in the [official example](https://docs.ray.io/en/latest/rllib/rllib-models.html#autoregressive-action-distributions) the model uses a Categorical distribution, which is not a new gym space. While in my case, I have a totally new gym space and distribution to sample from, as the Multinomial is not currently on RLLIB nor in gym.

I will be very grateful for any pointer.

Thanks!

---

<div class="post-metadata">

**Author:** ![PrasannaMaddila](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/prasannamaddila/32/4829_2.png) [@PrasannaMaddila](https://discuss.ray.io/u/PrasannaMaddila)\
**Post date:** [July 27, 2023, 3:08pm UTC](https://discuss.ray.io/t/custom-action-space/11496/2 "2023-07-27T15:08:03Z")

</div>

Hello !

Just a fellow user like you, but I had a similar use-case recently. I had to simulate if an agent has detected one of many objects in their current observations set. We ended up using `spaces.MultiBinary` for the observation space, but implemented the detection using `numpy` and the `action_mask` ( so essentially, `space.sample(action_mask)`).

I’d love to help you out if I can 🙂 Do you have an example in mind? Also, is this what you had in mind?

---

<div class="post-metadata">

**Author:** ![Username1](https://avatars.discourse-cdn.com/v4/letter/u/ee59a6/32.png) [@Username1](https://discuss.ray.io/u/Username1)\
**Post date:** [July 31, 2023, 10:01am UTC](https://discuss.ray.io/t/custom-action-space/11496/3 "2023-07-31T10:01:52Z")

</div>

> [@PrasannaMaddila](#):
>
> e you, but I had a similar use-case recently. I had to simulate if an agent has detected one of many objects in their current observations set. We ended up using `spaces.MultiBinary` for the observation space, but implemented the detection using `numpy` and the `action_mask` ( so essentially, `space.sample(action_mask)`).

Hi, thank you very much for your answer!.  
My problem is that after Ray \> 2.0 everything has changed and now I don’t know how to pass a custom action space, a custom policy and a custom loss. The old methods seem to be deprecated and the documentation hasn’t catch up.

I have a code running in Pytorch and I wanted to convert it to RLLIB, as simple as that. I can share my GitHub if you would like, it has a running code in Pytorch. It’s a private repo so I’d need your github user. It can be over PM if you want.

Thanks a lot!

---

<div class="post-metadata">

**Author:** ![PrasannaMaddila](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/prasannamaddila/32/4829_2.png) [@PrasannaMaddila](https://discuss.ray.io/u/PrasannaMaddila)\
**Post date:** [July 31, 2023, 1:44pm UTC](https://discuss.ray.io/t/custom-action-space/11496/4 "2023-07-31T13:44:29Z")

</div>

> [@Username1](#):
>
> Hi, thank you very much for your answer!.  
> My problem is that after Ray \> 2.0 everything has changed and now I don’t know how to pass a custom action space, a custom policy and a custom loss. The old methods seem to be deprecated and the documentation hasn’t catch up.

Hello, I’m a beginner too, but let’s do our best to solve this 🙂 I’ve been going heavily through their examples in the repository for my own code and they seem like good places to start, for example

1. [(Example) Deploying Autoregressive model + action dist](https://github.com/ray-project/ray/blob/702c36da157e846d4b1b9780a4ba314157528027/rllib/examples/autoregressive_action_dist.py)
2. [(Docs) Deploy Custom Torch Policy](https://docs.ray.io/en/latest/rllib/key-concepts.html#policies) + [(Example) Custom Torch Policy](https://github.com/ray-project/ray/blob/master/rllib/examples/custom_torch_policy.py)
3. [(Example) Custom Loss Function](https://github.com/ray-project/ray/blob/702c36da157e846d4b1b9780a4ba314157528027/rllib/examples/models/custom_loss_model.py#L9)

Do you have an example we can go through, or otherwise, PM me?

---

<div class="post-metadata">

**Author:** ![Username1](https://avatars.discourse-cdn.com/v4/letter/u/ee59a6/32.png) [@Username1](https://discuss.ray.io/u/Username1)\
**Post date:** [July 31, 2023, 2:13pm UTC](https://discuss.ray.io/t/custom-action-space/11496/5 "2023-07-31T14:13:17Z")

</div>

Thank you very much @PrasannaMaddila ! I believe there are things that used to work on Ray 2.0 and not afterwards. For example, everything that uses “PPPOTrainer” has been demised, and everything that uses “with\_updates” has either different imports or different logic from the official documentation.

For example, everything that uses `build_policy_class` as a deprecation warning [here](https://github.com/ray-project/ray/blob/702c36da157e846d4b1b9780a4ba314157528027/rllib/policy/policy_template.py#L40).

I believe the “solution” now is just to subclass the TorchPolicy. I will PM you.

Thank you very much!
