# Customise policy to only do forward/backward pass for certain observations

**URL:** <https://discuss.ray.io/t/customise-policy-to-only-do-forward-backward-pass-for-certain-observations/4247>\
**Category:** RLlib\
**Created:** [November 25, 2021, 10:42am UTC](https://discuss.ray.io/t/customise-policy-to-only-do-forward-backward-pass-for-certain-observations/4247 "2021-11-25T10:42:05Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![bmanczak](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/bmanczak/32/1802_2.png) [@bmanczak](https://discuss.ray.io/u/bmanczak)\
**Post date:** [November 25, 2021, 10:42am UTC](https://discuss.ray.io/t/customise-policy-to-only-do-forward-backward-pass-for-certain-observations/4247/1 "2021-11-25T10:42:06Z")

</div>

Hi all,

I am new to rllib and for my research project I want to create a simple PPO baseline for a discrete control problem.

I managed to get everything up and running and now I would like to customise the training process such that the forward/backward pass is only executed when certain characteristic of the observation is met. If it is not met I would like the agent to take a predetermined action (do-nothing action).

What is the best way to go about this in rllib?  
Thanks for the help!

---

<div class="post-metadata">

**Author:** ![felipeeeantunes](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/felipeeeantunes/32/86_2.png) [@felipeeeantunes](https://discuss.ray.io/u/felipeeeantunes)\
**Post date:** [December 8, 2021, 4:12pm UTC](https://discuss.ray.io/t/customise-policy-to-only-do-forward-backward-pass-for-certain-observations/4247/2 "2021-12-08T16:12:57Z")

</div>

Hi! I guess you can use a custom\_model and add a mask in the forward function based on the observation. Similar to Action Mask as described [here](https://towardsdatascience.com/action-masking-with-rllib-5e4bec5e7505).

---

<div class="post-metadata">

**Author:** ![bmanczak](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/bmanczak/32/1802_2.png) [@bmanczak](https://discuss.ray.io/u/bmanczak)\
**Post date:** [December 9, 2021, 9:11am UTC](https://discuss.ray.io/t/customise-policy-to-only-do-forward-backward-pass-for-certain-observations/4247/3 "2021-12-09T09:11:07Z")

</div>

Hi!

Thanks for the reply.  
That’s what I ended up doing and it works.  
However, as far as I understand, this still requires forward/backward pass, causing an overhead. I tried to solve the issue by customising the `compute_single_action` in the PPO trainer ([post](https://discuss.ray.io/t/controlling-compute-actions-during-training/4260)) but that did not work.

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [December 9, 2021, 10:36am UTC](https://discuss.ray.io/t/customise-policy-to-only-do-forward-backward-pass-for-certain-observations/4247/4 "2021-12-09T10:36:40Z")

</div>

@bmanczak,

Couldn’t you do this in the environment in reset and step?

1. Rllib calls reset or step
2. In your environment evaluate if you need actions from the policy. If not, take predetermined actions otherwise return transiton info.

---

<div class="post-metadata">

**Author:** ![bmanczak](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/bmanczak/32/1802_2.png) [@bmanczak](https://discuss.ray.io/u/bmanczak)\
**Post date:** [December 9, 2021, 11:13am UTC](https://discuss.ray.io/t/customise-policy-to-only-do-forward-backward-pass-for-certain-observations/4247/5 "2021-12-09T11:13:32Z")

</div>

Yes, that’s also an option and can achieve the desired result.  
Though I think it would be more elegant and amenable to changes to implement from the policy side.
