# How tu use PPO agent with env with masked actions?

**URL:** <https://discuss.ray.io/t/how-tu-use-ppo-agent-with-env-with-masked-actions/5924>\
**Category:** RLlib\
**Created:** [April 25, 2022, 9:37pm UTC](https://discuss.ray.io/t/how-tu-use-ppo-agent-with-env-with-masked-actions/5924 "2022-04-25T21:37:36Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Peter\_Pirog](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/peter_pirog/32/234_2.png) [@Peter\_Pirog](https://discuss.ray.io/u/Peter_Pirog)\
**Post date:** [April 25, 2022, 9:37pm UTC](https://discuss.ray.io/t/how-tu-use-ppo-agent-with-env-with-masked-actions/5924/1 "2022-04-25T21:37:36Z")

</div>

I try to use PPO with env with action masking but I have some problem with configuration:

```auto
import gym
from gym.spaces import Box, Dict, Discrete
import numpy as np
import random

from ray import tune
from ray.tune.registry import register_env

class ParametricActionsCartPoleNoEmbeddings(gym.Env):
    """Same as the above ParametricActionsCartPole.
    However, action embeddings are not published inside observations,
    but will be learnt by the model.
    At each step, we emit a dict of:
        - the actual cart observation
        - a mask of valid actions (e.g., [0, 0, 1, 0, 0, 1] for 6 max avail)
        - action embeddings (w/ "dummy embedding" for invalid actions) are
          outsourced in the model and will be learned.
    """

    def __init__ (self, max_avail_actions):
        # Randomly set which two actions are valid and available.
        self.left_idx, self.right_idx = random.sample(range(max_avail_actions), 2)
        self.valid_avail_actions_mask = np.array(
            [0.0] * max_avail_actions, dtype=np.float32
        )
        self.valid_avail_actions_mask[self.left_idx] = 1
        self.valid_avail_actions_mask[self.right_idx] = 1
        self.action_space = Discrete(max_avail_actions)
        self.wrapped = gym.make("CartPole-v0")
        self.observation_space = Dict(
            {
                "valid_avail_actions_mask": Box(0, 1, shape=(max_avail_actions,)),
                "cart": self.wrapped.observation_space,
            }
        )
        self._skip_env_checking = True

    def reset(self):
        return {
            "valid_avail_actions_mask": self.valid_avail_actions_mask,
            "cart": self.wrapped.reset(),
        }

    def step(self, action):
        if action == self.left_idx:
            actual_action = 0
        elif action == self.right_idx:
            actual_action = 1
        else:
            raise ValueError(
                "Chosen action was not one of the non-zero action embeddings",
                action,
                self.valid_avail_actions_mask,
                self.left_idx,
                self.right_idx,
            )
        orig_obs, rew, done, info = self.wrapped.step(actual_action)
        obs = {
            "valid_avail_actions_mask": self.valid_avail_actions_mask,
            "cart": orig_obs,
        }
        return obs, rew, done, info

if __name__ == " __main__":

    def env_creator(env_config={}):
        return ParametricActionsCartPoleNoEmbeddings(max_avail_actions=6) # return an env instance

    register_env("my_env", env_creator)

    tune.run("PPO",
             # algorithm specific configuration
             config={"env": "my_env",
                     "evaluation_interval": 2,
                     "evaluation_num_episodes": 20},
             local_dir="cartpole_v1", # directory to save results
             checkpoint_freq=2, # frequency between checkpoints
             keep_checkpoints_num=6, )

```

The error is:

> (PPOTrainer pid=330560) raise ValueError(  
> (PPOTrainer pid=330560) ValueError: (‘Chosen action was not one of the non-zero action embeddings’, 5, array([0., 0., 1., 0., 1., 0.], dtype=float32), 4, 2)  
> Traceback (most recent call last):

Does someone know how to chenge agent configuration to use env with masking?

---

<div class="post-metadata">

**Author:** ![sven1977](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sven1977/32/53_2.png) [@sven1977](https://discuss.ray.io/u/sven1977)\
**Post date:** [April 26, 2022, 7:32am UTC](https://discuss.ray.io/t/how-tu-use-ppo-agent-with-env-with-masked-actions/5924/2 "2022-04-26T07:32:18Z")

</div>

Hey @Peter_Pirog , you would need a custom model for this to work with e.g. PPO/APPO/IMPALA.

You can take a look here at the example script (`ray.rllib.examples.action_masking.py`), which uses the custom model(s): `ray.rllib.examples.models.action_mask_model.py::ActionMaskModel|TorchActionMaskModel`.

Then with:

```auto
config:
  model:
    custom_model: ActionMaskModel

```

it should work.

Your custom model basically has to make sure that the `available_actions` information from your observation is properly “translated” into very negative logits such that the exploration component (in case of PP: stochastic sampling) will never sample the unavailable actions.

---

<div class="post-metadata">

**Author:** ![sirjay](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sirjay/32/2486_2.png) [@sirjay](https://discuss.ray.io/u/sirjay)\
**Post date:** [April 28, 2022, 8:57am UTC](https://discuss.ray.io/t/how-tu-use-ppo-agent-with-env-with-masked-actions/5924/3 "2022-04-28T08:57:03Z")

</div>

@sven1977 what’s the difference between ` class ParametricActionsModel(DistributionalQTFModel)` and `class ActionMaskModel(TFModelV2)` ? I guess both performs masking invalid actions.

- rllib/examples/models/parametric\_actions\_model.py
- rllib/examples/models/action\_mask\_model.py

---

<div class="post-metadata">

**Author:** ![Peter\_Pirog](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/peter_pirog/32/234_2.png) [@Peter\_Pirog](https://discuss.ray.io/u/Peter_Pirog)\
**Post date:** [May 3, 2022, 7:32am UTC](https://discuss.ray.io/t/how-tu-use-ppo-agent-with-env-with-masked-actions/5924/4 "2022-05-03T07:32:46Z")

</div>

@sven1977 Thank You for the answer
