# Env precheck inconsistent with Trainer

**URL:** <https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268>\
**Category:** RLlib\
**Created:** [May 25, 2022, 6:57pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268 "2022-05-25T18:57:21Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![rusu24edward](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rusu24edward/32/333_2.png) [@rusu24edward](https://discuss.ray.io/u/rusu24edward)\
**Post date:** [May 25, 2022, 6:57pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/1 "2022-05-25T18:57:21Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- Low: It annoys or frustrates me for a moment.

Recently, I’ve been upgrading my rllib dependency, and I ran into a number of issues starting with version 1.10. I see that in this version there was a bit of an overhaul to the MultiAgentEnv class. I’ve been making some code changes on my end to match up with these changes, and I’m running into some issues. I’m trying to connect with the newest version of ray, which is 1.12.1 as of this post.

I have a simple game called [MultiCorridor](https://github.com/LLNL/Abmarl/blob/main/abmarl/sim/corridor/multi_corridor.py) that I use for testing. This is a multiagent game, and each agent has the following observation space:

```auto
observation_space={
    'position': Box(0, self.end-1, (1,), int),
    'left': MultiBinary(1),
    'right': MultiBinary(1)
}

```

Here is a snippet of the function that returns the observation, which is called from my step and reset functions:

```python
agent_position = self.agents[agent_id].position
if agent_position == 0 or self.corridor[agent_position-1] is None:
    left = False
else:
    left = True
if agent_position == self.end-1 or self.corridor[agent_position+1] is None:
    right = False
else:
    right = True
return {
    'position': [agent_position],
    'left': [left],
    'right': [right],
}

```

I’ve broken down my issue into two parts:

## The checker makes design harder

In previous version of rllib, the trainer was smart enough to see that the observations are in the observation space, even though the types don’t match up exactly.

When I attempt to run this with rllib 1.12.1, I get:

```auto
ValueError: The observation collected from env.reset was not contained within your env's observation space. Its possible that there was a typemismatch (for example observations of np.float32 and a space ofnp.float64 observations), or that one of the sub-observations wasout of bounds

 reset_obs: {'agent0': {'position': [0], 'left': [False], 'right': [False]}}

 env.observation_space_sample(): {'agent1': OrderedDict([('left', array([1], dtype=int8)), ('position', array([4])), ('right', array([0], dtype=int8))]), 'agent2': OrderedDict([('left', array([0], dtype=int8)), ('position', array([0])), ('right', array([1], dtype=int8))]), 'agent3': OrderedDict([('left', array([0], dtype=int8)), ('position', array([2])), ('right', array([0], dtype=int8))]), 'agent4': OrderedDict([('left', array([0], dtype=int8)), ('position', array([1])), ('right', array([0], dtype=int8))]), 'agent0': OrderedDict([('left', array([1], dtype=int8)), ('position', array([1])), ('right', array([0], dtype=int8))])}

```

So I tried making this match up by changing my code. I got to this point:

```python
agent_position = self.agents[agent_id].position
if agent_position == 0 or self.corridor[agent_position-1] is None:
    left = False
else:
    left = True
if agent_position == self.end-1 or self.corridor[agent_position+1] is None:
    right = False
else:
    right = True
out = OrderedDict()
out['left'] = np.array([int(left)], dtype=np.int8)
out['position'] = np.array([agent_position])
out['right'] = np.array([int(right)], dtype=np.int8)
return out

```

and I still get this error:

```auto
ValueError: The observation collected from env.reset was not contained within your env's observation space. Its possible that there was a typemismatch (for example observations of np.float32 and a space ofnp.float64 observations), or that one of the sub-observations wasout of bounds

 reset_obs: {'agent0': OrderedDict([('left', array([1], dtype=int8)), ('position', array([7])), ('right', array([0], dtype=int8))])}

 env.observation_space_sample(): {'agent4': OrderedDict([('left', array([0], dtype=int8)), ('position', array([5])), ('right', array([1], dtype=int8))]), 'agent2': OrderedDict([('left', array([1], dtype=int8)), ('position', array([2])), ('right', array([1], dtype=int8))]), 'agent0': OrderedDict([('left', array([0], dtype=int8)), ('position', array([9])), ('right', array([0], dtype=int8))]), 'agent3': OrderedDict([('left', array([1], dtype=int8)), ('position', array([5])), ('right', array([1], dtype=int8))]), 'agent1': OrderedDict([('left', array([0], dtype=int8)), ('position', array([8])), ('right', array([1], dtype=int8))])}

```

At this point, I’m not really sure how to further modify the observation to match the type any closer. And besides that, it’s a bit ridiculous to be so detailed in the observation output instead of relying on a smart trainer that can match the types the way it did before.

## The checker is inconsistent with the trainer

I turned off the environment checker and ran my env with the super-detailed observation output above. I was able to train and achieve the results that I reached with previous version of rllib. It seems strange to me that the environment checker would fail but the trainer would still run.

Furthermore, I was able to simplify the observations to this:

```auto
return {
    'position': np.array([agent_position]),
    'left': np.array([int(left)]),
    'right': np.array([int(right)]),
}

```

and the training runs. I could not simplify it further (i.e. back to what I had it before).

## Suggestions

1. The environment checker should not be stricter than the trainer.
2. The trainers shouldn’t expect to receive **exactly** the same type as specified in the space. They should be smarter, at least as smart as they used to be.

---

<div class="post-metadata">

**Author:** ![hossein836](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hossein836/32/2559_2.png) [@hossein836](https://discuss.ray.io/u/hossein836)\
**Post date:** [May 26, 2022, 8:26am UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/2 "2022-05-26T08:26:12Z")

</div>

Great, I’m asking just out of curiosity, your env is not gym and you are not subclassed `Base.Env`. your `reset` and `step` methods doesn’t have return. I thought we at least should subclass `Base.Env` if we don’t use gym. is that the env you are using?

---

<div class="post-metadata">

**Author:** ![rusu24edward](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rusu24edward/32/333_2.png) [@rusu24edward](https://discuss.ray.io/u/rusu24edward)\
**Post date:** [May 26, 2022, 3:20pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/3 "2022-05-26T15:20:56Z")

</div>

Hi @hossein836, thanks for your quick reply. I use my own simulation framework which does not inherit anything from gym or rllib. This allows me flexibility in my design. In order to connect to Rllib, I wrap my sim in a custom [Turn Based Manager](https://github.com/LLNL/Abmarl/blob/main/abmarl/managers/turn_based_manager.py), and then I wrap that object in my [Multi Agent Wrapper](https://github.com/LLNL/Abmarl/blob/abmarl-epic-263-upgrade-ray-version/abmarl/external/rllib_multiagentenv_wrapper.py). It’s a bit of a layered onion, but this design enables me to connect to a few different learning libraries as needed.

My environment is a MultiAgentEnv when it gets plugged into tune.

---

<div class="post-metadata">

**Author:** ![hossein836](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hossein836/32/2559_2.png) [@hossein836](https://discuss.ray.io/u/hossein836)\
**Post date:** [May 26, 2022, 5:54pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/4 "2022-05-26T17:54:29Z")

</div>

now I see, good experiment overall. good luck 👍

---

<div class="post-metadata">

**Author:** ![rusu24edward](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rusu24edward/32/333_2.png) [@rusu24edward](https://discuss.ray.io/u/rusu24edward)\
**Post date:** [May 26, 2022, 6:26pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/5 "2022-05-26T18:26:28Z")

</div>

Upon more experimentation, I discovered that the issue is with my [TurnBasedManager](https://github.com/LLNL/Abmarl/blob/main/abmarl/managers/turn_based_manager.py). My game is designed to return only a single agent’s (obs, reward, done, info) in each step since I want my agents to take turns and since RLlib only produces actions for agents that report an observation.

Apparently, the env checker doesn’t like this and complains that it only saw the observations from a single agent and not the observations from all available agents. Is this the way env checker is supposed to work?

---

<div class="post-metadata">

**Author:** ![hossein836](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hossein836/32/2559_2.png) [@hossein836](https://discuss.ray.io/u/hossein836)\
**Post date:** [May 26, 2022, 6:59pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/6 "2022-05-26T18:59:03Z")

</div>

so what should you pass when you don’t want to pass anything to an agent? none? I hardly remember but I think that was ok to not pass anything related to an agent.

---

<div class="post-metadata">

**Author:** ![rusu24edward](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rusu24edward/32/333_2.png) [@rusu24edward](https://discuss.ray.io/u/rusu24edward)\
**Post date:** [May 26, 2022, 8:44pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/7 "2022-05-26T20:44:41Z")

</div>

@hossein836 RLlib will generate actions only for those agents who reported an observation from step, so it’s okay to only report output from a subset of agents in each step (e.g. turn based games). This is why it is strange to me that the env checker seems to require output from all agents because it violates this feature.

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [May 27, 2022, 3:07pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/8 "2022-05-27T15:07:51Z")

</div>

Hey @hossein836 ,

For now, you can turn off env checking with `disable_env_checking=True`.  
I’ll find out if this a design issue or a bug.

Cheers

---

<div class="post-metadata">

**Author:** ![avnishn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/avnishn/32/1544_2.png) [@avnishn](https://discuss.ray.io/u/avnishn)\
**Post date:** [May 27, 2022, 8:17pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/9 "2022-05-27T20:17:05Z")

</div>

Author of the env pre-checker here:

Yeah this is an unnecessary detail that we check for. We should be be able to check that spaces are corrected if only specific agents return observations in your environment. Do you mind opening a ticket for this on github?

I’m not sure if we can make the pre-checker as lax as the trainer generally. Right now we use gym directly to do element and space checking, although, you can overload this checking in your environment by writing your own checking functions,

```auto
def observation_space_contains(self, x: MultiAgentDict) -> bool:

def action_space_contains(self, x: MultiAgentDict) -> bool:

def action_space_sample(self, agent_ids: list = None) -> MultiAgentDict:

def observation_space_sample(self, agent_ids: list = None) -> MultiEnvDict:
    

```

---

<div class="post-metadata">

**Author:** ![rusu24edward](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rusu24edward/32/333_2.png) [@rusu24edward](https://discuss.ray.io/u/rusu24edward)\
**Post date:** [June 1, 2022, 4:25pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/10 "2022-06-01T16:25:36Z")

</div>

Thanks @avnishn, I will do as you suggest.

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [June 6, 2022, 1:12pm UTC](https://discuss.ray.io/t/env-precheck-inconsistent-with-trainer/6268/11 "2022-06-06T13:12:21Z")

</div>

I’ve set up a PR that deals with this issue by warning instead of raising an error in your case @rusu24edward. I’ve used your library to test this.  
You might want to add a super().\_\_init\_\_() to your MultiAgentWrapper, since that’s also something our env checker looks out for!

@avnishn is the expert here and I’ll ask him for review.
