# Register a custom environment and runing PPOTrainer on that environment not working

**URL:** <https://discuss.ray.io/t/register-a-custom-environment-and-runing-ppotrainer-on-that-environment-not-working/6143>\
**Category:** RLlib\
**Created:** [May 15, 2022, 6:37am UTC](https://discuss.ray.io/t/register-a-custom-environment-and-runing-ppotrainer-on-that-environment-not-working/6143 "2022-05-15T06:37:32Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Shirelle\_Marcus](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/shirelle_marcus/32/1970_2.png) [@Shirelle\_Marcus](https://discuss.ray.io/u/Shirelle_Marcus)\
**Post date:** [May 15, 2022, 6:37am UTC](https://discuss.ray.io/t/register-a-custom-environment-and-runing-ppotrainer-on-that-environment-not-working/6143/1 "2022-05-15T06:37:32Z")

</div>

It blocks me to complete my task.

I’m trying to run the PPO algorithm on my custom gym environment (I’m new to new to RL). first I wrote a gyn env for my robotic dog, you can see it here:

```auto
import gym
from gym import error, spaces, utils
from gym.utils import seeding
import numpy as np
import random
from gym_dog.envs.mujoco import mujoco_env
from ray.rllib.agents.ppo import PPOTrainer

class DogEnv2(mujoco_env.MujocoEnv, utils.EzPickle, gym.Env):

    def __init__ (self):
        mujoco_env.MujocoEnv. __init__ (self, "./go1/xml/go1.xml", 5)
        utils.EzPickle. __init__ (self)

    def step(self, action):
        xposbefore = self.sim.data.qpos[0]
        self.do_simulation(action, self.frame_skip)
        xposafter = self.sim.data.qpos[0]
        ob = self._get_obs()
        reward_ctrl = -0.1 * np.square(action).sum()
        reward_run = (xposafter - xposbefore) / self.dt
        reward = reward_ctrl + reward_run
        done = False
        return ob, reward, done, dict(reward_run=reward_run, reward_ctrl=reward_ctrl)

    def _get_obs(self):
        return np.concatenate(
            [
                self.sim.data.qpos.flat[1:],
                self.sim.data.qvel.flat,
            ]
        )

    def reset_model(self):
        return self._get_obs()

    def viewer_setup(self):
        self.viewer.cam.distance = self.model.stat.extent * 0.5

```

then, I wrote my main file like this:

```auto
import numpy as np
import gym
import mujoco_py
import random
import sys

sys.path.insert(1, '/home/mosaic-challenge/shirelle_ws/gym_shirelle/gym-dog')
import gym_dog
from gym_dog.envs.dog_env_2 import DogEnv2
import ray
from ray.rllib.agents import ppo
import tensorflow as tf
from ray.tune.registry import register_env
from ray import tune
from ray.tune.logger import pretty_print
import time

def create_my_env():
    import gym
    # from gym_dog.envs.dog_env_2 import DogEnv2
    env = gym.make('dog-v2')
    return env

ray.init()
env_creator = lambda config: create_my_env()
register_env('DogEnv2', env_creator)
ppo_config = ppo.DEFAULT_CONFIG.copy()
trainer = ppo.PPOTrainer(config=ppo_config, env="DogEnv2")
for _ in range(10):
    result = trainer.train()
    print(pretty_print(result))

ray.shutdown()

```

I’m getting- a raise error.UnregisteredEnv(‘No registered env with id: {}’.format(id))  
what is wrong with my registration?

I also tried the following:

```auto
ray.init()
env_creator = lambda config: create_my_env()
register_env('DogEnv2', env_creator)

tune.run(
        "PPO",
        stop={"episode_reward_mean": 200},
        config={
            "env": env_creator,
            "num_workers": 1,
            },
        )

```

and then I’m getting this error:

```auto
2022-05-15 09:54:40,617	INFO trial_runner.py:803 -- starting PPO_<function <lambda> at 0x7f43ecac4550>_e473d_00000
== Status ==
Current time: 2022-05-15 09:54:42 (running for 00:00:01.94)
Memory usage on this node: 8.1/62.6 GiB
Using FIFO scheduling algorithm.
Resources requested: 2.0/16 CPUs, 0/1 GPUs, 0.0/33.81 GiB heap, 0.0/16.9 GiB objects
Result logdir: /home/mosaic-challenge/ray_results/PPO
Number of trials: 1/1 (1 RUNNING)
+-------------------------------------------------------+----------+-------+
| Trial name | status | loc |
|-------------------------------------------------------+----------+-------|
| PPO_<function <lambda> at 0x7f43ecac4550>_e473d_00000 | RUNNING | |
+-------------------------------------------------------+----------+-------+

2022-05-15 09:54:42,410	ERROR trial_runner.py:876 -- Trial PPO_<function <lambda> at 0x7f43ecac4550>_e473d_00000: Error processing event.
NoneType: None
Result for PPO_<function <lambda> at 0x7f43ecac4550>_e473d_00000:
  trial_id: e473d_00000
  
== Status ==
Current time: 2022-05-15 09:54:42 (running for 00:00:01.95)
Memory usage on this node: 8.1/62.6 GiB
Using FIFO scheduling algorithm.
Resources requested: 0/16 CPUs, 0/1 GPUs, 0.0/33.81 GiB heap, 0.0/16.9 GiB objects
Result logdir: /home/mosaic-challenge/ray_results/PPO
Number of trials: 1/1 (1 ERROR)
+-------------------------------------------------------+----------+-------+
| Trial name | status | loc |
|-------------------------------------------------------+----------+-------|
| PPO_<function <lambda> at 0x7f43ecac4550>_e473d_00000 | ERROR | |
+-------------------------------------------------------+----------+-------+
Number of errored trials: 1
+-------------------------------------------------------+--------------+------------------------------------------------------------------------------------------------------------------------------+
| Trial name | # failures | error file |
|-------------------------------------------------------+--------------+------------------------------------------------------------------------------------------------------------------------------|
| PPO_<function <lambda> at 0x7f43ecac4550>_e473d_00000 | 1 | /home/mosaic-challenge/ray_results/PPO/PPO_<function <lambda> at 0x7f43ecac4550>_e473d_00000_0_2022-05-15_09-54-40/error.txt |
+-------------------------------------------------------+--------------+------------------------------------------------------------------------------------------------------------------------------+

2022-05-15 09:54:42,412	ERROR ray_trial_executor.py:102 -- An exception occurred when trying to stop the Ray actor:Traceback (most recent call last):
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/tune/ray_trial_executor.py", line 93, in post_stop_cleanup
    ray.get(future, timeout=0)
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/_private/client_mode_hook.py", line 105, in wrapper
    return func(*args, **kwargs)
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/worker.py", line 1811, in get
    raise value
ray.exceptions.RayActorError: The actor died because of an error raised in its creation task, ray::PPOTrainer. __init__ () (pid=29183, ip=132.72.112.217, repr=PPOTrainer)
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/agents/trainer.py", line 767, in __init__
    self._env_id: Optional[str] = self._register_if_needed(
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/agents/trainer.py", line 2810, in _register_if_needed
    raise ValueError(
ValueError: <function <lambda> at 0x7f3fbd2ba310> is an invalid env specification. You can specify a custom env as either a class (e.g., YourEnvCls) or a registered env id (e.g., "your_env").

(PPOTrainer pid=29183) 2022-05-15 09:54:42,407	ERROR worker.py:449 -- Exception raised in creation task: The actor died because of an error raised in its creation task, ray::PPOTrainer. __init__ () (pid=29183, ip=132.72.112.217, repr=PPOTrainer)
(PPOTrainer pid=29183) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/agents/trainer.py", line 767, in __init__
(PPOTrainer pid=29183) self._env_id: Optional[str] = self._register_if_needed(
(PPOTrainer pid=29183) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/agents/trainer.py", line 2810, in _register_if_needed
(PPOTrainer pid=29183) raise ValueError(
(PPOTrainer pid=29183) ValueError: <function <lambda> at 0x7f3fbd2ba310> is an invalid env specification. You can specify a custom env as either a class (e.g., YourEnvCls) or a registered env id (e.g., "your_env").
Traceback (most recent call last):
  File "main.py", line 44, in <module>
    tune.run(
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/tune/tune.py", line 695, in run
    raise TuneError("Trials did not complete", incomplete_trials)
ray.tune.error.TuneError: ('Trials did not complete', [PPO_<function <lambda> at 0x7f43ecac4550>_e473d_00000])

```

---

<div class="post-metadata">

**Author:** ![Peter\_Pirog](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/peter_pirog/32/234_2.png) [@Peter\_Pirog](https://discuss.ray.io/u/Peter_Pirog)\
**Post date:** [May 15, 2022, 8:10pm UTC](https://discuss.ray.io/t/register-a-custom-environment-and-runing-ppotrainer-on-that-environment-not-working/6143/2 "2022-05-15T20:10:31Z")

</div>

Maybe try:

```auto
    def env_creator(env_config={}):
        return DogEnv2() 

    register_env("my_env", env_creator)

```

---

<div class="post-metadata">

**Author:** ![Shirelle\_Marcus](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/shirelle_marcus/32/1970_2.png) [@Shirelle\_Marcus](https://discuss.ray.io/u/Shirelle_Marcus)\
**Post date:** [May 16, 2022, 6:22am UTC](https://discuss.ray.io/t/register-a-custom-environment-and-runing-ppotrainer-on-that-environment-not-working/6143/3 "2022-05-16T06:22:37Z")

</div>

thank you for your reply.  
I tried what you offered:

```auto
def create_my_env(env_config={}):
    return DogEnv2()

ray.init()
env_creator = lambda config: create_my_env()
register_env("my_env", env_creator)

tune.run(
        "PPO",
        stop={"episode_reward_mean": 200},
        config={
            "env": "my_env",
            "num_workers": 1,
            },
        )

```

but I’m getting this error now:

```auto
2022-05-16 09:18:52,062	INFO trial_runner.py:803 -- starting PPO_my_env_0e406_00000
(PPOTrainer pid=13071) 2022-05-16 09:18:53,848	INFO trainer.py:2295 -- Your framework setting is 'tf', meaning you are using static-graph mode. Set framework='tf2' to enable eager execution with tf2.x. You may also then want to set eager_tracing=True in order to reach similar execution speed as with static-graph mode.
(PPOTrainer pid=13071) 2022-05-16 09:18:54,332	INFO ppo.py:268 -- In multi-agent mode, policies will be optimized sequentially by the multi-GPU optimizer. Consider setting simple_optimizer=True if this doesn't work for you.
(PPOTrainer pid=13071) 2022-05-16 09:18:54,332	INFO trainer.py:864 -- Current log_level is WARN. For more information, set 'log_level': 'INFO' / 'DEBUG' or use the -v and -vv flags.
== Status ==
Current time: 2022-05-16 09:18:56 (running for 00:00:04.88)
Memory usage on this node: 6.5/62.6 GiB
Using FIFO scheduling algorithm.
Resources requested: 2.0/16 CPUs, 0/1 GPUs, 0.0/34.94 GiB heap, 0.0/17.47 GiB objects
Result logdir: /home/mosaic-challenge/ray_results/PPO
Number of trials: 1/1 (1 RUNNING)
+------------------------+----------+-------+
| Trial name | status | loc |
|------------------------+----------+-------|
| PPO_my_env_0e406_00000 | RUNNING | |
+------------------------+----------+-------+

2022-05-16 09:18:56,831	ERROR trial_runner.py:876 -- Trial PPO_my_env_0e406_00000: Error processing event.
NoneType: None
Result for PPO_my_env_0e406_00000:
  trial_id: 0e406_00000
  
== Status ==
Current time: 2022-05-16 09:18:56 (running for 00:00:04.88)
Memory usage on this node: 6.5/62.6 GiB
Using FIFO scheduling algorithm.
Resources requested: 0/16 CPUs, 0/1 GPUs, 0.0/34.94 GiB heap, 0.0/17.47 GiB objects
Result logdir: /home/mosaic-challenge/ray_results/PPO
Number of trials: 1/1 (1 ERROR)
+------------------------+----------+-------+
| Trial name | status | loc |
|------------------------+----------+-------|
| PPO_my_env_0e406_00000 | ERROR | |
+------------------------+----------+-------+
Number of errored trials: 1
+------------------------+--------------+-----------------------------------------------------------------------------------------------+
| Trial name | # failures | error file |
|------------------------+--------------+-----------------------------------------------------------------------------------------------|
| PPO_my_env_0e406_00000 | 1 | /home/mosaic-challenge/ray_results/PPO/PPO_my_env_0e406_00000_0_2022-05-16_09-18-52/error.txt |
+------------------------+--------------+-----------------------------------------------------------------------------------------------+

2022-05-16 09:18:56,833	ERROR ray_trial_executor.py:102 -- An exception occurred when trying to stop the Ray actor:Traceback (most recent call last):
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/tune/ray_trial_executor.py", line 93, in post_stop_cleanup
    ray.get(future, timeout=0)
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/_private/client_mode_hook.py", line 105, in wrapper
    return func(*args, **kwargs)
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/worker.py", line 1811, in get
    raise value
ray.exceptions.RayActorError: The actor died because of an error raised in its creation task, ray::PPOTrainer. __init__ () (pid=13071, ip=132.72.112.217, repr=PPOTrainer)
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/agents/trainer.py", line 1035, in _init
    raise NotImplementedError
NotImplementedError

During handling of the above exception, another exception occurred:

ray::PPOTrainer. __init__ () (pid=13071, ip=132.72.112.217, repr=PPOTrainer)
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/agents/trainer.py", line 830, in __init__
    super(). __init__ (
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/tune/trainable.py", line 149, in __init__
    self.setup(copy.deepcopy(self.config))
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/agents/trainer.py", line 911, in setup
    self.workers = WorkerSet(
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/evaluation/worker_set.py", line 134, in __init__
    remote_spaces = ray.get(
ray.exceptions.RayActorError: The actor died because of an error raised in its creation task, ray::RolloutWorker. __init__ () (pid=13141, ip=132.72.112.217, repr=<ray.rllib.evaluation.rollout_worker.RolloutWorker object at 0x7fcaf52ecc10>)
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/evaluation/rollout_worker.py", line 507, in __init__
    check_env(self.env)
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/utils/pre_checks/env.py", line 65, in check_env
    check_gym_environments(env)
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/utils/pre_checks/env.py", line 135, in check_gym_environments
    sampled_observation = env.observation_space.sample()
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/gym/spaces/box.py", line 42, in sample
    return self.np_random.uniform(low=self.low, high=high, size=self.shape).astype(self.dtype)
  File "mtrand.pyx", line 1133, in numpy.random.mtrand.RandomState.uniform
OverflowError: Range exceeds valid bounds

(PPOTrainer pid=13071) 2022-05-16 09:18:56,828	ERROR worker.py:449 -- Exception raised in creation task: The actor died because of an error raised in its creation task, ray::PPOTrainer. __init__ () (pid=13071, ip=132.72.112.217, repr=PPOTrainer)
(PPOTrainer pid=13071) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/agents/trainer.py", line 1035, in _init
(PPOTrainer pid=13071) raise NotImplementedError
(PPOTrainer pid=13071) NotImplementedError
(PPOTrainer pid=13071) 
(PPOTrainer pid=13071) During handling of the above exception, another exception occurred:
(PPOTrainer pid=13071) 
(PPOTrainer pid=13071) ray::PPOTrainer. __init__ () (pid=13071, ip=132.72.112.217, repr=PPOTrainer)
(PPOTrainer pid=13071) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/agents/trainer.py", line 830, in __init__
(PPOTrainer pid=13071) super(). __init__ (
(PPOTrainer pid=13071) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/tune/trainable.py", line 149, in __init__
(PPOTrainer pid=13071) self.setup(copy.deepcopy(self.config))
(PPOTrainer pid=13071) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/agents/trainer.py", line 911, in setup
(PPOTrainer pid=13071) self.workers = WorkerSet(
(PPOTrainer pid=13071) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/evaluation/worker_set.py", line 134, in __init__
(PPOTrainer pid=13071) remote_spaces = ray.get(
(PPOTrainer pid=13071) ray.exceptions.RayActorError: The actor died because of an error raised in its creation task, ray::RolloutWorker. __init__ () (pid=13141, ip=132.72.112.217, repr=<ray.rllib.evaluation.rollout_worker.RolloutWorker object at 0x7fcaf52ecc10>)
(PPOTrainer pid=13071) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/evaluation/rollout_worker.py", line 507, in __init__
(PPOTrainer pid=13071) check_env(self.env)
(PPOTrainer pid=13071) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/utils/pre_checks/env.py", line 65, in check_env
(PPOTrainer pid=13071) check_gym_environments(env)
(PPOTrainer pid=13071) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/utils/pre_checks/env.py", line 135, in check_gym_environments
(PPOTrainer pid=13071) sampled_observation = env.observation_space.sample()
(PPOTrainer pid=13071) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/gym/spaces/box.py", line 42, in sample
(PPOTrainer pid=13071) return self.np_random.uniform(low=self.low, high=high, size=self.shape).astype(self.dtype)
(PPOTrainer pid=13071) File "mtrand.pyx", line 1133, in numpy.random.mtrand.RandomState.uniform
(PPOTrainer pid=13071) OverflowError: Range exceeds valid bounds
(RolloutWorker pid=13141) 2022-05-16 09:18:56,824	WARNING rollout_worker.py:498 -- We've added a module for checking environments that are used in experiments. It will cause your environment to fail if your environment is not set upcorrectly. You can disable check env by setting `disable_env_checking` to True in your experiment config dictionary. You can run the environment checking module standalone by calling ray.rllib.utils.check_env(env).
(RolloutWorker pid=13141) 2022-05-16 09:18:56,824	WARNING env.py:120 -- Your env doesn't have a .spec.max_episode_steps attribute. This is fine if you have set 'horizon' in your config dictionary, or `soft_horizon`. However, if you haven't, 'horizon' will default to infinity, and your environment will not be reset.
(RolloutWorker pid=13141) 2022-05-16 09:18:56,825	ERROR worker.py:449 -- Exception raised in creation task: The actor died because of an error raised in its creation task, ray::RolloutWorker. __init__ () (pid=13141, ip=132.72.112.217, repr=<ray.rllib.evaluation.rollout_worker.RolloutWorker object at 0x7fcaf52ecc10>)
(RolloutWorker pid=13141) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/evaluation/rollout_worker.py", line 507, in __init__
(RolloutWorker pid=13141) check_env(self.env)
(RolloutWorker pid=13141) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/utils/pre_checks/env.py", line 65, in check_env
(RolloutWorker pid=13141) check_gym_environments(env)
(RolloutWorker pid=13141) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/rllib/utils/pre_checks/env.py", line 135, in check_gym_environments
(RolloutWorker pid=13141) sampled_observation = env.observation_space.sample()
(RolloutWorker pid=13141) File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/gym/spaces/box.py", line 42, in sample
(RolloutWorker pid=13141) return self.np_random.uniform(low=self.low, high=high, size=self.shape).astype(self.dtype)
(RolloutWorker pid=13141) File "mtrand.pyx", line 1133, in numpy.random.mtrand.RandomState.uniform
(RolloutWorker pid=13141) OverflowError: Range exceeds valid bounds
Traceback (most recent call last):
  File "main.py", line 26, in <module>
    tune.run(
  File "/home/mosaic-challenge/anaconda3/envs/python3.8/lib/python3.8/site-packages/ray/tune/tune.py", line 695, in run
    raise TuneError("Trials did not complete", incomplete_trials)
ray.tune.error.TuneError: ('Trials did not complete', [PPO_my_env_0e406_00000])

```

any suggestions?

---

<div class="post-metadata">

**Author:** ![fardinabbasi](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/fardinabbasi/32/5085_2.png) [@fardinabbasi](https://discuss.ray.io/u/fardinabbasi)\
**Post date:** [September 21, 2023, 2:46pm UTC](https://discuss.ray.io/t/register-a-custom-environment-and-runing-ppotrainer-on-that-environment-not-working/6143/4 "2023-09-21T14:46:54Z")

</div>

Did you find any solution? I have the same problem.  
Registering costum environments that take env\_config is a common problem!

---

<div class="post-metadata">

**Author:** ![PhilippWillms](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/philippwillms/32/3885_2.png) [@PhilippWillms](https://discuss.ray.io/u/PhilippWillms)\
**Post date:** [September 23, 2023, 11:04am UTC](https://discuss.ray.io/t/register-a-custom-environment-and-runing-ppotrainer-on-that-environment-not-working/6143/5 "2023-09-23T11:04:22Z")

</div>

@fardinabbasi : Keep in mind that the class of your custom environment, inheriting from gym.Env., must also be “ready” to take an env\_config dictionary and parse it accordingly to attributes. See below a code snippet developed by me.

```auto
def assign_env_config(self, args, kwargs):
    """Configure instance based on args and keyword args."""
    # In order to ensure compatbility to multiple RL libraries, this method is supposed
    # to flexible treat the input as "args" or as "kwargs"

    # First path: Parameters are available as a dictionary, which is delivered under the keyword "env_config".
    # This case is occurring while gym.make().
    if kwargs is not None:
        for key, value in kwargs.items():
            setattr(self, key, value)
        if hasattr(self, "env_config"):
            for key, value in self.env_config.items():
                # Check types based on default settings
                if hasattr(self, key):
                    if type(getattr(self, key)) == np.ndarray:
                        setattr(self, key, value)
                    else:
                        setattr(self, key, type(getattr(self, key))(value))
                else:
                    raise AttributeError(f"{self} has no attribute {key}")

    # Second path: "env_config" is passed as flattened dictionary, which is part of a tuple.
    # This case is occurring in e.g. ray rllib.
    # While ray provides EnvContext to capture that properly, we want to avoid dependency on ray.
    if args is not None:
        for i in range(len(args)):
            args_item = args[i]
            for key, value in args_item.items():
                # Check types based on default settings
                if hasattr(self, key):
                    if type(getattr(self, key)) == np.ndarray:
                        setattr(self, key, value)
                    else:
                        setattr(self, key, type(getattr(self, key))(value))
                else:
                    raise AttributeError(f"{self} has no attribute {key}")

```

---

<div class="post-metadata">

**Author:** ![fardinabbasi](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/fardinabbasi/32/5085_2.png) [@fardinabbasi](https://discuss.ray.io/u/fardinabbasi)\
**Post date:** [September 23, 2023, 8:02pm UTC](https://discuss.ray.io/t/register-a-custom-environment-and-runing-ppotrainer-on-that-environment-not-working/6143/6 "2023-09-23T20:02:06Z")

</div>

@PhilippWillms :  
Thank you, Philipp, for your response, but I’m not entirely sure I understand what you mean by being “ready” to take an env\_config.  
This is my custom environment which only takes a env\_config, but when I try to run it with ray tune I get `Cannot create PPOConfig from given `config\_dict`! Property __stdout_file__ not supported.`  
Here is the environment:

> <https://github.com/ray-project/ray/issues/39807>
>
> \### What happened + What you expected to happen
> 
> I want to train a PPO agent i…n my custom environment called RankingEnv, but I'm encountering several errors and warnings that result in the agent's termination. The majority of these issues are:
> \`
> observation = OrderedDict(sorted(observation.items()))
> AttributeError: 'NoneType' object has no attribute 'items'
> \`
> \`\`\`
> WARNING deprecation.py:50 -- DeprecationWarning: \`ValueNetworkMixin\` has been deprecated. This will raise an error in the future!
> WARNING deprecation.py:50 -- DeprecationWarning: \`LearningRateSchedule\` has been deprecated. This will raise an error in the future!
> WARNING deprecation.py:50 -- DeprecationWarning: \`EntropyCoeffSchedule\` has been deprecated. This will raise an error in the future!
> WARNING deprecation.py:50 -- DeprecationWarning: \`KLCoeffMixin\` has been deprecated. This will raise an error in the future!
> \`\`\`
> \`\`\`
> DeprecationWarning: \`DirectStepOptimizer\` has been deprecated. This will raise an error in the future!
> \`\`\`
> \`\`\`
> WARNING algorithm\_config.py:672 -- Cannot create PPOConfig from given \`config\_dict\`! Property \_\_stdout\_file\_\_ not supported.
> \`\`\`
> \### Costum Environment
> \`\`\`ruby
> class RankingEnv(gym.Env):
> def \_\_init\_\_(self, config: dict):
> super().\_\_init\_\_()
> self.coins = config\['df'\]\['tic'\].unique()
> features\_col = config\['df'\].columns.difference(\['date','tic','score','close\_growth(%)'\])
> 
> self.df=config\['df'\].groupby('date')
> self.dates = list(self.df.groups.keys())
> 
> self.observation\_space = gym.spaces.Dict({coin: gym.spaces.Box(low=-np.inf, high=np.inf, shape=(len(features\_col),), dtype=np.float64) for coin in self.coins})
> self.action\_space = gym.spaces.Dict({coin: gym.spaces.Box(low=np.float64(1.0), high=np.float64(len(self.coins)), shape=(1,), dtype=np.float64) for coin in self.coins})#actions are the scores
> 
> def step(self, action):
> group = self.df.get\_group(self.dates\[self.time\])
> true\_scores = group.set\_index('tic')\['score'\].to\_dict()
> # The indexes are according to predicted score and the values are true score!(Sorting true score by predicted score)
> scores = \[true\_scores\[coin\] for coin in sorted(action, key=action.get, reverse=True)\]
> ideal\_scores = sorted(true\_scores.values(), reverse=True) #This is just the True\_score as list
> 
> dcg = self.calculate\_dcg(scores)
> idcg = self.calculate\_dcg(ideal\_scores)
> reward = dcg/idcg # reward = ndcg
> 
> self.time+=1
> terminated = self.time \>= len(self.dates)
> info = {}
> return self.\_get\_obs() if not terminated else None, reward, terminated, False, info
> 
> def calculate\_dcg(self, scores):
> dcg = 0.0
> for i in range(len(scores)):
> dcg += (2 \*\* scores\[i\] - 1) / np.log2(i + 2)
> return dcg
> 
> def reset(self, \*, seed: Optional\[int\] = None, options: Optional\[dict\] = None):
> super().reset(seed=seed,options=options)
> self.time = 0
> info={}
> return self.\_get\_obs() ,info
> 
> def \_get\_obs(self):
> group = self.df.get\_group(self.dates\[self.time\])
> obs = group.drop(\['date','score','close\_growth(%)'\], axis=1)
> 
> obs = obs.set\_index('tic').agg(lambda x: np.array(x.tolist()), axis=1).to\_dict()
>       
> return obs
> 
> def render(self):
> pass
> 
> def close(self):
> pass
> \`\`\`
> \### Agent
> 
> \`\`\`ruby
> class DRLlibv2:
> 
> 
> def \_\_init\_\_(
> self,
> trainable: str | Any,
> params: dict,
> train\_env=None,
> run\_name: str = "tune\_run",
> local\_dir: str = "tune\_results",
> search\_alg=None,
> concurrent\_trials: int = 0,
> num\_samples: int = 0,
> scheduler\_=None,
> num\_cpus: float | int = 2,
> dataframe\_save: str = "tune.csv",
> metric: str = "episode\_reward\_mean",
> mode: str | list\[str\] = "max",
> max\_failures: int = 0,
> training\_iterations: int = 100,
> checkpoint\_num\_to\_keep: None | int = None,
> checkpoint\_freq: int = 0,
> reuse\_actors: bool = True
> ):
> self.params = params
> 
> self.train\_env = train\_env
> self.run\_name = run\_name
> self.local\_dir = local\_dir
> self.search\_alg = search\_alg
> if concurrent\_trials != 0:
> self.search\_alg = ConcurrencyLimiter(
> self.search\_alg, max\_concurrent=concurrent\_trials
> )
> self.scheduler\_ = scheduler\_
> self.num\_samples = num\_samples
> self.trainable = trainable
> if isinstance(self.trainable, str):
> self.trainable = self.trainable.upper()
> self.num\_cpus = num\_cpus
> self.dataframe\_save = dataframe\_save
> self.metric = metric
> self.mode = mode
> self.max\_failures = max\_failures
> self.training\_iterations = training\_iterations
> self.checkpoint\_freq = checkpoint\_freq
> self.checkpoint\_num\_to\_keep = checkpoint\_num\_to\_keep
> self.reuse\_actors = reuse\_actors
> 
> def train\_tune\_model(self):
> 
> if ray.is\_initialized():
> ray.shutdown()
> 
> ray.init(num\_cpus=self.num\_cpus, num\_gpus=self.params\['num\_gpus'\], ignore\_reinit\_error=True)
> 
> if self.train\_env is not None:
> register\_env(self.params\['env'\], lambda env\_config: self.train\_env)
> 
> 
> tuner = tune.Tuner(
> self.trainable,
> param\_space=self.params,
> tune\_config=TuneConfig(
> search\_alg=self.search\_alg,
> scheduler=self.scheduler\_,
> num\_samples=self.num\_samples,
> # metric=self.metric,
> # mode=self.mode,
> \*\*({'metric': self.metric, 'mode': self.mode} if self.scheduler\_ is None else {}),
> reuse\_actors=self.reuse\_actors,
> 
> ),
> run\_config=RunConfig(
> name=self.run\_name,
> storage\_path=self.local\_dir,
> failure\_config=FailureConfig(
> max\_failures=self.max\_failures, fail\_fast=False
> ),
> stop={"training\_iteration": self.training\_iterations},
> checkpoint\_config=CheckpointConfig(
> num\_to\_keep=self.checkpoint\_num\_to\_keep,
> checkpoint\_score\_attribute=self.metric,
> checkpoint\_score\_order=self.mode,
> checkpoint\_frequency=self.checkpoint\_freq,
> checkpoint\_at\_end=True,
> ),
> verbose=3,#Verbosity mode. 0 = silent, 1 = default, 2 = verbose, 3 = detailed
> ),
> )
> 
> self.results = tuner.fit()
> if self.search\_alg is not None:
> self.search\_alg.save\_to\_dir(self.local\_dir)
> # ray.shutdown()
> return self.results
> 
> def infer\_results(self, to\_dataframe: str = None, mode: str = "a"):
> 
> results\_df = self.results.get\_dataframe()
> 
> if to\_dataframe is None:
> to\_dataframe = self.dataframe\_save
> 
> results\_df.to\_csv(to\_dataframe, mode=mode)
> 
> best\_result = self.results.get\_best\_result()
> # best\_result = self.results.get\_best\_result()
> # best\_metric = best\_result.metrics
> # best\_checkpoint = best\_result.checkpoint
> # best\_trial\_dir = best\_result.log\_dir
> # results\_df = self.results.get\_dataframe()
> 
> return results\_df, best\_result
> 
> def restore\_agent(
> self,
> checkpoint\_path: str = "",
> restore\_search: bool = False,
> resume\_unfinished: bool = True,
> resume\_errored: bool = False,
> restart\_errored: bool = False,
> ):
> 
> # if restore\_search:
> # self.search\_alg = self.search\_alg.restore\_from\_dir(self.local\_dir)
> if checkpoint\_path == "":
> checkpoint\_path = self.results.get\_best\_result().checkpoint.\_local\_path
> 
> restored\_agent = tune.Tuner.restore(
> checkpoint\_path,
> restart\_errored=restart\_errored,
> resume\_unfinished=resume\_unfinished,
> resume\_errored=resume\_errored,
> )
> print(restored\_agent)
> self.results = restored\_agent.fit()
> 
> if self.search\_alg is not None:
> self.search\_alg.save\_to\_dir(self.local\_dir)
> return self.results
> 
> def get\_test\_agent(self, test\_env\_name: str, test\_env=None, checkpoint=None):
> 
> # if test\_env is not None:
> # register\_env(test\_env\_name, lambda config: \[test\_env\])
> 
> if checkpoint is None:
> checkpoint = self.results.get\_best\_result().checkpoint
> 
> testing\_agent = Algorithm.from\_checkpoint(checkpoint)
> # testing\_agent.config\['env'\] = test\_env\_name
> 
> return testing\_agent
> \`\`\`
> 
> \### Versions / Dependencies
> 
> \- Operating system: Google Colab
> \- Python version: 3.10.12
> \- Ray version: 2.7.0
> 
> \### Reproduction script
> 
> \`\`\`ruby
> train\_env\_config = {'df': train\_data}
> train\_config = (PPOConfig()
> .training(lr=tune.loguniform(5e-5, 0.001), entropy\_coeff=tune.loguniform(0.00000001, 0.1),sgd\_minibatch\_size=tune.choice(\[32, 64, 128, 256, 512\]),lambda\_=tune.choice(\[0.1,0.3,0.5,0.7,0.9,1.0\]))
> .resources(num\_gpus=0)
> .debugging(log\_level="DEBUG", seed = 1234)
> .rollouts(num\_rollout\_workers=1)
> .framework("torch")
> .environment(env=RankingEnv, disable\_env\_checking=True, env\_config=train\_env\_config)
> )
> train\_config.model\['fcnet\_hiddens'\] = \[256, 256\]
> 
> search\_alg = OptunaSearch(metric="episode\_reward\_mean",mode="max")#what if metric=step\_reward??
> scheduler\_ = ASHAScheduler(metric="episode\_reward\_mean",mode="max",max\_t=5,grace\_period=1,reduction\_factor=2)#max\_t: The maximum budget for each trial(hyperparameter), in seconds. grace\_period: The number of seconds to wait before terminating a trial that has not reported any results.
> \# wandb\_callback = WandbLoggerCallback(project="Ray Tune Trial Run",log\_config=True,save\_checkpoints=True)
> 
> drl\_agent = DRLlibv2(
> trainable="PPO",
> # train\_env = RankingEnv,
> run\_name = "PPO\_TRAIN",
> local\_dir = "/content/PPO\_TRAIN",
> params = train\_config.to\_dict(),
> num\_samples = 1,#Number of samples of hyperparameters config to run
> training\_iterations=5,
> checkpoint\_freq=5,
> # scheduler\_=scheduler\_,
> search\_alg=search\_alg,
> metric = "episode\_reward\_mean",
> mode = "max"
> # callbacks=\[wandb\_callback\]
> )
> 
> res = drl\_agent.train\_tune\_model()
> results\_df, best\_result = drl\_agent.infer\_results()
> \`\`\`
> 
> \`\`\`2023-09-22 12:20:05,276	INFO worker.py:1633 -- Started a local Ray instance. View the dashboard at 127.0.0.1:8265 
> 2023-09-22 12:20:08,561	INFO tune.py:654 -- \[output\] This will use the new output engine with verbosity 2. To disable the new output and use the legacy output engine, set the environment variable RAY\_AIR\_NEW\_OUTPUT=0. For more information, please see https://github.com/ray-project/ray/issues/36949
> 2023-09-22 12:20:08,640	WARNING deprecation.py:50 -- DeprecationWarning: \`build\_tf\_policy\` has been deprecated. This will raise an error in the future!
> 2023-09-22 12:20:08,650	WARNING deprecation.py:50 -- DeprecationWarning: \`build\_policy\_class\` has been deprecated. This will raise an error in the future!
> 2023-09-22 12:20:08,737	WARNING algorithm\_config.py:2578 -- Setting \`exploration\_config={}\` because you set \`\_enable\_rl\_module\_api=True\`. When RLModule API are enabled, exploration\_config can not be set. If you want to implement custom exploration behaviour, please modify the \`forward\_exploration\` method of the RLModule at hand. On configs that have a default exploration config, this must be done with \`config.exploration\_config={}\`.
> /usr/local/lib/python3.10/dist-packages/gymnasium/spaces/box.py:130: UserWarning: WARN: Box bound precision lowered by casting to float32
> gym.logger.warn(f"Box bound precision lowered by casting to {self.dtype}")
> /usr/local/lib/python3.10/dist-packages/gymnasium/utils/passive\_env\_checker.py:164: UserWarning: WARN: The obs returned by the \`reset()\` method was expecting numpy array dtype to be float32, actual type: float64
> logger.warn(
> /usr/local/lib/python3.10/dist-packages/gymnasium/utils/passive\_env\_checker.py:188: UserWarning: WARN: The obs returned by the \`reset()\` method is not within the observation space.
> logger.warn(f"{pre} is not within the observation space.")
> 2023-09-22 12:20:08,838	WARNING algorithm\_config.py:2578 -- Setting \`exploration\_config={}\` because you set \`\_enable\_rl\_module\_api=True\`. When RLModule API are enabled, exploration\_config can not be set. If you want to implement custom exploration behaviour, please modify the \`forward\_exploration\` method of the RLModule at hand. On configs that have a default exploration config, this must be done with \`config.exploration\_config={}\`.
> \[I 2023-09-22 12:20:08,896\] A new study created in memory with name: optuna
> 2023-09-22 12:20:08,944	WARNING algorithm\_config.py:2578 -- Setting \`exploration\_config={}\` because you set \`\_enable\_rl\_module\_api=True\`. When RLModule API are enabled, exploration\_config can not be set. If you want to implement custom exploration behaviour, please modify the \`forward\_exploration\` method of the RLModule at hand. On configs that have a default exploration config, this must be done with \`config.exploration\_config={}\`.
> +--------------------------------------------------+
> | Configuration for experiment PPO\_TRAIN |
> +--------------------------------------------------+
> | Search algorithm SearchGenerator |
> | Scheduler FIFOScheduler |
> | Number of trials 1 |
> +--------------------------------------------------+
> 
> View detailed results here: /content/PPO\_TRAIN/PPO\_TRAIN
> To visualize your results with TensorBoard, run: \`tensorboard --logdir /root/ray\_results/PPO\_TRAIN\`
> 
> Trial status: 1 PENDING
> Current time: 2023-09-22 12:20:09. Total running time: 0s
> Logical resource usage: 0/2 CPUs, 0/0 GPUs
> +------------------------------------------------------------------------------------------------------+
> | Trial name status lr sgd\_minibatch\_size entropy\_coeff lambda |
> +------------------------------------------------------------------------------------------------------+
> | PPO\_RankingEnv\_8689c516 PENDING 0.000166759 256 4.82692e-07 0.1 |
> +------------------------------------------------------------------------------------------------------+
> (pid=2755) /usr/local/lib/python3.10/dist-packages/tensorflow\_probability/python/\_\_init\_\_.py:57: DeprecationWarning: distutils Version classes are deprecated. Use packaging.version instead.
> (pid=2755) if (distutils.version.LooseVersion(tf.\_\_version\_\_) \<
> (pid=2755) DeprecationWarning: \`DirectStepOptimizer\` has been deprecated. This will raise an error in the future!
> (pid=2755) /usr/local/lib/python3.10/dist-packages/google/rpc/\_\_init\_\_.py:20: DeprecationWarning: Deprecated call to \`pkg\_resources.declare\_namespace('google.rpc')\`.
> (pid=2755) Implementing implicit namespace packages (as specified in PEP 420) is preferred to \`pkg\_resources.declare\_namespace\`. See https://setuptools.pypa.io/en/latest/references/keywords.html#keyword-namespace-packages
> (pid=2755) pkg\_resources.declare\_namespace(\_\_name\_\_)
> (pid=2755) /usr/local/lib/python3.10/dist-packages/pkg\_resources/\_\_init\_\_.py:2349: DeprecationWarning: Deprecated call to \`pkg\_resources.declare\_namespace('google')\`.
> (pid=2755) Implementing implicit namespace packages (as specified in PEP 420) is preferred to \`pkg\_resources.declare\_namespace\`. See https://setuptools.pypa.io/en/latest/references/keywords.html#keyword-namespace-packages
> (pid=2755) declare\_namespace(parent)
> (PPO pid=2755) 2023-09-22 12:20:18,560	WARNING algorithm\_config.py:2578 -- Setting \`exploration\_config={}\` because you set \`\_enable\_rl\_module\_api=True\`. When RLModule API are enabled, exploration\_config can not be set. If you want to implement custom exploration behaviour, please modify the \`forward\_exploration\` method of the RLModule at hand. On configs that have a default exploration config, this must be done with \`config.exploration\_config={}\`.
> (PPO pid=2755) 2023-09-22 12:20:18,561	WARNING algorithm\_config.py:672 -- Cannot create PPOConfig from given \`config\_dict\`! Property \_\_stdout\_file\_\_ not supported.
> (pid=2819) /usr/local/lib/python3.10/dist-packages/tensorflow\_probability/python/\_\_init\_\_.py:57: DeprecationWarning: distutils Version classes are deprecated. Use packaging.version instead.
> (pid=2819) if (distutils.version.LooseVersion(tf.\_\_version\_\_) \<
> (pid=2819) DeprecationWarning: \`DirectStepOptimizer\` has been deprecated. This will raise an error in the future!
> (pid=2819) /usr/local/lib/python3.10/dist-packages/google/rpc/\_\_init\_\_.py:20: DeprecationWarning: Deprecated call to \`pkg\_resources.declare\_namespace('google.rpc')\`.
> (pid=2819) Implementing implicit namespace packages (as specified in PEP 420) is preferred to \`pkg\_resources.declare\_namespace\`. See https://setuptools.pypa.io/en/latest/references/keywords.html#keyword-namespace-packages
> (pid=2819) pkg\_resources.declare\_namespace(\_\_name\_\_)
> (pid=2819) /usr/local/lib/python3.10/dist-packages/pkg\_resources/\_\_init\_\_.py:2349: DeprecationWarning: Deprecated call to \`pkg\_resources.declare\_namespace('google')\`.
> (pid=2819) Implementing implicit namespace packages (as specified in PEP 420) is preferred to \`pkg\_resources.declare\_namespace\`. See https://setuptools.pypa.io/en/latest/references/keywords.html#keyword-namespace-packages
> (pid=2819) declare\_namespace(parent)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,813	DEBUG rollout\_worker.py:1761 -- Creating policy for default\_policy
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,813	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,814	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,814	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,815	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,815	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,815	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,816	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,816	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,816	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,816	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,817	DEBUG catalog.py:789 -- Created preprocessor \<ray.rllib.models.preprocessors.DictFlatteningPreprocessor object at 0x7e19968ff280\>: Dict('AAVEUSDT': Box(-inf, inf, (25,), float64), 'AVAXUSDT': Box(-inf, inf, (25,), float64), 'BTCUSDT': Box(-inf, inf, (25,), float64), 'ETHUSDT': Box(-inf, inf, (25,), float64), 'LINKUSDT': Box(-inf, inf, (25,), float64), 'LTCUSDT': Box(-inf, inf, (25,), float64), 'MATICUSDT': Box(-inf, inf, (25,), float64), 'NEARUSDT': Box(-inf, inf, (25,), float64), 'SOLUSDT': Box(-inf, inf, (25,), float64), 'UNIUSDT': Box(-inf, inf, (25,), float64)) -\> (250,)
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,820	WARNING algorithm\_config.py:2578 -- Setting \`exploration\_config={}\` because you set \`\_enable\_rl\_module\_api=True\`. When RLModule API are enabled, exploration\_config can not be set. If you want to implement custom exploration behaviour, please modify the \`forward\_exploration\` method of the RLModule at hand. On configs that have a default exploration config, this must be done with \`config.exploration\_config={}\`.
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,885	INFO policy.py:1294 -- Policy (worker=1) running on CPU.
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,886	INFO torch\_policy\_v2.py:113 -- Found 0 visible cuda devices.
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,886	WARNING deprecation.py:50 -- DeprecationWarning: \`ValueNetworkMixin\` has been deprecated. This will raise an error in the future!
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,886	WARNING deprecation.py:50 -- DeprecationWarning: \`LearningRateSchedule\` has been deprecated. This will raise an error in the future!
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,886	WARNING deprecation.py:50 -- DeprecationWarning: \`EntropyCoeffSchedule\` has been deprecated. This will raise an error in the future!
> (RolloutWorker pid=2819) 2023-09-22 12:20:29,886	WARNING deprecation.py:50 -- DeprecationWarning: \`KLCoeffMixin\` has been deprecated. This will raise an error in the future!
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,074	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,075	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,075	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,075	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,075	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,075	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,076	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,076	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,076	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,076	DEBUG preprocessors.py:304 -- Creating sub-preprocessor for Box(-inf, inf, (25,), float64)
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,077	INFO util.py:118 -- Using connectors:
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,077	INFO util.py:119 -- AgentConnectorPipeline
> (RolloutWorker pid=2819) ObsPreprocessorConnector
> (RolloutWorker pid=2819) StateBufferConnector
> (RolloutWorker pid=2819) ViewRequirementAgentConnector
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,078	INFO util.py:120 -- ActionConnectorPipeline
> (RolloutWorker pid=2819) ConvertToNumpyConnector
> (RolloutWorker pid=2819) NormalizeActionsConnector
> (RolloutWorker pid=2819) ImmutableActionsConnector
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,078	DEBUG rollout\_worker.py:645 -- Created rollout worker with env \<ray.rllib.env.vector\_env.VectorEnvWrapper object at 0x7e19969901f0\> (\<RankingEnv instance\>), policies \<PolicyMap lru-caching-capacity=100 policy-IDs=\['default\_policy'\]\>
> (PPO pid=2755) 2023-09-22 12:20:30,116	INFO worker\_set.py:297 -- Inferred observation/action spaces from remote worker (local worker has no env): {'default\_policy': (Dict('AAVEUSDT': Box(-inf, inf, (25,), float64), 'AVAXUSDT': Box(-inf, inf, (25,), float64), 'BTCUSDT': Box(-inf, inf, (25,), float64), 'ETHUSDT': Box(-inf, inf, (25,), float64), 'LINKUSDT': Box(-inf, inf, (25,), float64), 'LTCUSDT': Box(-inf, inf, (25,), float64), 'MATICUSDT': Box(-inf, inf, (25,), float64), 'NEARUSDT': Box(-inf, inf, (25,), float64), 'SOLUSDT': Box(-inf, inf, (25,), float64), 'UNIUSDT': Box(-inf, inf, (25,), float64)), Dict('AAVEUSDT': Box(1.0, 10.0, (1,), float64), 'AVAXUSDT': Box(1.0, 10.0, (1,), float64), 'BTCUSDT': Box(1.0, 10.0, (1,), float64), 'ETHUSDT': Box(1.0, 10.0, (1,), float64), 'LINKUSDT': Box(1.0, 10.0, (1,), float64), 'LTCUSDT': Box(1.0, 10.0, (1,), float64), 'MATICUSDT': Box(1.0, 10.0, (1,), float64), 'NEARUSDT': Box(1.0, 10.0, (1,), float64), 'SOLUSDT': Box(1.0, 10.0, (1,), float64), 'UNIUSDT': Box(1.0, 10.0, (1,), float64))), '\_\_env\_\_': (Dict('AAVEUSDT': Box(-inf, inf, (25,), float64), 'AVAXUSDT': Box(-inf, inf, (25,), float64), 'BTCUSDT': Box(-inf, inf, (25,), float64), 'ETHUSDT': Box(-inf, inf, (25,), float64), 'LINKUSDT': Box(-inf, inf, (25,), float64), 'LTCUSDT': Box(-inf, inf, (25,), float64), 'MATICUSDT': Box(-inf, inf, (25,), float64), 'NEARUSDT': Box(-inf, inf, (25,), float64), 'SOLUSDT': Box(-inf, inf, (25,), float64), 'UNIUSDT': Box(-inf, inf, (25,), float64)), Dict('AAVEUSDT': Box(1.0, 10.0, (1,), float64), 'AVAXUSDT': Box(1.0, 10.0, (1,), float64), 'BTCUSDT': Box(1.0, 10.0, (1,), float64), 'ETHUSDT': Box(1.0, 10.0, (1,), float64), 'LINKUSDT': Box(1.0, 10.0, (1,), float64), 'LTCUSDT': Box(1.0, 10.0, (1,), float64), 'MATICUSDT': Box(1.0, 10.0, (1,), float64), 'NEARUSDT': Box(1.0, 10.0, (1,), float64), 'SOLUSDT': Box(1.0, 10.0, (1,), float64), 'UNIUSDT': Box(1.0, 10.0, (1,), float64)))}
> (PPO pid=2755) 2023-09-22 12:20:30,151	INFO policy.py:1294 -- Policy (worker=local) running on CPU.
> (PPO pid=2755) 2023-09-22 12:20:30,173	INFO rollout\_worker.py:1742 -- Built policy map: \<PolicyMap lru-caching-capacity=100 policy-IDs=\['default\_policy'\]\>
> (PPO pid=2755) 2023-09-22 12:20:30,173	INFO rollout\_worker.py:1743 -- Built preprocessor map: {'default\_policy': None}
> (PPO pid=2755) 2023-09-22 12:20:30,173	INFO rollout\_worker.py:550 -- Built filter map: defaultdict(\<class 'ray.rllib.utils.filter.NoFilter'\>, {})
> (PPO pid=2755) 2023-09-22 12:20:30,173	DEBUG rollout\_worker.py:645 -- Created rollout worker with env None (None), policies \<PolicyMap lru-caching-capacity=100 policy-IDs=\['default\_policy'\]\>
> 
> Trial PPO\_RankingEnv\_8689c516 started with configuration:
> +---------------------------------------------------------------------------+
> | Trial PPO\_RankingEnv\_8689c516 config |
> +---------------------------------------------------------------------------+
> | \_AlgorithmConfig\_\_prior\_exploration\_config/type StochasticSampling |
> | \_disable\_action\_flattening False |
> | \_disable\_execution\_plan\_api True |
> | \_disable\_initialize\_loss\_from\_dummy\_batch False |
> | \_disable\_preprocessor\_api False |
> | \_enable\_learner\_api True |
> | \_enable\_rl\_module\_api True |
> | \_fake\_gpus False |
> | \_is\_atari |
> | \_learner\_class |
> | \_tf\_policy\_handles\_more\_than\_one\_loss False |
> | action\_mask\_key action\_mask |
> | action\_space |
> | actions\_in\_input\_normalized False |
> | always\_attach\_evaluation\_results False |
> | auto\_wrap\_old\_gym\_envs True |
> | batch\_mode truncate\_episodes |
> | callbacks ...efaultCallbacks'\> |
> | checkpoint\_trainable\_policies\_only False |
> | clip\_actions False |
> | clip\_param 0.3 |
> | clip\_rewards |
> | compress\_observations False |
> | count\_steps\_by env\_steps |
> | create\_env\_on\_driver False |
> | custom\_eval\_function |
> | delay\_between\_worker\_restarts\_s 60. |
> | disable\_env\_checking True |
> | eager\_max\_retraces 20 |
> | eager\_tracing True |
> | enable\_async\_evaluation False |
> | enable\_connectors True |
> | enable\_tf1\_exec\_eagerly False |
> | entropy\_coeff 0. |
> | entropy\_coeff\_schedule |
> | env ...in\_\_.RankingEnv'\> |
> | env\_config/df ...ows x 29 columns\] |
> | env\_runner\_cls |
> | env\_task\_fn |
> | evaluation\_config |
> | evaluation\_duration 10 |
> | evaluation\_duration\_unit episodes |
> | evaluation\_interval |
> | evaluation\_num\_workers 0 |
> | evaluation\_parallel\_to\_training False |
> | evaluation\_sample\_timeout\_s 180. |
> | explore True |
> | export\_native\_model\_files False |
> | fake\_sampler False |
> | framework torch |
> | gamma 0.99 |
> | grad\_clip |
> | grad\_clip\_by global\_norm |
> | ignore\_worker\_failures False |
> | in\_evaluation False |
> | input sampler |
> | keep\_per\_episode\_custom\_metrics False |
> | kl\_coeff 0.2 |
> | kl\_target 0.01 |
> | lambda 0.1 |
> | local\_gpu\_idx 0 |
> | local\_tf\_session\_args/inter\_op\_parallelism\_threads 8 |
> | local\_tf\_session\_args/intra\_op\_parallelism\_threads 8 |
> | log\_level DEBUG |
> | log\_sys\_usage True |
> | logger\_config |
> | logger\_creator |
> | lr 0.00017 |
> | lr\_schedule |
> | max\_num\_worker\_restarts 1000 |
> | max\_requests\_in\_flight\_per\_sampler\_worker 2 |
> | metrics\_episode\_collection\_timeout\_s 60. |
> | metrics\_num\_episodes\_for\_smoothing 100 |
> | min\_sample\_timesteps\_per\_iteration 0 |
> | min\_time\_s\_per\_iteration |
> | min\_train\_timesteps\_per\_iteration 0 |
> | model/\_disable\_action\_flattening False |
> | model/\_disable\_preprocessor\_api False |
> | model/\_time\_major False |
> | model/\_use\_default\_native\_models -1 |
> | model/always\_check\_shapes False |
> | model/attention\_dim 64 |
> | model/attention\_head\_dim 32 |
> | model/attention\_init\_gru\_gate\_bias 2.0 |
> | model/attention\_memory\_inference 50 |
> | model/attention\_memory\_training 50 |
> | model/attention\_num\_heads 1 |
> | model/attention\_num\_transformer\_units 1 |
> | model/attention\_position\_wise\_mlp\_dim 32 |
> | model/attention\_use\_n\_prev\_actions 0 |
> | model/attention\_use\_n\_prev\_rewards 0 |
> | model/conv\_activation relu |
> | model/conv\_filters |
> | model/custom\_action\_dist |
> | model/custom\_model |
> | model/custom\_preprocessor |
> | model/dim 84 |
> | model/encoder\_latent\_dim |
> | model/fcnet\_activation tanh |
> | model/fcnet\_hiddens \[256, 256\] |
> | model/framestack True |
> | model/free\_log\_std False |
> | model/grayscale False |
> | model/lstm\_cell\_size 256 |
> | model/lstm\_use\_prev\_action False |
> | model/lstm\_use\_prev\_action\_reward -1 |
> | model/lstm\_use\_prev\_reward False |
> | model/max\_seq\_len 20 |
> | model/no\_final\_linear False |
> | model/post\_fcnet\_activation relu |
> | model/post\_fcnet\_hiddens \[\] |
> | model/use\_attention False |
> | model/use\_lstm False |
> | model/vf\_share\_layers False |
> | model/zero\_mean True |
> | normalize\_actions True |
> | num\_consecutive\_worker\_failures\_tolerance 100 |
> | num\_cpus\_for\_driver 1 |
> | num\_cpus\_per\_learner\_worker 1 |
> | num\_cpus\_per\_worker 1 |
> | num\_envs\_per\_worker 1 |
> | num\_gpus 0 |
> | num\_gpus\_per\_learner\_worker 0 |
> | num\_gpus\_per\_worker 0 |
> | num\_learner\_workers 0 |
> | num\_sgd\_iter 30 |
> | num\_workers 1 |
> | observation\_filter NoFilter |
> | observation\_fn |
> | observation\_space |
> | offline\_sampling False |
> | ope\_split\_batch\_by\_episode True |
> | output |
> | output\_compress\_columns \['obs', 'new\_obs'\] |
> | output\_max\_file\_size 67108864 |
> | placement\_strategy PACK |
> | policies/default\_policy ...None, None, None) |
> | policies\_to\_train |
> | policy\_map\_cache -1 |
> | policy\_map\_capacity 100 |
> | policy\_mapping\_fn ...t 0x7ad1d496b910\> |
> | policy\_states\_are\_swappable False |
> | postprocess\_inputs False |
> | preprocessor\_pref deepmind |
> | recreate\_failed\_workers False |
> | remote\_env\_batch\_wait\_ms 0 |
> | remote\_worker\_envs False |
> | render\_env False |
> | replay\_sequence\_length |
> | restart\_failed\_sub\_environments False |
> | rl\_module\_spec |
> | rollout\_fragment\_length auto |
> | sample\_async False |
> | sample\_collector ...leListCollector'\> |
> | sampler\_perf\_stats\_ema\_coef |
> | seed 1234 |
> | sgd\_minibatch\_size 256 |
> | shuffle\_buffer\_size 0 |
> | shuffle\_sequences True |
> | simple\_optimizer -1 |
> | sync\_filters\_on\_rollout\_workers\_timeout\_s 60. |
> | synchronize\_filters -1 |
> | tf\_session\_args/allow\_soft\_placement True |
> | tf\_session\_args/device\_count/CPU 1 |
> | tf\_session\_args/gpu\_options/allow\_growth True |
> | tf\_session\_args/inter\_op\_parallelism\_threads 2 |
> | tf\_session\_args/intra\_op\_parallelism\_threads 2 |
> | tf\_session\_args/log\_device\_placement False |
> | torch\_compile\_learner False |
> | torch\_compile\_learner\_dynamo\_backend inductor |
> | torch\_compile\_learner\_dynamo\_mode |
> | torch\_compile\_learner\_what\_to\_compile ...ile.FORWARD\_TRAIN |
> | torch\_compile\_worker False |
> | torch\_compile\_worker\_dynamo\_backend onnxrt |
> | torch\_compile\_worker\_dynamo\_mode |
> | train\_batch\_size 4000 |
> | update\_worker\_filter\_stats True |
> | use\_critic True |
> | use\_gae True |
> | use\_kl\_loss True |
> | use\_worker\_filter\_stats True |
> | validate\_workers\_after\_construction True |
> | vf\_clip\_param 10. |
> | vf\_loss\_coeff 1. |
> | vf\_share\_layers -1 |
> | worker\_cls -1 |
> | worker\_health\_probe\_timeout\_s 60 |
> | worker\_restore\_timeout\_s 1800 |
> +---------------------------------------------------------------------------+
> (PPO pid=2755) Trainable.setup took 11.358 seconds. If your trainable is slow to initialize, consider setting reuse\_actors=True to reduce actor creation overheads.
> (PPO pid=2755) Install gputil for GPU system monitoring.
> (RolloutWorker pid=2819) 2023-09-22 12:20:30,361	INFO rollout\_worker.py:690 -- Generating sample batch of size 4000
> 2023-09-22 12:20:30,496	ERROR tune\_controller.py:1502 -- Trial task failed for trial PPO\_RankingEnv\_8689c516
> Traceback (most recent call last):
> File "/usr/local/lib/python3.10/dist-packages/ray/air/execution/\_internal/event\_manager.py", line 110, in resolve\_future
> result = ray.get(future)
> File "/usr/local/lib/python3.10/dist-packages/ray/\_private/auto\_init\_hook.py", line 24, in auto\_init\_wrapper
> return fn(\*args, \*\*kwargs)
> File "/usr/local/lib/python3.10/dist-packages/ray/\_private/client\_mode\_hook.py", line 103, in wrapper
> return func(\*args, \*\*kwargs)
> File "/usr/local/lib/python3.10/dist-packages/ray/\_private/worker.py", line 2547, in get
> raise value.as\_instanceof\_cause()
> ray.exceptions.RayTaskError(AttributeError): ray::PPO.train() (pid=2755, ip=172.28.0.12, actor\_id=334c1673a75a460bc3f3ad2101000000, repr=PPO)
> File "/usr/local/lib/python3.10/dist-packages/ray/tune/trainable/trainable.py", line 400, in train
> raise skipped from exception\_cause(skipped)
> File "/usr/local/lib/python3.10/dist-packages/ray/tune/trainable/trainable.py", line 397, in train
> result = self.step()
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/algorithms/algorithm.py", line 853, in step
> results, train\_iter\_ctx = self.\_run\_one\_training\_iteration()
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/algorithms/algorithm.py", line 2838, in \_run\_one\_training\_iteration
> results = self.training\_step()
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/algorithms/ppo/ppo.py", line 429, in training\_step
> train\_batch = synchronous\_parallel\_sample(
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/execution/rollout\_ops.py", line 85, in synchronous\_parallel\_sample
> sample\_batches = worker\_set.foreach\_worker(
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/worker\_set.py", line 680, in foreach\_worker
> handle\_remote\_call\_result\_errors(remote\_results, self.\_ignore\_worker\_failures)
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/worker\_set.py", line 76, in handle\_remote\_call\_result\_errors
> raise r.get()
> ray.exceptions.RayTaskError(AttributeError): ray::RolloutWorker.apply() (pid=2819, ip=172.28.0.12, actor\_id=fc63482e3750d98323a830f401000000, repr=\<ray.rllib.evaluation.rollout\_worker.RolloutWorker object at 0x7e19975be8c0\>)
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/utils/actor\_manager.py", line 185, in apply
> raise e
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/utils/actor\_manager.py", line 176, in apply
> return func(self, \*args, \*\*kwargs)
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/execution/rollout\_ops.py", line 86, in \<lambda\>
> lambda w: w.sample(), local\_worker=False, healthy\_only=True
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/rollout\_worker.py", line 696, in sample
> batches = \[self.input\_reader.next()\]
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/sampler.py", line 92, in next
> batches = \[self.get\_data()\]
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/sampler.py", line 277, in get\_data
> item = next(self.\_env\_runner)
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/env\_runner\_v2.py", line 344, in run
> outputs = self.step()
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/env\_runner\_v2.py", line 370, in step
> active\_envs, to\_eval, outputs = self.\_process\_observations(
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/env\_runner\_v2.py", line 637, in \_process\_observations
> processed = policy.agent\_connectors(acd\_list)
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/connectors/agent/pipeline.py", line 41, in \_\_call\_\_
> ret = c(ret)
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/connectors/connector.py", line 265, in \_\_call\_\_
> return \[self.transform(d) for d in acd\_list\]
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/connectors/connector.py", line 265, in \<listcomp\>
> return \[self.transform(d) for d in acd\_list\]
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/connectors/agent/obs\_preproc.py", line 58, in transform
> d\[SampleBatch.NEXT\_OBS\] = self.\_preprocessor.transform(
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/models/preprocessors.py", line 319, in transform
> self.write(observation, array, 0)
> File "/usr/local/lib/python3.10/dist-packages/ray/rllib/models/preprocessors.py", line 325, in write
> observation = OrderedDict(sorted(observation.items()))
> AttributeError: 'NoneType' object has no attribute 'items'
> (PPO pid=2755) 2023-09-22 12:20:30,490	ERROR actor\_manager.py:500 -- Ray error, taking actor 1 out of service. ray::RolloutWorker.apply() (pid=2819, ip=172.28.0.12, actor\_id=fc63482e3750d98323a830f401000000, repr=\<ray.rllib.evaluation.rollout\_worker.RolloutWorker object at 0x7e19975be8c0\>)
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/utils/actor\_manager.py", line 185, in apply
> (PPO pid=2755) raise e
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/utils/actor\_manager.py", line 176, in apply
> (PPO pid=2755) return func(self, \*args, \*\*kwargs)
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/execution/rollout\_ops.py", line 86, in \<lambda\>
> (PPO pid=2755) lambda w: w.sample(), local\_worker=False, healthy\_only=True
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/rollout\_worker.py", line 696, in sample
> (PPO pid=2755) batches = \[self.input\_reader.next()\]
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/sampler.py", line 92, in next
> (PPO pid=2755) batches = \[self.get\_data()\]
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/sampler.py", line 277, in get\_data
> (PPO pid=2755) item = next(self.\_env\_runner)
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/env\_runner\_v2.py", line 344, in run
> (PPO pid=2755) outputs = self.step()
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/env\_runner\_v2.py", line 370, in step
> (PPO pid=2755) active\_envs, to\_eval, outputs = self.\_process\_observations(
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/evaluation/env\_runner\_v2.py", line 637, in \_process\_observations
> (PPO pid=2755) processed = policy.agent\_connectors(acd\_list)
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/connectors/agent/pipeline.py", line 41, in \_\_call\_\_
> (PPO pid=2755) ret = c(ret)
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/connectors/connector.py", line 265, in \_\_call\_\_
> (PPO pid=2755) return \[self.transform(d) for d in acd\_list\]
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/connectors/connector.py", line 265, in \<listcomp\>
> (PPO pid=2755) return \[self.transform(d) for d in acd\_list\]
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/connectors/agent/obs\_preproc.py", line 58, in transform
> (PPO pid=2755) d\[SampleBatch.NEXT\_OBS\] = self.\_preprocessor.transform(
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/models/preprocessors.py", line 319, in transform
> (PPO pid=2755) self.write(observation, array, 0)
> (PPO pid=2755) File "/usr/local/lib/python3.10/dist-packages/ray/rllib/models/preprocessors.py", line 325, in write
> (PPO pid=2755) observation = OrderedDict(sorted(observation.items()))
> (PPO pid=2755) AttributeError: 'NoneType' object has no attribute 'items'
> 2023-09-22 12:20:30,568	WARNING experiment\_state.py:371 -- Experiment checkpoint syncing has been triggered multiple times in the last 30.0 seconds. A sync will be triggered whenever a trial has checkpointed more than \`num\_to\_keep\` times since last sync or if 300 seconds have passed since last sync. If you have set \`num\_to\_keep\` in your \`CheckpointConfig\`, consider increasing the checkpoint frequency or keeping more checkpoints. You can supress this warning by changing the \`TUNE\_WARN\_EXCESSIVE\_EXPERIMENT\_CHECKPOINT\_SYNC\_THRESHOLD\_S\` environment variable.
> 
> Trial PPO\_RankingEnv\_8689c516 errored after 0 iterations at 2023-09-22 12:20:30. Total running time: 21s
> Error file: /root/ray\_results/PPO\_TRAIN/PPO\_RankingEnv\_8689c516\_1\_type=StochasticSampling,disable\_action\_flattening=False,disable\_execution\_plan\_api=True,disable\_initiali\_2023-09-22\_12-20-08/error.txt
> 
> Trial status: 1 ERROR
> Current time: 2023-09-22 12:20:30. Total running time: 21s
> Logical resource usage: 2.0/2 CPUs, 0/0 GPUs
> +------------------------------------------------------------------------------------------------------+
> | Trial name status lr sgd\_minibatch\_size entropy\_coeff lambda |
> +------------------------------------------------------------------------------------------------------+
> | PPO\_RankingEnv\_8689c516 ERROR 0.000166759 256 4.82692e-07 0.1 |
> +------------------------------------------------------------------------------------------------------+
> 
> Number of errored trials: 1
> +---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
> | Trial name # failures error file |
> +---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
> | PPO\_RankingEnv\_8689c516 1 /root/ray\_results/PPO\_TRAIN/PPO\_RankingEnv\_8689c516\_1\_type=StochasticSampling,disable\_action\_flattening=False,disable\_execution\_plan\_api=True,disable\_initiali\_2023-09-22\_12-20-08/error.txt |
> +---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
> 2023-09-22 12:20:30,939	ERROR tune.py:1139 -- Trials did not complete: \[PPO\_RankingEnv\_8689c516\]
> \`\`\`
> 
> \### Issue Severity
> 
> High: It blocks me from completing my task.

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [September 24, 2023, 12:01am UTC](https://discuss.ray.io/t/register-a-custom-environment-and-runing-ppotrainer-on-that-environment-not-working/6143/7 "2023-09-24T00:01:20Z")

</div>

Hi @fardinabbasi,

This is where your error is coming from  
`return self._get_obs() if not terminated else None`

You cannot return None. You will need to come up with some dummy observation to return.

---

<div class="post-metadata">

**Author:** ![fardinabbasi](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/fardinabbasi/32/5085_2.png) [@fardinabbasi](https://discuss.ray.io/u/fardinabbasi)\
**Post date:** [September 24, 2023, 8:50am UTC](https://discuss.ray.io/t/register-a-custom-environment-and-runing-ppotrainer-on-that-environment-not-working/6143/8 "2023-09-24T08:50:46Z")

</div>

Hi @mannyv,

You’re correct; I’ve made the change to `self._get_obs() if not terminated else self.observation_space.sample()`. However, I’m still encountering the error message: `Cannot create PPOConfig from the given config_dict! Property stdout_file is not supported.`

I’m new to RLlib, and I’m uncertain whether this issue is due to not registering env and pass RankingEnv directly as `train_config["env"] = RankingEnv`, or if it’s related to how I retrieve my configuration within RankingEnv.
