# PPO.train incorrect result

**URL:** <https://discuss.ray.io/t/ppo-train-incorrect-result/10303>\
**Category:** RLlib\
**Created:** [April 19, 2023, 8:27am UTC](https://discuss.ray.io/t/ppo-train-incorrect-result/10303 "2023-04-19T08:27:38Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![dan\_phi](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/dan_phi/32/4258_2.png) [@dan\_phi](https://discuss.ray.io/u/dan_phi)\
**Post date:** [April 19, 2023, 8:27am UTC](https://discuss.ray.io/t/ppo-train-incorrect-result/10303/1 "2023-04-19T08:27:39Z")

</div>

PPO.train incorrect result

PPO.train returns the number of iterations  
in episode\_reward\_max, episode\_reward\_min

Am I doing wrong or is it a bug?

if it’s a bug, how do I install RAY 2.2.0?

```auto
pip install -U "ray[default, tune, rllib, air, serve]" # 2.3.1
pip install tensorflow

class MockEnv(gymnasium.Env):

    def __init__ (self, env_config):
        self.episode_length = env_config["episode_length"]
        self.config = env_config
        self.i = 0
        self.observation_space = gymnasium.spaces.Discrete(20)
        self.action_space = gymnasium.spaces.Discrete(2)

    def reset(self, *, seed=None, options=None):
        self.i = 0
        return 0, {}

    def step(self, action):
        self.i += 1
        mock_obs = 12
        mock_reward = 9.0

        terminated = truncated = self.i >= self.episode_length
        return mock_obs, mock_reward, terminated, truncated, {}
		
if __name__ == ' __main__':
    # ray.init(
    # local_mode = True #local_mode
    # # , logging_level = 'DEBUG',
    # , ignore_reinit_error = True
    # # , num_cpus = 15
    # )

    def env_creator(env_config:EnvContext):
        env = MockEnv(env_config)
        return env

    env_id = "MockEnv01"
    register_env(env_id, env_creator)

    env_cfg = {
        'env_id': env_id,
        'env_creator': env_creator,
        'episode_length': 3
    }

    ppoconfig = PPOConfig()
    ppoconfig.disable_env_checking = True
    ppoconfig.auto_wrap_old_gym_envs = False
    ppoconfig.train_batch_size = 10 # for speed
    ppoconfig.sgd_minibatch_size = 5 #for speed

    ppoconfig.environment(env=env_id)
    ppoconfig.env_config = env_cfg

    ppoconfig.log_level = "WARNING" # "DEBUG"

    ppoconfig.ignore_worker_failures = True
    # ppoconfig.framework_str = "torch"
    ppoconfig.framework_str = "tf2"

    ppoconfig.lr = 8e-6
    ppoconfig.num_gpus = 0
    ppoconfig.lr_schedule = [
                [0, 1e-1],
                [int(1e2), 1e-2],
                [int(1e3), 1e-3],
                [int(1e4), 1e-4],
                [int(1e5), 1e-5],
                [int(1e6), 1e-6],
                [int(1e7), 1e-7]
            ]
    ppoconfig.clip_rewards = True
    ppoconfig.gamma = 0.99
    ppoconfig.vf_loss_coeff = 0.5
    ppoconfig.vf_share_layers = True
    ppoconfig.entropy_coeff = 0.01

    # ppoconfig.checkpoint_freq = 1000
    ppoconfig.keep_checkpoints_num = 3
    ppoconfig.verbose = 1
    ppoconfig.log_to_file = False

    algo = PPO(ppoconfig)
    result = algo.train()

    print(f'episode_length:{env_cfg["episode_length"]}')
    print(f'episode_reward_max: {result["episode_reward_max"]}, episode_reward_min: {result["episode_reward_min"]}')

```

---

<div class="post-metadata">

**Author:** ![Rohan138](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rohan138/32/1160_2.png) [@Rohan138](https://discuss.ray.io/u/Rohan138)\
**Post date:** [May 23, 2023, 12:25am UTC](https://discuss.ray.io/t/ppo-train-incorrect-result/10303/2 "2023-05-23T00:25:34Z")

</div>

clip\_rewards = True clips the reward at each timestep to -1.0, 0.0, or 1.0 based on its sign. [ray/algorithm\_config.py at master · ray-project/ray · GitHub](https://github.com/ray-project/ray/blob/master/rllib/algorithms/algorithm_config.py#L1341)

As a side note, I would reccommend using the AlgorithmConfig API to avoid issues like these, since it provides documentation for the arguments you pass in to the config object’s functions.
