# Issue with multiple environments training one PPO policy

**URL:** <https://discuss.ray.io/t/issue-with-multiple-environments-training-one-ppo-policy/22554>\
**Category:** RLlib\
**Created:** [May 25, 2025, 1:24am UTC](https://discuss.ray.io/t/issue-with-multiple-environments-training-one-ppo-policy/22554 "2025-05-25T01:24:30Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![blackpanther](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/blackpanther/32/5003_2.png) [@blackpanther](https://discuss.ray.io/u/blackpanther)\
**Post date:** [May 25, 2025, 1:24am UTC](https://discuss.ray.io/t/issue-with-multiple-environments-training-one-ppo-policy/22554/1 "2025-05-25T01:24:30Z")

</div>

**1. Severity of the issue: (select one)**  
 High: Completely blocks me.

**2. Environment:**

- Ray version: rllib 2.4.0
- Python version: 3.6
- OS: Ubuntu 18
- Cloud/Infrastructure: local running in a docker container
- Other libs/tools (if relevant):

**3. What happened vs. what you expected:**

- Expected: 4 parallel workers (one for each env) training a single PPO policy
- Actual: 2 workers created but only one environment/worker is used for training

ray.init()  
register\_my\_stuff()

def env\_creator(env\_config):  
worker\_index = env\_config.worker\_index

```
env_id = (worker_index - 1) % 4  

print(f"[ENV] Worker {worker_index} assigned env_id: {env_id}")
if env_id == 0:
    return OhlcvEnv1()
elif env_id == 1:
    return OhlcvEnv2()
elif env_id == 2:
    return OhlcvEnv3()
else:
    return OhlcvEnv()

```

register\_env(“MultiParallelEnv”, lambda config: env\_creator(config))

config = (  
PPOConfig()  
.environment(

```
    env="MultiParallelEnv",##
    env_config={"env_id": tune.grid_search([0, 1, 2, 3])}

```

)  
.resources(  
num\_gpus=1, # use 1 GPU for the local (training) worker  
num\_cpus\_per\_worker=1

```
)
.framework("tf2") 
.rollouts(num_rollout_workers=4,##
          num_envs_per_worker=1)##
.training(
    model={
        "custom_model": "my_model"
        
    },
    gamma=0.99,
    lr=1e-4, # Actor LR
    train_batch_size=6000,

    entropy_coeff=0.02,  
    entropy_coeff_schedule=[
         [0, 0.02],   
         [1000, 0.001] 
     ],
)

.evaluation(
    evaluation_interval=1,             
    evaluation_duration=1,             
    evaluation_duration_unit="episodes",  
    evaluation_parallel_to_training=False, 
    evaluation_num_workers=1
)

```

)

algo = config.build()

max\_iterations = 100000  
for i in range(max\_iterations):  
result = algo.train()  
eval\_results = algo.evaluate()

ray.shutdown()
