# Num\_env & agent\_steps\_trained 0 even though steps sampled?

**URL:** <https://discuss.ray.io/t/num-env-agent-steps-trained-0-even-though-steps-sampled/11730>\
**Category:** RLlib\
**Created:** [August 9, 2023, 3:37pm UTC](https://discuss.ray.io/t/num-env-agent-steps-trained-0-even-though-steps-sampled/11730 "2023-08-09T15:37:40Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![JenHa](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jenha/32/4898_2.png) [@JenHa](https://discuss.ray.io/u/JenHa)\
**Post date:** [August 9, 2023, 3:37pm UTC](https://discuss.ray.io/t/num-env-agent-steps-trained-0-even-though-steps-sampled/11730/1 "2023-08-09T15:37:40Z")

</div>

High: It blocks me to complete my task.

Hi everyone -  
I am currently building a multi-agent RL setup using rllib with PPO - see the script below. The script generally runs for both tune and train but in either case the number of **trained** steps and gents remains zero throughout the process.  
I was wondering whether someone’s faced a similar issue?

\*\*‘counters’: \*\*  
**{‘num\_env\_steps\_sampled’: 3584,**  
\*\* ‘num\_env\_steps\_trained’: 0, \*\*  
\*\*‘num\_agent\_steps\_sampled’: 10752, \*\*  
**‘num\_agent\_steps\_trained’: 0},**

Current mains script:

import os  
import sys

import platform  
if platform.system() != “Linux”:  
if ‘SUMO\_HOME’ in os.environ:  
tools = os.path.join(os.environ[‘SUMO\_HOME’], ‘tools’)  
sys.path.append(tools) # we need to import python modules from the $SUMO\_HOME/tools directory  
else:  
sys.exit(“please declare environment variable ‘SUMO\_HOME’”)  
import traci  
else:  
import libsumo as traci

import numpy as np  
import pandas as pd  
import ray  
import traci  
from ray import tune  
#from ray.rllib.algorithms.ppo import PPOConfig  
from ray.rllib.algorithms.ppo import (  
PPOConfig,  
PPOTF1Policy,  
PPOTF2Policy,  
PPOTorchPolicy,  
)  
from ray.rllib.env.wrappers.pettingzoo\_env import ParallelPettingZooEnv, PettingZooEnv  
from ray.tune.registry import register\_env  
from ray.tune.logger import pretty\_print

from stable\_baselines3.common.evaluation import evaluate\_policy  
from stable\_baselines3 import PPO

import ma\_environment.custom\_envs as custom\_env  
import supersuit as ss

def env\_creator(args):  
env = custom\_env.MA\_grid\_new(  
net\_file = “…”,  
route\_file =“…”,  
use\_gui=False,  
num\_seconds=30000,  
begin\_time=19800,  
time\_to\_teleport=300,  
reward\_fn=‘combined\_emission’,  
sumo\_warnings=False)  
return env

if **name** == “ **main** ”:  
ray.init()

```
env_name = "MA_grid_new"

register_env(env_name, lambda config: ParallelPettingZooEnv(env_creator(config)))
env = ParallelPettingZooEnv(env_creator({}))
#get obs and action space
obs_space = env.observation_space
act_space = env.action_space

config = (
    PPOConfig()
    .environment(env=env_name, disable_env_checking=True)
    .rollouts(num_rollout_workers=3, rollout_fragment_length='auto')
    .training(
        train_batch_size=512,
        lr=2e-5,
        gamma=0.95,
        lambda_=0.9,
        use_gae=True,
        clip_param=0.4,
        grad_clip=None,
        entropy_coeff=0.1,
        vf_loss_coeff=0.25,
        sgd_minibatch_size=64,
        num_sgd_iter=10,
    )
    .debugging(log_level="ERROR")
    .framework(framework="torch")
    .resources(num_gpus=int(os.environ.get("RLLIB_NUM_GPUS", "0")))
    .evaluation(evaluation_num_workers=1)
)

algo = config.build()  

for _ in range(10):
    
    print('Training iteration: ', _)
    print(algo.train())  

algo.evaluate()

```

---

<div class="post-metadata">

**Author:** ![kuza55](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kuza55/32/4931_2.png) [@kuza55](https://discuss.ray.io/u/kuza55)\
**Post date:** [August 17, 2023, 3:18pm UTC](https://discuss.ray.io/t/num-env-agent-steps-trained-0-even-though-steps-sampled/11730/2 "2023-08-17T15:18:10Z")

</div>

Did you make any progress on this?

I am noticing the same thing with my custom environment and even the custom\_env.py example without changes.

---

<div class="post-metadata">

**Author:** ![JenHa](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jenha/32/4898_2.png) [@JenHa](https://discuss.ray.io/u/JenHa)\
**Post date:** [August 17, 2023, 6:18pm UTC](https://discuss.ray.io/t/num-env-agent-steps-trained-0-even-though-steps-sampled/11730/3 "2023-08-17T18:18:58Z")

</div>

Unfortunately not - I decided to switch back to using stable baselines for the moment since I am on a deadline and was also facing issues with ray and a linux server.

---

<div class="post-metadata">

**Author:** ![wcthibault](https://avatars.discourse-cdn.com/v4/letter/w/cdc98d/32.png) [@wcthibault](https://discuss.ray.io/u/wcthibault)\
**Post date:** [August 17, 2023, 6:27pm UTC](https://discuss.ray.io/t/num-env-agent-steps-trained-0-even-though-steps-sampled/11730/4 "2023-08-17T18:27:43Z")

</div>

I have been experiencing a similar issue with off policy algorithms like DDPG and SAC when using replay buffers with storage units set to episodes. I made a post about it here: [Replay buffer with episodes as storage unit not training](https://discuss.ray.io/t/replay-buffer-with-episodes-as-storage-unit-not-training/11822)

---

<div class="post-metadata">

**Author:** ![abyaadrafid](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/abyaadrafid/32/5086_2.png) [@abyaadrafid](https://discuss.ray.io/u/abyaadrafid)\
**Post date:** [September 25, 2023, 1:46am UTC](https://discuss.ray.io/t/num-env-agent-steps-trained-0-even-though-steps-sampled/11730/5 "2023-09-25T01:46:41Z")

</div>

Seeing similar issues with my runs on a custom env too. However the agents seem to learn “something” at least.  
Edit : No they don’t.

---

<div class="post-metadata">

**Author:** ![abyaadrafid](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/abyaadrafid/32/5086_2.png) [@abyaadrafid](https://discuss.ray.io/u/abyaadrafid)\
**Post date:** [September 28, 2023, 11:35pm UTC](https://discuss.ray.io/t/num-env-agent-steps-trained-0-even-though-steps-sampled/11730/6 "2023-09-28T23:35:52Z")

</div>

So it turns out my custom env had a bug. My action and observation spaces were not defined in the correct range.

---

<div class="post-metadata">

**Author:** ![JerryHao-art](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jerryhao-art/32/6036_2.png) [@JerryHao-art](https://discuss.ray.io/u/JerryHao-art)\
**Post date:** [April 23, 2024, 10:29am UTC](https://discuss.ray.io/t/num-env-agent-steps-trained-0-even-though-steps-sampled/11730/7 "2024-04-23T10:29:50Z")

</div>

i miss the same problem, do you have any process on it

---

<div class="post-metadata">

**Author:** ![VisionZUS29](https://avatars.discourse-cdn.com/v4/letter/v/ea666f/32.png) [@VisionZUS29](https://discuss.ray.io/u/VisionZUS29)\
**Post date:** [April 25, 2024, 1:01am UTC](https://discuss.ray.io/t/num-env-agent-steps-trained-0-even-though-steps-sampled/11730/8 "2024-04-25T01:01:03Z")

</div>

@JerryHao-art @wcthibault @JenHa @kuza55  
Try specifically defining the polices, the mapping function, and the policies to be trained in the algo config… See below.

```auto
config = (
    PPOConfig()
    .environment(env=env_name, disable_env_checking=True)
    .rollouts(num_rollout_workers=3, rollout_fragment_length='auto')
    .training(
        train_batch_size=512,
        lr=2e-5,
        gamma=0.95,
        lambda_=0.9,
        use_gae=True,
        clip_param=0.4,
        grad_clip=None,
        entropy_coeff=0.1,
        vf_loss_coeff=0.25,
        sgd_minibatch_size=64,
        num_sgd_iter=10,
    )
    .debugging(log_level="ERROR")
    .framework(framework="torch")
    .resources(num_gpus=int(os.environ.get("RLLIB_NUM_GPUS", "0")))
    .evaluation(evaluation_num_workers=1)
    .multi_agent(
            policies=["policy0", "policy1"],
            policy_mapping_fn=lambda agent_id, episode, worker, **kwargs: "policy0" if agent_id == "agent_0" else "policy1",
            policies_to_train=["policy0", "policy1"]
        )
)

```

agent id must be defined in your env, for example in my env I have agent ids defined in `self._agent_ids = {f'agent_{i}' for i in range(self.num_agents)}`  
and you can expand your training script depending on num\_agents as well
