# Understanding agent\_timesteps\_total

**URL:** https://discuss.ray.io/t/understanding-agent-timesteps-total/9232
**Category:** RLlib
**Created:** [February 3, 2023, 11:52am UTC](https://discuss.ray.io/t/understanding-agent-timesteps-total/9232 "2023-02-03T11:52:10Z")
**Posts on this page:** 3
**Page:** 1

<div class="post-metadata">

### Author: ![Archana\_R](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/archana_r/32/3829_2.png) [@Archana\_R](https://discuss.ray.io/u/Archana_R)
#### Post date: [February 3, 2023, 11:52am UTC](https://discuss.ray.io/t/understanding-agent-timesteps-total/9232/1 "2023-02-03T11:52:10Z")

</div>

Hi ,

Below is a snapshot of my output

agent\_timesteps\_total: 4000  
counters:  
num\_agent\_steps\_sampled: 4000  
num\_agent\_steps\_trained: 4000  
num\_env\_steps\_sampled: 4000  
num\_env\_steps\_trained: 4000  
custom\_metrics: {}  
date: 2023-02-03\_12-46-22  
done: false  
episode\_len\_mean: .nan  
episode\_media: {}  
episode\_reward\_max: .nan  
episode\_reward\_mean: .nan  
episode\_reward\_min: .nan  
episodes\_this\_iter: 0  
episodes\_total: 0

My Code:

from ray.rllib.agents.ppo import PPOTrainer, DEFAULT\_CONFIG  
from ray.tune.logger import pretty\_print  
config = DEFAULT\_CONFIG.copy()

agent = PPOTrainer(config, env=“fss-v1”) #custom environment

for \_ in range(1):  
print(“Entered \_ :”,\_)  
result = agent.train()

My question:

1. Why does it show : episodes\_total = 0 ?
2. Why would the episode reward be NAN
3. What is agent\_timesteps\_total = 4000 mean ?

I checked Horizon config ( it is None - I do not understand this either and should i change its value )  
Urgently need your inputs please.

Thank you!

---

<div class="post-metadata">

### Author: ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)
#### Post date: [February 3, 2023, 2:24pm UTC](https://discuss.ray.io/t/understanding-agent-timesteps-total/9232/2 "2023-02-03T14:24:37Z")

</div>

Thus means that in one call to train, which samples 4000 steps from your environment(s), your environment did not terminate. Return done=True. The episode count and the mean reward do not update until episodes terminate. Training, by which I mean updating the policy, will occur every time it collects 4000 new environment steps.

---

<div class="post-metadata">

### Author: ![Archana\_R](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/archana_r/32/3829_2.png) [@Archana\_R](https://discuss.ray.io/u/Archana_R)
#### Post date: [February 3, 2023, 2:56pm UTC](https://discuss.ray.io/t/understanding-agent-timesteps-total/9232/3 "2023-02-03T14:56:05Z")

</div>

So it looks like my episode is not terminating . How do i get it so ? If my actions are always non legal actions , the game does not end.
