# Cannot get a simple Evaluation to work as intended

**URL:** <https://discuss.ray.io/t/cannot-get-a-simple-evaluation-to-work-as-intended/7308>\
**Category:** RLlib\
**Created:** [August 24, 2022, 12:15am UTC](https://discuss.ray.io/t/cannot-get-a-simple-evaluation-to-work-as-intended/7308 "2022-08-24T00:15:24Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![hridayns](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hridayns/32/1222_2.png) [@hridayns](https://discuss.ray.io/u/hridayns)\
**Post date:** [August 24, 2022, 12:15am UTC](https://discuss.ray.io/t/cannot-get-a-simple-evaluation-to-work-as-intended/7308/1 "2022-08-24T00:15:24Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- High: It blocks me to complete my task.

I have a saved checkpoint from training, and I would like to run just evaluation episodes now with exploration set to False (which I have configured correctly). The problem is that I am unable to control the number of episodes for which this happens.  
.  
.  
.  
policy\_conf[‘evaluation\_interval’] = 1 # every 1 episode?  
policy\_conf[‘evaluation\_duration’] = 1 # for 1 episode  
policy\_conf[‘evaluation\_duration\_unit’] = ‘episodes’  
.  
.  
.

stop = {  
“episodes\_total”: 1,  
}

result = tune.run(  
‘PPO’,  
name=‘ma-pheromone’,  
config=policy\_conf,  
metric=‘episode\_reward\_mean’,  
mode=‘max’,  
trial\_name\_creator=trial\_str\_creator,  
stop=stop,  
verbose=1,  
log\_to\_file=(‘test.log’,‘error.log’),  
local\_dir=‘./tune-eval-results’,  
restore=‘./tune-results/ma-pheromone/PPO-\_100-2-4-1000.0-0.1-0.1-0-1\_\_0\_2022-08-23\_20-23-55/checkpoint\_000030/checkpoint-30’,  
num\_samples=10,  
)

I have set the evaluation duration to 1 and even the **stop condition** for episodes\_total to 1. I don’t understand why I am unable to do a simple evaluation of 10 episodes, so I tried to just stick to run 1 episode of evaluation but it just does not work. I set num\_samples to 10 because I want to repeat the evaluation 10 times for consistency. Please help me.

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [September 4, 2022, 3:28pm UTC](https://discuss.ray.io/t/cannot-get-a-simple-evaluation-to-work-as-intended/7308/2 "2022-09-04T15:28:37Z")

</div>

It looks like you are using tune to evaluate. The stop condition you set does not apply to **evaluated episodes** but to **trained episodes.**. Try using the algorithm’s `evaluate` function instead, since you don’t seem to be interesting in training and therefore not interested in tuning anything here.

---

<div class="post-metadata">

**Author:** ![hridayns](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hridayns/32/1222_2.png) [@hridayns](https://discuss.ray.io/u/hridayns)\
**Post date:** [September 4, 2022, 4:10pm UTC](https://discuss.ray.io/t/cannot-get-a-simple-evaluation-to-work-as-intended/7308/3 "2022-09-04T16:10:41Z")

</div>

I have tried in the past and I have had trouble using evaluate. Also, not very sure where the evaluation results get stored on the system for visualization. Any workaround to keep using tune but for evaluation?

---

<div class="post-metadata">

**Author:** ![RaymondK](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/raymondk/32/2652_2.png) [@RaymondK](https://discuss.ray.io/u/RaymondK)\
**Post date:** [September 5, 2022, 10:21am UTC](https://discuss.ray.io/t/cannot-get-a-simple-evaluation-to-work-as-intended/7308/4 "2022-09-05T10:21:12Z")

</div>

This is a bug. I had the same behavior that the number of episodes was more than the evaluation\_duration that was specified. This is fixed in Ray 2.0.0.  
See also: [[RLlib] Excessive evaluation if rollout\_fragment\_length \< timesteps\_per\_iteration · Issue #27821 · ray-project/ray · GitHub](https://github.com/ray-project/ray/issues/27821)

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [September 5, 2022, 3:34pm UTC](https://discuss.ray.io/t/cannot-get-a-simple-evaluation-to-work-as-intended/7308/5 "2022-09-05T15:34:07Z")

</div>

@hridayns The results would not be visualized. But you can record the statistics yourself and for most metrics an average will suffice - and does not need visualization. Maybe you can calculate percentiles of distributions of rewards or something like that.

@RaymondK Are you sure that this is the referenced bug or even a bug? The way I understand hridayns is what I would expect to happen.

---

<div class="post-metadata">

**Author:** ![RaymondK](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/raymondk/32/2652_2.png) [@RaymondK](https://discuss.ray.io/u/RaymondK)\
**Post date:** [September 5, 2022, 6:06pm UTC](https://discuss.ray.io/t/cannot-get-a-simple-evaluation-to-work-as-intended/7308/6 "2022-09-05T18:06:00Z")

</div>

To be honest, I am not completely clear what @hridayns is asking. Maybe I should have stated it less directly. At least I did experience the bug I referenced in the Github issue and it was solved by Ray 2.0.0 but I am not sure this solves the problem hridayns has.

---

<div class="post-metadata">

**Author:** ![hridayns](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hridayns/32/1222_2.png) [@hridayns](https://discuss.ray.io/u/hridayns)\
**Post date:** [September 5, 2022, 9:50pm UTC](https://discuss.ray.io/t/cannot-get-a-simple-evaluation-to-work-as-intended/7308/7 "2022-09-05T21:50:09Z")

</div>

I have actually not yet tried to evaluate using tune.run() in Ray 2.0.0. Since I switched to using evaluate() from 1.11 and have continued to do the same as it seemed to work fine until I ran into another issue later which I made a post about
