# \[Tune PBT\] Population Based Training :: Questions & Errors

**URL:** <https://discuss.ray.io/t/tune-pbt-population-based-training-questions-errors/1466>\
**Category:** Ray Tune\
**Created:** [March 30, 2021, 5:51am UTC](https://discuss.ray.io/t/tune-pbt-population-based-training-questions-errors/1466 "2021-03-30T05:51:12Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Kai\_Yun](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kai_yun/32/167_2.png) [@Kai\_Yun](https://discuss.ray.io/u/Kai_Yun)\
**Post date:** [March 30, 2021, 5:51am UTC](https://discuss.ray.io/t/tune-pbt-population-based-training-questions-errors/1466/1 "2021-03-30T05:51:12Z")

</div>

(The questions I’m posting here will be fairly simple to answer for any experienced user. I’m still in the beginning stages with Ray and am just playing around with many of its utilities. So, any help will be nice!)

Hi,

I’ve run a PBT experiment on PPO with my custom simulator.  
Among the 4 trials I ran, only 1 survived after six hours, as you can see from the figure below. I need help with understanding why the errors occurred. I have included the link to the error logs below. If someone can explain the reasons, it will be really helpful.

- [Trial 00001 Error LOG](https://gist.githubusercontent.com/kaiyun717/ab0fc94c355b68790a409fe24ea2b2ef/raw/7436ad1013f2ca978c950277f3deb1d9b58a0221/trial00001_error.json)
- [Trial 00002 Error LOG](https://gist.githubusercontent.com/kaiyun717/2253294a95e2d88582d3b9762eb1e691/raw/a65a11dbee2308d60bedd8a5bf5a3e7a13a7fe42/trial00002_error.json)
- [Trial 00003 Error LOG](https://gist.githubusercontent.com/kaiyun717/46b9fe998aae3c626e8690f6bc1f4365/raw/bcce9e2d5f3d8b3b28c9d99a66c291fc3468f272/trail00003_error.json)

 ![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/1X/1a987c3c2a9736fdfa0436799aa4869c79bd2668.png)

Furthermore, there are some things I do not understand about Ray’s implementation of PBT.

Below is part of my code that is concerned with PBT. As you can see, I’m trying to optimize these six hyper-parameters: `lambda`, `clip_param`, `lr`, `num_sgd_iter`, `sgd_minibatch_size`, `train_batch_size`.  
The questions I have are:

1. I used `tune.qrandint(128, 1024, 128)`, hoping to have the candidates in the search space to be rounded to the integer increments of 128 as the Tune API states. But, in the `pbt_global.txt` file, I found values such as 307 and 153. How is this possible?
2. Can someone help me understand how to interpret the `pbt_global.txt` file? I’m very lost with it. I Here is the [link to `pbt_global.txt`](https://gist.githubusercontent.com/kaiyun717/8f61fb3a70206fc6faafc48fdf0cb070/raw/8b168fedefbf407627e1401180adeee4d5c252e8/pbt_global.json). I mainly want to know what to look to figure out where each trials changed their hyper-parameters using exploration and exploitation. This will be fairly simple for any experience user, I assume.

```python
pbt = PopulationBasedTraining(
        time_attr="training_iteration",
        perturbation_interval=50,
        metric="episode_reward_max",
        mode="max",
        hyperparam_mutations={
            "lambda": lambda: random.uniform(0.95, 1.0),
            "clip_param": lambda: random.uniform(0.01, 0.5),
            "lr": [1e-3, 5e-4, 1e-4, 5e-5, 1e-5],
            "num_sgd_iter": lambda: random.randint(1, 30),
            "sgd_minibatch_size": tune.qrandint(128, 1024, 128),
            "train_batch_size": tune.qrandint(2_500, 7_500, 2_500)
        }
    )

results = tune.run(
        "PPO",
        name="PBT_PPO",
        config=config,
        checkpoint_freq=1,
        stop={
            "time_total_s": 43_200
        },
        checkpoint_score_attr="episode_reward_max",
        scheduler=pbt,
        num_samples=4,
        local_dir=args.save_dir
    )

```

Thank you!

---

<div class="post-metadata">

**Author:** ![kai](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kai/32/3380_2.png) [@kai](https://discuss.ray.io/u/kai)\
**Post date:** [March 30, 2021, 7:56am UTC](https://discuss.ray.io/t/tune-pbt-population-based-training-questions-errors/1466/2 "2021-03-30T07:56:45Z")

</div>

Hi @Kai_Yun,

for `tune.qrandint` - this sampler is used for the initial sampling of hyperparameter values. In population based training, hyperparameters are _mutated_ when a trial exploits another trial - and per the original paper this means that the parameter values are multiplied with 0.8 or 1.2. Hence the 153 - this is just 128\*1.2 = 153.6 (rounded down to 153). (And 256 \* 1.2 = 307.2 ~= 307).

There’s a couple of things you could do to change this behavior, and the most straightforward one would be to just pass `tune.choice` instead - e.g. `"sgd_minibatch_size": tune.choice(list(range(128, 1025, 128))),` - then the mutations will also only use values supplied from the list of categories.

Instead of looking at the hyperparameter globals file, I’d suggest looking at the files of the individual trials. Generally the row is like this: `old_tag, new_tag, old_step, new_step, old_conf, new_conf` - so each time a trial exploits another trial, you get its old tag, new tag, the step the original trial was at when it exploitet the new trial, the step the new trial had (and the trial will henceforth have as well), the old configuration, and the new (mutated) configuration.

For the tensorflow errors, this might be related to your search space definition and it’s hard to tell without looking at your custom simulator. Though cc @sven1977 if you’ve seen something like this before.

---

<div class="post-metadata">

**Author:** ![Kai\_Yun](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kai_yun/32/167_2.png) [@Kai\_Yun](https://discuss.ray.io/u/Kai_Yun)\
**Post date:** [March 31, 2021, 5:46am UTC](https://discuss.ray.io/t/tune-pbt-population-based-training-questions-errors/1466/3 "2021-03-31T05:46:34Z")

</div>

Thank you @kai !

I have some other follow-up questions. I made a summary of all my `pbt_policy_<trial-number>.txt` files based on your answer as you can see below.  
The questions I have are:

1. Why does the first checkpoint of `policy_00000` have Tag 2 as its “old\_tag”? Shouldn’t it supposed to be Tag 0 since this is its first checkpoint? If you can explain how this tagging works, it’d be helpful.
2. The first three checkpoints of all policies have the exact same values. Does this simply just mean that they are all taking on the hyperparameters of the same tag since trials in PBT take the best config among the trials?
3. The `pbt_global.txt` file has 9 rows. My PBT experiment had 15 checkpoint and 9 perturbations in total. As you can see in the figure below, the total number of checkpoints from all individual pbt files is 15. What should I make of this?

Thanks for the kind answers again.

 ![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/1X/0bd4f8bc72a71f5bd46809b119a2886922738ff1.png)

---

<div class="post-metadata">

**Author:** ![kai](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kai/32/3380_2.png) [@kai](https://discuss.ray.io/u/kai)\
**Post date:** [April 1, 2021, 3:49pm UTC](https://discuss.ray.io/t/tune-pbt-population-based-training-questions-errors/1466/4 "2021-04-01T15:49:01Z")

</div>

Hi,

it is a bit tricky to parse the policy file. You can also take a look at the code here to see how we parse this for replay: [ray/pbt.py at master · ray-project/ray · GitHub](https://github.com/ray-project/ray/blob/master/python/ray/tune/schedulers/pbt.py#L669)

1. When a trial exploits another trial, it also copies the exploitation history. As you can see in your PBT global log, trial 2 exploits trial 1 first (second row). Then trial 0 exploits trial 2. So the history here is 1-\>2-\>0. In your case, trial 0 still performs bad and later exploits trial 3. However, before then trial 3 never exploited anything else. Thus, if you were to replay that trial, you would just run the trial 3 configuration and to nothing else.

2. The first three perturbations are the same because trial 0 exploited trial 3, and all other trials exploited trial 0 in the last iteration, copying the full history of trial 0

3. pbt\_global.txt logs all perturbations. A trial might save several checkpoints before exploiting another trial, so the number of checkpoints is unrelated to the number of perturbations

I hope this helps!
