# PB2 seems stuck in space margins and raises exceptions with lambdas

**URL:** <https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467>\
**Category:** Ray Tune\
**Created:** [January 16, 2021, 4:10pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467 "2021-01-16T16:10:33Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [January 16, 2021, 4:10pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/1 "2021-01-16T16:10:33Z")

</div>

As of title, both @Peter_Pirog and I have encountered this bug where the PB2 remains fixed on the minimum and maximum values of the given hyper-parameters space. When providing a lambda, it crashes saying that it must be a tuple with length 2.

---

<div class="post-metadata">

**Author:** ![kai](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kai/32/3380_2.png) [@kai](https://discuss.ray.io/u/kai)\
**Post date:** [January 18, 2021, 1:56pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/2 "2021-01-18T13:56:52Z")

</div>

cc @amogkam who helped making the PB2 integration happen

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [January 18, 2021, 2:52pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/3 "2021-01-18T14:52:22Z")

</div>

Perfect! I will prepare an example on Colab to easily reproduce this peculiar behaviour.

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [January 18, 2021, 3:02pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/4 "2021-01-18T15:02:37Z")

</div>

Here is the Colab reproducing the aforementioned error: [https://colab.research.google.com/drive/1rnSJr2r-hCuDHTyeqOS9fM9J4kF4r7PA?usp=sharing](https://colab.research.google.com/drive/1rnSJr2r-hCuDHTyeqOS9fM9J4kF4r7PA?usp=sharing)

---

<div class="post-metadata">

**Author:** ![amogkam](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amogkam/32/17_2.png) [@amogkam](https://discuss.ray.io/u/amogkam)\
**Post date:** [January 18, 2021, 5:22pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/5 "2021-01-18T17:22:35Z")

</div>

Hi @LucaCappelletti94, thanks for sharing the Colab notebook.

Since no `config` parameter was passed to `tune.run`, the `hyperparam_bounds` is used for the initial search space, so all the `x` values for all the trials will either be 1 or 100. And the reason why these values never change is because no perturbations are actually occurring.  
 ![image](https://us1.discourse-cdn.com/flex020/uploads/ray/original/1X/79c927e8aeb5895e896c197a5b4a0e6a099a3d04.png)

This is because only each trial only runs 1 iteration. Instead, you need to add a loop to your loss function so that multiple training iterations are run:

```auto
def loss(config):
    for _ in range(100):
        tune.report(loss=config.get("x")**2)

```

And then, you also have to reduce the value of your perturbation\_interval. As it’s set up now, no perturbations occur until the trial runs for 1000 seconds. Instead you should lower this number, or change the setup to use iterations as the interval instead:

```auto
time_attr='training_iteration',
perturbation_interval=2,

```

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [January 18, 2021, 5:38pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/6 "2021-01-18T17:38:16Z")

</div>

Ok I see, is there an example on how to use this perturbation on something like a Sklearn or Keras model? I am trying to picture what models that I am usually familiar with may benefit from this particular optimization approach.

Thanks!

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [January 18, 2021, 6:03pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/7 "2021-01-18T18:03:42Z")

</div>

I have tried the following but still, it keeps jumping back and forth from -100 to 100. What should I change?

 ![Schermata 2021-01-18 alle 19.02.26](https://us1.discourse-cdn.com/flex020/uploads/ray/original/1X/35d48c048cee9a5db276ea476d0b35764c87008b.png)

---

<div class="post-metadata">

**Author:** ![amogkam](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amogkam/32/17_2.png) [@amogkam](https://discuss.ray.io/u/amogkam)\
**Post date:** [January 18, 2021, 11:16pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/8 "2021-01-18T23:16:57Z")

</div>

@LucaCappelletti94 ah yes for PB2 you also need to specify an initial hyperparameter search space by passing in a `config` to `tune.run`, per haps something like:

```auto
tune.run(
  loss,
  ...,
  config={
    'x': tune.uniform(-100, 100)
})

```

If you don’t specify this, then the initial hyperparameter values for all trials will be either -100 or 100, and since PB2 uses the information from previous trials to inform future hyperparameters, all future trials will be at these bound values as well. Try this out, and let me know if it works for you!

Yes, you can use these algorithms with sklearn and keras. Tune will work with anything that can be specified in a training function. You can see here for an example of Keras: [tune\_mnist\_keras — Ray v2.0.0.dev0](https://docs.ray.io/en/master/tune/examples/tune_mnist_keras.html)

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [January 19, 2021, 10:51am UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/9 "2021-01-19T10:51:03Z")

</div>

Ok I see, I will try this ASAP.

For the notes on the Keras model, I fail to see how the perturbation at the various training epochs may be applied. How would the PB2 method work there?

---

<div class="post-metadata">

**Author:** ![amogkam](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amogkam/32/17_2.png) [@amogkam](https://discuss.ray.io/u/amogkam)\
**Post date:** [January 19, 2021, 6:39pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/10 "2021-01-19T18:39:45Z")

</div>

In the example, we add a `TuneReportCallback` to `model.fit` inside our training function. This will automatically report results to Tune after every training epoch. With this, we can use Keras with Tune and you can pass in any scheduler when you call `tune.run`, whether that’s PB2, PBT, ASHA, etc. These schedulers will work just like with any other training function.

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [January 20, 2021, 1:21pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/11 "2021-01-20T13:21:43Z")

</div>

I understand how the TuneReporter works, what I am not understanding is how the PB2 would update the parameters of the model as the epochs proceed. Would this be like ASHA? Is PB2 also a median-based early stopping mechanism?

---

<div class="post-metadata">

**Author:** ![amogkam](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amogkam/32/17_2.png) [@amogkam](https://discuss.ray.io/u/amogkam)\
**Post date:** [January 21, 2021, 5:50pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/12 "2021-01-21T17:50:23Z")

</div>

PB2 is very similar to Population Based Training, it’s not an early stopping algorithm. Multiple trials are run in parallel, and at a certain interval, the bottom percentile of trials copy the hyperparameters and state (i.e. model weights) of the top percentile, slightly perturbs the hyperparameters and then continues training until the stopping condition is met.

The difference is that PBT uses a heuristic based approach to perturb the hyperparameters, while PB2 uses a Bayesian-style Gaussian model and leverages previous results to perturb the hyperparameters once they are copied from the top trials.

Population Based Training generally is a good HPO approach and has shown to perform well across a variety of domains (Image processing, NLP/transformer models, RL), and is also more resource efficient than other approaches.

PB2 is particularly beneficial for cases where you have to use a smaller population size, perhaps due to resource constraints.

Does this help answer your question?

For more info you can also check out:  
[The PB2 documentation](https://docs.ray.io/en/master/tune/api_docs/schedulers.html#tune-scheduler-pb2)  
[PB2 Blog Post](https://www.anyscale.com/blog/population-based-bandits)

And for more in-depth info for context on standard Population Based Training and the motivation on why we need PB2:  
[This Anyscale Connect Talk](https://www.anyscale.com/blog/population-based-bandits)

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [January 21, 2021, 6:04pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/13 "2021-01-21T18:04:47Z")

</div>

I don’t understand how the tuning of something like the number of layers of a model can be tuned if the weights are kept as a different number of layers would imply a different shape for the weights. Also, how are the weights kept? Is the checkpointing system automatic?

---

<div class="post-metadata">

**Author:** ![kai](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kai/32/3380_2.png) [@kai](https://discuss.ray.io/u/kai)\
**Post date:** [January 22, 2021, 3:24pm UTC](https://discuss.ray.io/t/pb2-seems-stuck-in-space-margins-and-raises-exceptions-with-lambdas/467/14 "2021-01-22T15:24:02Z")

</div>

Your trainable function will have to take care of checkpoint saving and restoring, but that’s no different to other trainables.

As for number of layers, this is no problem as long as the number of layers remains constant within a trial (i.e. they can’t be mutated). So let’s say you have 4 trials A, B, C and D. A to C use 2 layers each and D uses 3 layers. Trial A performs badly and should be stopped. Trial D currently performs best. Then Trial A will copy all hyperparameters from Trial D (including the number of layers, 3) and restores from the latest checkpoint of D (which has weights for all three layers). It also perturbs some of the hyperparameters, e.g. the learning rate or so, and then continues training.

So in a way it does stop bad performing trials early, but the resources are then used to continue training on a modified copy of a well performing trial. Does this help?
