# Tuning process with PBT is killed after a very small number of iterations (6/500))

**URL:** <https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227>\
**Category:** Ray Tune\
**Created:** [December 14, 2020, 9:27am UTC](https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227 "2020-12-14T09:27:08Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [December 14, 2020, 9:27am UTC](https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227/1 "2020-12-14T09:27:08Z")

</div>

Hello,

I understand that this question will be vague, and it is mainly though to my ignorance on the PBT mechanism.

I am trying to tune a CNN model on genomic sequence data on a relatively easy task: the first model I came up with in 5 minutes achieved a 0.92 AUPRC and 0.94 AUROC validation scores

 ![photo_2020-12-14 10.03.42](https://us1.discourse-cdn.com/flex020/uploads/ray/original/1X/6a20c1dfd954f1abfe4c92afa2d9b51157761273.jpeg)

After having established that the task is relatively easy (hence most models will achieve a decent result), I wanted to try the PBT method to tune the meta-model to learn how to use PBT.

Having defined the parameters as the number of filters/kernel size plus some activation regularization weight, I have tried to use the PBT as follows:

First I have defined a space of parameters using uniform lambdas [as shown in the tutorial here](https://docs.ray.io/en/master/tune/tutorials/tune-advanced-tutorial.html). Here I use double lambdas just to capture the local variables defining the range.

```python
space = {
    key: (lambda vr: lambda: np.random.uniform(*vr))(val_range)
    for key, val_range in model.space().items()
}

```

Then I define the train method as follows:

```python
def train_convnet(config):
    import silence_tensorflow.auto
    window_size = 256
    train, test = create_training_sequence(window_size)
    meta_model: Model = build_model(window_size)
    meta_model.space()
    model = meta_model.build(**config)
    model.compile(
        optimizer='nadam',
        loss="binary_crossentropy",
        metrics=[
            "accuracy",
            AUC(curve="PR", name="auprc"),
            AUC(curve="ROC", name="auroc")
        ]
    )
    model.fit(
        train,
        validation_data=test,
        epochs=1000,
        verbose=False,
        callbacks=[
            TuneReportCallback(metrics="val_auprc"),
            EarlyStopping(monitor="auprc", patience=5)
        ]
    )

```

And finally I call the tuning process:

```auto
from ray.tune.stopper import EarlyStopping as TuneEarlyStopping

scheduler = PopulationBasedTraining(
    time_attr="training_iteration",
    perturbation_interval=5,
    hyperparam_mutations=space
)

analysis = tune.run(
    train_convnet,
    name="pbt_test",
    scheduler=scheduler,
    metric="val_auprc",
    mode="max",
    verbose=1,
    stop=TuneEarlyStopping("val_auprc", ),
    resources_per_trial={
        "cpu": cpu_count()//4,
        "gpu": 1
    },
    num_samples=500
)

```

The tuning processes are then all terminated and they achieve at most an AUPRC of 0.45.

What am I doing wrong? What information is needed to properly resolve this issue?

Though to the reserved nature of the training labels, I cannot share an example of the dataset but I believe that the issue at hand has little to do with the considered task and more to do with how I am using tune and PBT.

Thank you,  
Luca

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [December 15, 2020, 8:22am UTC](https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227/2 "2020-12-15T08:22:34Z")

</div>

Can you try removing the EarlyStopping parameters?

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [December 15, 2020, 8:24am UTC](https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227/3 "2020-12-15T08:24:49Z")

</div>

Which of the Early Stopping ones? The ones within the Keras’s Early Stopping?

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [December 15, 2020, 8:27am UTC](https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227/4 "2020-12-15T08:27:23Z")

</div>

Can you try removing both at first?

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [December 15, 2020, 8:29am UTC](https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227/5 "2020-12-15T08:29:11Z")

</div>

I am now trying to run the experiment with neither of them. I am wondering if I am not passing the mode of optimization somewhere, and maybe I am seeing always the minima of the AUPRC being reported instead of the max.

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [December 15, 2020, 8:37am UTC](https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227/6 "2020-12-15T08:37:21Z")

</div>

So far I am seeing all the models converging to the very same value of val\_auprc, with extraordinary precision. I don’t think overfitting of the model is an issue, seeing how easy the task at hand is.

If you’d like to have a call to see first hand the complete notebook please do let me know on Slack.

Thanks!

 ![Schermata 2020-12-15 alle 09.34.51](https://us1.discourse-cdn.com/flex020/uploads/ray/original/1X/8ef3532e50a5255f763fb9309c69c048c9ed1b52.png)

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [December 15, 2020, 11:05pm UTC](https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227/7 "2020-12-15T23:05:28Z")

</div>

> [@LucaCappelletti94](#):
>
> So far I am seeing all the models converging to the very same value of val\_auprc, with extraordinary precision. I don’t think overfitting of the model is an issue, seeing how easy the task at hand is.

Hmm, can you set `verbose=3` for tune.run to see which metrics are actually being updated?

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [January 13, 2021, 8:25am UTC](https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227/8 "2021-01-13T08:25:03Z")

</div>

Sorry for the delay, got sidetracked with another project.  
Apparently, the space of hyper-parameters was too vast and the BO, even with 100 initial random steps.  
If I significantly restrict the hyper-parameters space the performance increase, but they don’t achieve the performance of the quickly hand-picked model even with over 600 iterations.  
I don’t understand why is this happening.

---

<div class="post-metadata">

**Author:** ![LucaCappelletti94](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lucacappelletti94/32/144_2.png) [@LucaCappelletti94](https://discuss.ray.io/u/LucaCappelletti94)\
**Post date:** [January 13, 2021, 8:26am UTC](https://discuss.ray.io/t/tuning-process-with-pbt-is-killed-after-a-very-small-number-of-iterations-6-500/227/9 "2021-01-13T08:26:14Z")

</div>

The early stopping class from tune is what still kills the process after the patience number of iterations (even with the restricted hyper-parameters space). I guess there is something wrong in there, as I see that the performance are getting better.
