# Saving best model at the end of the training

**URL:** <https://discuss.ray.io/t/saving-best-model-at-the-end-of-the-training/742>\
**Category:** Ray Tune\
**Created:** [February 5, 2021, 11:25am UTC](https://discuss.ray.io/t/saving-best-model-at-the-end-of-the-training/742 "2021-02-05T11:25:51Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![J\_J](https://avatars.discourse-cdn.com/v4/letter/j/3ab097/32.png) [@J\_J](https://discuss.ray.io/u/J_J)\
**Post date:** [February 5, 2021, 11:25am UTC](https://discuss.ray.io/t/saving-best-model-at-the-end-of-the-training/742/1 "2021-02-05T11:25:51Z")

</div>

Hi guys, I am trying to run hyperparamsearch using tune.with\_parameters, because my data is too big. My training function takes parameters config, data and checkpoint\_dir - I am currently saving the model after every trial. Is there a way, using tune.run and tune.with\_parameters to save the best model on the disk after all the trials are run? Can I achieve it with checkpointing? My first idea was to use Trainable class like in this example: [[tune] How to checkpoint best model · Issue #10290 · ray-project/ray · GitHub](https://github.com/ray-project/ray/issues/10290), but as far as I understand I can’t mix up Trainable class and tune.with\_parameters. Moreover when I tried, it didn’t work.

my current solution code snippet, where config - pipeline hyperparameters (for transformers and models):

```
def train_model(config, data, checkpoint_dir=None):
    (train_data, y_train, dev_data, y_valid) = data
    model = ModelPipeline.from_config(config)

    model.fit(train_data, y_train, validation_data=(dev_data, y_valid))
    dev_metric_results = Evaluation(metrics=['custom_metrics']) \
        .evaluate(model=model, X=dev_data, y_true=y_valid)

    with tune.checkpoint_dir(step=0) as checkpoint_dir:
        model.save(checkpoint_dir)

    tune.report(custom_metrics=dev_metric_results['custom_metrics'])
        
analysis = tune.run(tune.with_parameters(train_model, data=data),
            name = name,
            config = config,
            num_samples=num_samples,
            time_budget_s = time_budget,
            verbose = verbose,
            resources_per_trial = resources,
            metric = 'custom_metrics',
            mode = 'max',
            keep_checkpoints_num=1,
            checkpoint_freq=1,
            checkpoint_score_attr='custom_metrics')
```

---

<div class="post-metadata">

**Author:** ![kai](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kai/32/3380_2.png) [@kai](https://discuss.ray.io/u/kai)\
**Post date:** [February 5, 2021, 12:00pm UTC](https://discuss.ray.io/t/saving-best-model-at-the-end-of-the-training/742/2 "2021-02-05T12:00:26Z")

</div>

Hi, you can access the checkpoint of the best performing trial like this:

```auto
best_checkpoint_dir = analysis.best_checkpoint

```

For more information you can take a look here: [Analysis (tune.analysis) — Ray v1.1.0](https://docs.ray.io/en/latest/tune/api_docs/analysis.html#experimentanalysis-tune-experimentanalysis)

---

<div class="post-metadata">

**Author:** ![J\_J](https://avatars.discourse-cdn.com/v4/letter/j/3ab097/32.png) [@J\_J](https://discuss.ray.io/u/J_J)\
**Post date:** [February 5, 2021, 12:32pm UTC](https://discuss.ray.io/t/saving-best-model-at-the-end-of-the-training/742/3 "2021-02-05T12:32:59Z")

</div>

Hi, thanks for the answer.My question isn’t about finding the path of the best checkpoint (which would still require checkpointing model after each trial), but to have only one model (the best one) saved at the disk, when all the trials are done.

---

<div class="post-metadata">

**Author:** ![xwjiang2010](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/xwjiang2010/32/1476_2.png) [@xwjiang2010](https://discuss.ray.io/u/xwjiang2010)\
**Post date:** [January 18, 2022, 8:33pm UTC](https://discuss.ray.io/t/saving-best-model-at-the-end-of-the-training/742/4 "2022-01-18T20:33:03Z")

</div>

I am afraid this is not currently supported. There is at least one checkpoint per trial.

Would it work if you have a separate customized process to monitor trial results and proactively delete checkpoints of trials less performant?

In the long run, it may be helpful for tune to provide something API to interact with ongoing experiment (deleting less performant trial checkpoints can be instrumented using such APIs.)

---

<div class="post-metadata">

**Author:** ![poro1301](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/poro1301/32/6267_2.png) [@poro1301](https://discuss.ray.io/u/poro1301)\
**Post date:** [June 28, 2024, 1:59pm UTC](https://discuss.ray.io/t/saving-best-model-at-the-end-of-the-training/742/5 "2024-06-28T13:59:20Z")

</div>

May be the it is not supported but in my case, i save the model according to f1 using Checkpoint Config:

```auto
tuner = tune.Tuner(
    trainable_with_resources,
    param_space=search_space,
    tune_config=tune.TuneConfig(
        num_samples=1,
        mode='max',
        metric='eval_seq_f1'
        # scheduler=scheduler,
    ),
    run_config=RunConfig(
        name="tune_transformer_pbt",
        storage_path='/data-gpu/trungct/tmp',
        log_to_file=True,
        progress_reporter=reporter,
        checkpoint_config=CheckpointConfig(
            num_to_keep=2,
            checkpoint_score_attribute="eval_seq_f1",
            checkpoint_score_order='max',
        ),
    ),
)

```

And then the best checkpoint paths (defined in `num_to_keep`), … is available at:

```auto
tuner = Tuner(...)
tuner.fit()
results.get_best_result(scope='all').best_checkpoints

```
