# Unable to restore fully trained checkpoint

**URL:** <https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259>\
**Category:** RLlib\
**Created:** [November 14, 2022, 8:00am UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259 "2022-11-14T08:00:49Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Xorgress\_Grox](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/xorgress_grox/32/2982_2.png) [@Xorgress\_Grox](https://discuss.ray.io/u/Xorgress_Grox)\
**Post date:** [November 14, 2022, 8:00am UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/1 "2022-11-14T08:00:49Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- High: It blocks me to complete my task.

I’ve finished training with a bunch of algorithms using the Tuner() API and air library and they all have their appropriate checkpoint folders and files. However I can’t seem to restore those checkpoints. I tried using Tuner.restore() and run(restore=), both didn’t work.  
When using Tuner.restore() I got this error:

> (ApexDQN pid=476180) 2022-11-14 14:39:07,333 INFO trainable.py:715 – Checkpoint path was not available, trying to recover from latest available checkpoint instead. Unavailable checkpoint path: G:\Repos\ML\_CIV6\models(3w2s)-d(2w1s)\_default\APEX\APEX\_my\_env\_5ed3c\_00000\_0\_2022-09-27\_20-07-42\checkpoint\_004000\checkpoint-4000

And for run(restore=) I got this error:

> RuntimeError: Could not find Tuner state in restore directory. Did you passthe correct path (including experiment directory?) Got: G:\Repos\ML\_CIV6\models(3w2s)-d(2w1s)\_default\APEX\APEX\_my\_env\_5ed3c\_00000\_0\_2022-09-27\_20-07-42

The training code:

> <https://github.com/TheGroxEmpire/ML_CIV6/blob/pettingzoo/train_apex-dqn.py>

I’ve also tried referring to the folders above the checkpoint file, it all resulted in the same error output.

Thank you in advance.

---

<div class="post-metadata">

**Author:** ![varunjammula](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/varunjammula/32/2065_2.png) [@varunjammula](https://discuss.ray.io/u/varunjammula)\
**Post date:** [November 16, 2022, 10:21am UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/2 "2022-11-16T10:21:57Z")

</div>

hi, i think you need to restore from: `G:\Repos\ML_CIV6\models(3w2s)-d(2w1s)_default\APEX\APEX_my_env_5ed3c_00000_0_2022-09-27_20-07-42\` if you are using Tuner()

---

<div class="post-metadata">

**Author:** ![Xorgress\_Grox](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/xorgress_grox/32/2982_2.png) [@Xorgress\_Grox](https://discuss.ray.io/u/Xorgress_Grox)\
**Post date:** [November 16, 2022, 3:38pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/3 "2022-11-16T15:38:26Z")

</div>

I have tried that, and bunch of other folder path and none of them worked

---

<div class="post-metadata">

**Author:** ![varunjammula](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/varunjammula/32/2065_2.png) [@varunjammula](https://discuss.ray.io/u/varunjammula)\
**Post date:** [November 20, 2022, 6:30pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/4 "2022-11-20T18:30:53Z")

</div>

Can you upgrade to 2.1 and check?

---

<div class="post-metadata">

**Author:** ![Xorgress\_Grox](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/xorgress_grox/32/2982_2.png) [@Xorgress\_Grox](https://discuss.ray.io/u/Xorgress_Grox)\
**Post date:** [November 22, 2022, 3:29am UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/5 "2022-11-22T03:29:57Z")

</div>

I’ve tried 2.1, same issue persists.

---

<div class="post-metadata">

**Author:** ![Blubberblub](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/blubberblub/32/1396_2.png) [@Blubberblub](https://discuss.ray.io/u/Blubberblub)\
**Post date:** [December 21, 2022, 7:55am UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/6 "2022-12-21T07:55:04Z")

</div>

I found loading from checkpoint a little tricky too. This works for me:

```auto
algo_cls = get_algorithm_class("DQN")
algo = algo_cls(config=config)
checkpoint_path = "/path_to_folder/checkpoint_000225/rllib_checkpoint.json"
algo.restore(checkpoint_path)

```

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [December 21, 2022, 9:13am UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/7 "2022-12-21T09:13:39Z")

</div>

Hi, please consider upgrading to 2.2 and use the `Algorithm.from_checkpoint` API.  
Other than that: Can you post a reproduction script, please?

---

<div class="post-metadata">

**Author:** ![james116blue](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/james116blue/32/3551_2.png) [@james116blue](https://discuss.ray.io/u/james116blue)\
**Post date:** [December 29, 2022, 10:38am UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/8 "2022-12-29T10:38:34Z")

</div>

Same error for me after successful training

```auto
    config = (
        PPOConfig()
        .rollouts(
            num_rollout_workers=4,
            # rollout_fragment_length=512
        )
        .training(
            train_batch_size=512,
            lr=2e-5,
            gamma=0.99,
            lambda_=0.9,
            use_gae=True,
            clip_param=0.4,
            grad_clip=None,
            entropy_coeff=0.1,
            vf_loss_coeff=0.25,
            sgd_minibatch_size=64,
            num_sgd_iter=10,
        )
        .environment(env=env_name, clip_actions=True)
        .debugging(log_level="ERROR")
        .framework(framework="torch")
        .resources(num_gpus=int(os.environ.get("RLLIB_NUM_GPUS", "0")))
    )

    tune.Tuner(
        "PPO",
        run_config=air.RunConfig(stop={"timesteps_total": 5000000}),
        param_space=config.to_dict(),
    ).fit()

```

Version 2.2  
Also, there is a list of files in restore directory

```auto
error.txt  
events.out.tfevents.1671970856.pcname  
params.json  
params.pkl  
result.json

```

What is the correct way to load results of Tuner.fit() from directory in order to be analyzed?

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [December 29, 2022, 2:20pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/9 "2022-12-29T14:20:57Z")

</div>

@james116blue ,

Please post a complete reproduction script!

`Algorithm.load_checkpoint()` is deprecated.  
The “correct” way to load checkpoints is with the `Algorithm.from_checkpoint()` API.  
You can find multiple examples in the examples folder and [our documentation](https://docs.ray.io/en/latest/rllib/rllib-saving-and-loading-algos-and-policies.html).

You can, for example retrieve the best checkpoint (and later load it) with:

```auto
best_checkpoint = results.get_best_result(
            metric="episode_reward_mean",
            mode="max"
        ).checkpoint

```

Cheers

---

<div class="post-metadata">

**Author:** ![james116blue](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/james116blue/32/3551_2.png) [@james116blue](https://discuss.ray.io/u/james116blue)\
**Post date:** [December 31, 2022, 1:47pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/10 "2022-12-31T13:47:27Z")

</div>

Thank you for your detailed answer.  
I have found solution thanks to your reply and updated documentation

---

<div class="post-metadata">

**Author:** ![AI360](https://avatars.discourse-cdn.com/v4/letter/a/e5b9ba/32.png) [@AI360](https://discuss.ray.io/u/AI360)\
**Post date:** [February 6, 2023, 2:37pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/12 "2023-02-06T14:37:41Z")

</div>

> [@arturn](#):
>
> Algorithm.from\_checkpoint()

Algorithm.from\_checkpoint() can not load the checkpoints of Tuning in Ray 2.2. Could you please show how we can use it ? if it is possible in Colab.

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [February 6, 2023, 6:06pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/13 "2023-02-06T18:06:28Z")

</div>

Hi @AI360 ,

This is an example from our docs: [ray/saving\_and\_loading\_algos\_and\_policies.py at master · ray-project/ray · GitHub](https://github.com/ray-project/ray/blob/master/rllib/examples/documentation/saving_and_loading_algos_and_policies.py)

---

<div class="post-metadata">

**Author:** ![AI360](https://avatars.discourse-cdn.com/v4/letter/a/e5b9ba/32.png) [@AI360](https://discuss.ray.io/u/AI360)\
**Post date:** [February 6, 2023, 6:14pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/14 "2023-02-06T18:14:33Z")

</div>

Yes, I checked it, but I got this error:

File “/usr/local/lib/python3.10/site-packages/ray/rllib/algorithms/algorithm.py”, line 278, in from\_checkpoint  
return Algorithm.from\_state(state)  
File “/usr/local/lib/python3.10/site-packages/ray/rllib/algorithms/algorithm.py”, line 306, in from\_state  
new\_algo = algorithm\_class(config=config)  
File “/usr/local/lib/python3.10/site-packages/ray/rllib/algorithms/algorithm.py”, line 368, in **init**  
config.validate()  
File “/usr/local/lib/python3.10/site-packages/ray/rllib/algorithms/ppo/ppo.py”, line 222, in validate  
super().validate()  
File “/usr/local/lib/python3.10/site-packages/ray/rllib/algorithms/pg/pg.py”, line 91, in validate  
super().validate()  
File “/usr/local/lib/python3.10/site-packages/ray/rllib/algorithms/algorithm\_config.py”, line 556, in validate  
self.\_resolve\_tf\_settings(\_tf1, \_tfv)  
File “/usr/local/lib/python3.10/site-packages/ray/rllib/algorithms/algorithm\_config.py”, line 2490, in \_resolve\_tf\_settings  
\_tf1.enable\_eager\_execution()  
File “/usr/local/lib/python3.10/site-packages/tensorflow/python/framework/ops.py”, line 6155, in enable\_eager\_execution  
return enable\_eager\_execution\_internal(  
File “/usr/local/lib/python3.10/site-packages/tensorflow/python/framework/ops.py”, line 6223, in enable\_eager\_execution\_internal  
raise ValueError(  
ValueError: tf.enable\_eager\_execution must be called at program startup.

---

<div class="post-metadata">

**Author:** ![Lars\_Simon\_Zehnder](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lars_simon_zehnder/32/1185_2.png) [@Lars\_Simon\_Zehnder](https://discuss.ray.io/u/Lars_Simon_Zehnder)\
**Post date:** [February 19, 2023, 10:52pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/15 "2023-02-19T22:52:25Z")

</div>

@arturn This error has to do with the `_resolve_tf_settings()` from the `AlgorithmConfig`. It checks for `tf1` if `eager_execution` is enabled, if not it calls on `tf1` ènable\_eager\_execution()` which creates this error.

Can we change this somehow?

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [March 8, 2023, 8:09pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/16 "2023-03-08T20:09:53Z")

</div>

Are you getting this error on master? Because I can execute it without errors it seems.

---

<div class="post-metadata">

**Author:** ![Lars\_Simon\_Zehnder](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lars_simon_zehnder/32/1185_2.png) [@Lars\_Simon\_Zehnder](https://discuss.ray.io/u/Lars_Simon_Zehnder)\
**Post date:** [March 8, 2023, 8:24pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/17 "2023-03-08T20:24:13Z")

</div>

@arturn No, I got this error on Ray 2.2.0 as was also AI360 - this might have been fixed in the master and already in Ray 2.3.0

---

<div class="post-metadata">

**Author:** ![Lars\_Simon\_Zehnder](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lars_simon_zehnder/32/1185_2.png) [@Lars\_Simon\_Zehnder](https://discuss.ray.io/u/Lars_Simon_Zehnder)\
**Post date:** [March 8, 2023, 8:28pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/18 "2023-03-08T20:28:39Z")

</div>

@AI360 do you call `from_checkpoint()` in a Jupyter notebook? Also could you just try to put at the start of your code:

```auto
from ray.rllib.utils.framework import try_import_tf
tf1, tf, tfv = try_import_tf()

```

This at least should make your code work, I guess.

---

<div class="post-metadata">

**Author:** ![marc\_hollyoak](https://avatars.discourse-cdn.com/v4/letter/m/45deac/32.png) [@marc\_hollyoak](https://discuss.ray.io/u/marc_hollyoak)\
**Post date:** [June 5, 2023, 9:14am UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/19 "2023-06-05T09:14:44Z")

</div>

Good morning. Just to confirm, I am still getting this error on 2.4.0.

Code (in .py):  
checkpoint = algo.save()  
restored\_algo = Algorithm.from\_checkpoint(checkpoint)

Error:  
2023-06-05 09:46:28,769 WARNING checkpoints.py:109 – No `rllib_checkpoint.json` file found in checkpoint directory /home/marc/ray\_results/DQN\_GymEnvironment\_2023-06-05\_09-39-47ykvo2ont/checkpoint\_000010! Trying to extract checkpoint info from other files found in that dir.

Just checking - is there fix due imminently please?  
Many thanks.

---

<div class="post-metadata">

**Author:** ![hermmanhender](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hermmanhender/32/5691_2.png) [@hermmanhender](https://discuss.ray.io/u/hermmanhender)\
**Post date:** [October 20, 2023, 3:53pm UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/20 "2023-10-20T15:53:00Z")

</div>

Hi, I used the following and works, maybe it helps:

```auto
from ray.rllib.utils.framework import try_import_tf
tf1, tf, _ = try_import_tf()
tf1.enable_eager_execution()

```

Best regards.

---

<div class="post-metadata">

**Author:** ![fardinabbasi](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/fardinabbasi/32/5085_2.png) [@fardinabbasi](https://discuss.ray.io/u/fardinabbasi)\
**Post date:** [October 21, 2023, 4:54am UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259/21 "2023-10-21T04:54:59Z")

</div>

Hi @arturn ,  
I used `Algorithm.from_checkpoint` , but I still encountered the same issue. Would you please take a look at this [post](https://discuss.ray.io/t/restoring-agent-with-ray-tune/12525).
