# Restore and serve from remote checkpoit

**URL:** <https://discuss.ray.io/t/restore-and-serve-from-remote-checkpoit/6170>\
**Category:** Ray Serve\
**Created:** [May 17, 2022, 2:30pm UTC](https://discuss.ray.io/t/restore-and-serve-from-remote-checkpoit/6170 "2022-05-17T14:30:55Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![peterhaddad3121](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/peterhaddad3121/32/2547_2.png) [@peterhaddad3121](https://discuss.ray.io/u/peterhaddad3121)\
**Post date:** [May 17, 2022, 2:30pm UTC](https://discuss.ray.io/t/restore-and-serve-from-remote-checkpoit/6170/1 "2022-05-17T14:30:55Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- Medium: It contributes to significant difficulty to complete my task, but I can work around it.

I have used Ray Tune to train and sync checkpoints to S3.

I want to restore a PPO Agent with a checkpoint stored in S3.

The Trainable class in [ray/trainable.py at ray-1.12.0 · ray-project/ray · GitHub](https://github.com/ray-project/ray/blob/ray-1.12.0/python/ray/tune/trainable.py)

has the following code:

```python
    def restore(self, checkpoint_path):
        """Restores training state from a given model checkpoint.
        These checkpoints are returned from calls to save().
        Subclasses should override ``load_checkpoint()`` instead to
        restore state.
        This method restores additional metadata saved with the checkpoint.
        `checkpoint_path` should match with the return from ``save()``.
        `checkpoint_path` can be
        `~/ray_results/exp/MyTrainable_abc/
        checkpoint_00000/checkpoint`. Or,
        `~/ray_results/exp/MyTrainable_abc/checkpoint_00000`.
        `self.logdir` should generally be corresponding to `checkpoint_path`,
        for example, `~/ray_results/exp/MyTrainable_abc`.
        `self.remote_checkpoint_dir` in this case, is something like,
        `REMOTE_CHECKPOINT_BUCKET/exp/MyTrainable_abc`

```

However, I am unsure the best practice for implementing when using the following Serve deployment:

```python
from ray import serve
from starlette.requests import Request
import ray.rllib.agents.ppo as ppo

@serve.deployment(route_prefix="/cartpole-ppo")
class ServePPOModel:
    def __init__ (self, checkpoint_path) -> None:
        self.trainer = ppo.PPOTrainer(
            config={
                "framework": "torch",
                "num_workers": 0,
            },
            env="CartPole-v0",
        )
        self.uses_cloud_checkpointing = True
        self.remote_checkpoint_dir = checkpoint_path
        
        self.trainer.restore(checkpoint_path)

    async def __call__ (self, request: Request):
        json_input = await request.json()
        obs = json_input["observation"]

        action = self.trainer.compute_single_action(obs)
        return {"action": int(action)}

```

Are there best practices for doing this?

---

<div class="post-metadata">

**Author:** ![simon-mo](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/simon-mo/32/9_2.png) [@simon-mo](https://discuss.ray.io/u/simon-mo)\
**Post date:** [May 23, 2022, 7:55pm UTC](https://discuss.ray.io/t/restore-and-serve-from-remote-checkpoit/6170/2 "2022-05-23T19:55:27Z")

</div>

Hi @peterhaddad3121, thank you for your question!

I want to share two recommendation here:

- In the short term, your code looks great! It follows the pattern we recommend in our [Serve RLlib Tutorial](https://docs.ray.io/en/latest/serve/tutorials/rllib.html).
- We have recently introduced [Ray AI Runtime](https://docs.ray.io/en/master/ray-air/getting-started.html#) which unifies training → serving on top of Ray. Here’s an end to end example for serving RLlib models: [Serving reinforcement learning policy models — Ray 3.0.0.dev0](https://docs.ray.io/en/master/ray-air/examples/rl_serving_example.html)
