# How can I deploy my reinforcement learning model trained with tune using the new API?

**URL:** <https://discuss.ray.io/t/how-can-i-deploy-my-reinforcement-learning-model-trained-with-tune-using-the-new-api/23081>\
**Category:** Checkpointing, Restoring\
**Created:** [September 7, 2025, 10:28am UTC](https://discuss.ray.io/t/how-can-i-deploy-my-reinforcement-learning-model-trained-with-tune-using-the-new-api/23081 "2025-09-07T10:28:46Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Appleleo](https://avatars.discourse-cdn.com/v4/letter/a/e495f1/32.png) [@Appleleo](https://discuss.ray.io/u/Appleleo)\
**Post date:** [September 7, 2025, 10:28am UTC](https://discuss.ray.io/t/how-can-i-deploy-my-reinforcement-learning-model-trained-with-tune-using-the-new-api/23081/1 "2025-09-07T10:28:46Z")

</div>

I trained a reinforcement learning model using ray.tune and the PPO algorithm. A series of checkpoints were generated. When I tried to restore the **rlmodule** from the checkpoints using the example method and perform inference, I found that the accumulated reward was much smaller than the **episode\_return\_mean** given in progress.csv. Here is the script I used to apply the policy:

```python
    rl_module = RLModule.from_checkpoint(os.path.join(
            agentfile+'/checkpoint_000'+str(checkpoint)+'/',
            "learner_group",
            "learner",
            "rl_module",
            DEFAULT_MODULE_ID,
        )
    )

    env = gym.make(config["env"], config=config["env_config"])

    obs, info = env.reset()

    num_episodes = 0
    max_episodes = 9999
    episode_return = 0.0
    max_reward = 0

    while num_episodes < max_episodes:
        input_dict = {Columns.OBS: torch.from_numpy(obs).unsqueeze(0)}

        rl_module_out = rl_module.forward_inference(input_dict)

        action_dist_params = rl_module_out["action_dist_inputs"][0].numpy()
        greedy_action = np.clip(
            action_dist_params[0:22], 
            a_min=env.action_space.low[0],
            a_max=env.action_space.high[0],
        )

        obs, reward, terminated, truncated, _ = env.step(greedy_action)
        episode_return += reward

        if terminated or truncated:

            if to_save and episode_return > max_reward:
                max_reward = episode_return

            print('========')
            print(f"Episode done: Total reward = {episode_return}")
            obs, info = env.reset()
            num_episodes += 1
            episode_return = 0.0

```

In fact, in the old API, my environment can get reasonable rewards accumulation through **compute\_action**. However, in the new version, **compute\_single\_action** is disabled. I noticed that when **SingleAgentEnvRunner** performs sampling, whether during env reset or step, the **Connector** is frequently used. This means that the **action\_dist** obtained by **forward\_inference** needs to be processed multiple times before it can be used as the input of the env:

```python
#single_agent_env_runner.py Line 319  

                # Module-to-env connector.
                to_env = self._module_to_env(
                    rl_module=self.module,
                    batch=to_env,
                    episodes=episodes,
                    explore=explore,
                    shared_data=shared_data,
                    metrics=self.metrics,
                    metrics_prefix_key=(MODULE_TO_ENV_CONNECTOR,),
                )

```

Is this the reason why I can’t deploy the strategy correctly? Is there a more elegant solution than calling **\_module\_to\_env**?

Thanks so much!

---

<div class="post-metadata">

**Author:** ![MCW\_Lad](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mcw_lad/32/8009_2.png) [@MCW\_Lad](https://discuss.ray.io/u/MCW_Lad)\
**Post date:** [September 8, 2025, 7:17am UTC](https://discuss.ray.io/t/how-can-i-deploy-my-reinforcement-learning-model-trained-with-tune-using-the-new-api/23081/2 "2025-09-08T07:17:27Z")

</div>

I’ve got working inference code up [here](https://github.com/MatthewCWeston/rllib_sw). classes/inference\_helpers.py should have what you’re looking for.

---

<div class="post-metadata">

**Author:** ![Appleleo](https://avatars.discourse-cdn.com/v4/letter/a/e495f1/32.png) [@Appleleo](https://discuss.ray.io/u/Appleleo)\
**Post date:** [September 10, 2025, 6:36am UTC](https://discuss.ray.io/t/how-can-i-deploy-my-reinforcement-learning-model-trained-with-tune-using-the-new-api/23081/3 "2025-09-10T06:36:40Z")

</div>

Thanks [MCW\_Lad](https://discuss.ray.io/t/how-can-i-deploy-my-reinforcement-learning-model-trained-with-tune-using-the-new-api/23081/2) !!!

I also found a less elegant way, but it gets the exact result:

```python
    new_ppo = Algorithm.from_checkpoint(checkpoint_dir)

    config = new_ppo.get_config()

    env_runner_group = new_ppo.env_runner_group
    local_env_runner = env_runner_group.local_env_runner

    local_env_runner.config = config
    local_env_runner.make_env()

    episodes = local_env_runner.sample(
        num_episodes=1,
    )

    print(local_env_runner.get_metrics())

    local_env_runner.env.env.envs[0].env.env.CUSTOM_Function()

```

Hopefully there will be a more official package that implements this functionality.

Regards!!!

---

<div class="post-metadata">

**Author:** ![Daraan](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/daraan/32/7391_2.png) [@Daraan](https://discuss.ray.io/u/Daraan)\
**Post date:** [September 10, 2025, 10:25am UTC](https://discuss.ray.io/t/how-can-i-deploy-my-reinforcement-learning-model-trained-with-tune-using-the-new-api/23081/4 "2025-09-10T10:25:43Z")

</div>

I am not entirely sure if its relevant for your case. If not at least its a nice to know, do you know that the `episode_return_mean` is smoothed by `config.metrics_num_episodes_for_smoothing`? See the topic I just have posted:

> [@Change the episode metrics as the timeframe is a bit vague?](https://discuss.ray.io/t/change-the-episode-metrics-as-the-timeframe-is-a-bit-vague/23091):
>
> Currently episode metrics are logged like this with a window in the EnvRunner: win = config.metrics\_num\_episodes\_for\_smoothing logger.log\_value("return\_mean", ret, window=win) logger.log\_value("return\_min", ret, reduce="min", window=win) logger.log\_value("return\_max", ret, reduce="max", window=win) Now I don’t think that is the best way to do it, when we do not known the episodes, especially during evaluation this then is off when win != len(episodes\_seen). Case 1: win \< episo…

In short the min/mean/max you obtain by using `local_env_runner.get_metrics()` are from the last `metrics_num_episodes_for_smoothing` sampled episodes - not bound to an iteration.  
Furthermore (depending on your ray version), restoring metrics is broken, see [[RLlib] Checkpoint metrics loading with Tune is broken in 2.47.0 · Issue #53877 · ray-project/ray · GitHub](https://github.com/ray-project/ray/issues/53877). In your case I think the smoothing from the old episodes (if your windows reaches there) can be off / lost. So possibly you only get the smoothed value from _after_ you loaded the checkpoint.

Maybe you have to cross check if things are maybe correct but not logged like you would expect.  
Cheers, and good luck.
