# Get\_policy error when get an action from restored trained model- New API stack

**URL:** <https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287>\
**Category:** Uncategorized\
**Created:** [April 11, 2025, 8:40am UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287 "2025-04-11T08:40:35Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ali\_Zargarian](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ali_zargarian/32/7814_2.png) [@Ali\_Zargarian](https://discuss.ray.io/u/Ali_Zargarian)\
**Post date:** [April 11, 2025, 8:40am UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/1 "2025-04-11T08:40:36Z")

</div>

…\anaconda3\envs\rllibpy311\Lib\site-packages\ray\rllib\algorithms\algorithm.py", line 2109, in get\_policy  
return self.env\_runner.get\_policy(policy\_id)  
^^^^^^^^^^^^^^^^^^^^^^^^^^  
AttributeError: ‘SingleAgentEnvRunner’ object has no attribute ‘get\_policy’

Hi,  
This error blocks me for a while. As I searched here , there is no reliable answer up to now.  
Is there any standard way(or example) to restore and get an action from a trained model??  
this problem happens when I try to restore my model with PPO and DQN. any idea how it can be solved?

---

<div class="post-metadata">

**Author:** ![christina](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/christina/32/7542_2.png) [@christina](https://discuss.ray.io/u/christina)\
**Post date:** [April 11, 2025, 9:11pm UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/2 "2025-04-11T21:11:13Z")

</div>

Hi there, for debugging purposes, what version of RLlib API are you running right now? And is there a way for me to reproduce this error (do you have a snippet of code to share)?

---

<div class="post-metadata">

**Author:** ![Ali\_Zargarian](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ali_zargarian/32/7814_2.png) [@Ali\_Zargarian](https://discuss.ray.io/u/Ali_Zargarian)\
**Post date:** [April 13, 2025, 5:19pm UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/3 "2025-04-13T17:19:21Z")

</div>

Hi Christina,  
I refer you to this issue, that is the same somehow and still unsolved after long time:

> <https://github.com/ray-project/ray/issues/44475>
>
> \### What happened + What you expected to happen
> 
> After training Multi Agent PPO …with new New API Stack under the guidance of \[how-to-use-the-new-api-stack\](https://docs.ray.io/en/latest/rllib/rllib-new-api-stack.html#how-to-use-the-new-api-stack)
> I tried to compute actions:
> \`\`\`
> saved\_algorithm = Algorithm.from\_checkpoint(
> checkpoint=algorithm\_path,
> policy\_ids={"controlled\_vehicle\_0", "controlled\_vehicle\_1"},
> policy\_mapping\_fn=lambda agent\_id, episode, \*\*kwargs: f"controlled\_vehicle\_{agent\_id}",
> )
> print("saved\_algorithm type:", type(saved\_algorithm))
> # Evaluate the model
> obs, info = env.reset()
> print("obs:", obs)
> actions = {}
> for agent\_id, agent\_obs in obs.items():
> policy\_id = f"controlled\_vehicle\_{agent\_id}"
> action = saved\_algorithm.get\_policy(policy\_id).compute\_single\_action(agent\_obs)
> actions\[agent\_id\] = action
> print("action", actions)
> \`\`\`
> but I get the error message: 
> 
> \> AttributeError: 'MultiAgentEnvRunner' object has no attribute 'get\_policy'
> 
> I also tried some other way like:
> \`action = saved\_algorithm.compute\_single\_action(agent\_obs, policy\_id)\`
> but still get the same error message: AttributeError: 'MultiAgentEnvRunner' object has no attribute 'get\_policy'.
> I have seen a similar issue in #40312, are these two issues the same?
> 
> detailed error message are as follows:
> 
> \> Traceback (most recent call last):
> \> File "test\_evaluate.py", line 151, in \<module\>
> \> evaluate\_agent(saved\_algorithm, env)
> \> File "test\_evaluate.py", line 112, in evaluate\_agent
> \> action = saved\_algorithm.get\_policy(policy\_id).compute\_single\_action(agent\_obs)
> \> File "C:\\Users\\Ice Cream\\miniconda3\\envs\\env\_highway\\lib\\site-packages\\ray\\util\\tracing\\tracing\_helper.py", line 467, in \_resume\_span
> \> return method(self, \*\_args, \*\*\_kwargs)
> \> File "C:\\Users\\Ice Cream\\miniconda3\\envs\\env\_highway\\lib\\site-packages\\ray\\rllib\\algorithms\\algorithm.py", line 2051, in get\_policy
> \> return self.workers.local\_worker().get\_policy(policy\_id)
> \> AttributeError: 'MultiAgentEnvRunner' object has no attribute 'get\_policy'
> 
> and before I call this method, I also printed the relevant info, this part looks normal:
> \> saved\_algorithm type: \<class 'ray.rllib.algorithms.ppo.ppo.PPO'\>
> \> saved\_algorithm.get\_config() \<ray.rllib.algorithms.ppo.ppo.PPOConfig object at 0x0000014B0E4D4370\>
> 
> through the code:
> \`\`\`
> print("saved\_algorithm type:", type(saved\_algorithm))
> print("saved\_algorithm.get\_config()",saved\_algorithm.get\_config())
> \`\`\`
> 
> \### Versions / Dependencies
> 
> Ray 2.10.0
> Python 3.8.18
> Windows11
> 
> \### Reproduction script
> 
> the code used for training is as follows:
> \`\`\`
> register\_env("ray\_dict\_highway\_env", create\_env)
> config = (
> PPOConfig().environment(env="ray\_dict\_highway\_env")
> .experimental(\_enable\_new\_api\_stack=True)
> .rollouts(env\_runner\_cls=MultiAgentEnvRunner)
> .resources(
> num\_learner\_workers=0,
> num\_gpus\_per\_learner\_worker=1,
> num\_cpus\_for\_local\_worker=1,
> )
> .training(model={"uses\_new\_env\_runners": True})
> .multi\_agent(
> policies={
> "controlled\_vehicle\_0",
> "controlled\_vehicle\_1"
> },
> policy\_mapping\_fn=lambda agent\_id, episode, \*\*kwargs: f"controlled\_vehicle\_{agent\_id}",
> )
> .framework("torch")
> )
> current\_script\_directory = os.path.dirname(os.path.abspath(\_\_file\_\_))
> ray\_result\_path = os.path.join(current\_script\_directory, folder\_path, "ray\_results")
> tuner = tune.Tuner(
> "PPO",
> run\_config=RunConfig(
> storage\_path=ray\_result\_path,
> name="2-agent-PPO",
> stop={"timesteps\_total": 5e5}
> ),
> param\_space=config.to\_dict() 
> )
> results = tuner.fit()
> \`\`\`
> And the code for loading checkpoints:
> \`\`\`
> algorithm\_path = r"D:\\DRL\_Project\\DRL\_highway\\experiments\\hw-fast-ma-dict-v0\_rllib-mappo\\2024-04-01\_01-28\\ray\_results\\2-agent-PPO\\PPO\_ray\_dict\_highway\_env\_1c7ab\_00000\_0\_2024-04-01\_01-28-38\\checkpoint\_000000"
> saved\_algorithm = Algorithm.from\_checkpoint(
> checkpoint=algorithm\_path,
> policy\_ids={"controlled\_vehicle\_0", "controlled\_vehicle\_1"},
> policy\_mapping\_fn=lambda agent\_id, episode, \*\*kwargs: f"controlled\_vehicle\_{agent\_id}",
> )
> \`\`\`
> 
> \### Issue Severity
> 
> High: It blocks me from completing my task.

---

<div class="post-metadata">

**Author:** ![christina](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/christina/32/7542_2.png) [@christina](https://discuss.ray.io/u/christina)\
**Post date:** [April 14, 2025, 6:21pm UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/4 "2025-04-14T18:21:58Z")

</div>

Hi Ali,  
Thanks for posting the Github issue - let’s continue to track and discuss it there, in the meantime, if you have any more data about this issue I encourage you to post it in the Github issue link!  
Christina

---

<div class="post-metadata">

**Author:** ![Ali\_Zargarian](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ali_zargarian/32/7814_2.png) [@Ali\_Zargarian](https://discuss.ray.io/u/Ali_Zargarian)\
**Post date:** [April 15, 2025, 6:36am UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/5 "2025-04-15T06:36:24Z")

</div>

Hi Christina,  
as you see there, there wasn’t any reliable solution from April 2024 from Ray team.  
Thanks

---

<div class="post-metadata">

**Author:** ![Mike2](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mike2/32/7898_2.png) [@Mike2](https://discuss.ray.io/u/Mike2)\
**Post date:** [April 20, 2025, 6:55am UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/6 "2025-04-20T06:55:19Z")

</div>

Hello Cristina,  
Is it possible to increase the priority of issue mentioned by Ali\_Zargarian **up to P0**? Or provide an official workaround for the correct model/policy procedure of loading it from the checkpoint?  
Thanks in advance.

---

<div class="post-metadata">

**Author:** ![Mike2](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mike2/32/7898_2.png) [@Mike2](https://discuss.ray.io/u/Mike2)\
**Post date:** [April 20, 2025, 7:06am UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/7 "2025-04-20T07:06:06Z")

</div>

Hello Christina,  
some additional information from my side:

- ray 2.44.1
- new version of RLlib API
- checkpoint that I’m using for restoring model created by tune.Tuner

Please let me know if you need any additional information.  
Thanks in advance.

---

<div class="post-metadata">

**Author:** ![christina](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/christina/32/7542_2.png) [@christina](https://discuss.ray.io/u/christina)\
**Post date:** [April 21, 2025, 11:59pm UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/8 "2025-04-21T23:59:22Z")

</div>

Hi Mike, thank you for the additional info! I’ll go discuss with the team and see if there’s anything that I can do re: the Github issue.

---

<div class="post-metadata">

**Author:** ![Mike2](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mike2/32/7898_2.png) [@Mike2](https://discuss.ray.io/u/Mike2)\
**Post date:** [April 22, 2025, 6:38am UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/9 "2025-04-22T06:38:27Z")

</div>

Hello Cristina, I made a deep dive into source code. Please update documentation. The Github issue stays in the area of deprecated code (old API). You also have outdated documentation.

Here is example of the correct approach for new API.

Example. How to restore rl\_model with weights. IMPALA class can be replaced with other Algorithm class.

```auto
model = IMPALA.from_checkpoint(self.checkpoint_path)
rl_module = model.get_module()

```

Example. How to get action for Discrete action space

```auto
rl_module = self.rl_module
fwd_ins = {"obs": torch.Tensor(observation).unsqueeze(0)}
fwd_outputs = rl_module.forward_inference(fwd_ins)
action_dist_class = rl_module.get_inference_action_dist_cls()
action_dist = action_dist_class.from_logits(
      fwd_outputs["action_dist_inputs"]
)
action = action_dist.sample()[0].numpy()

```

Processing of fwd\_outputs should be appropriate for different type of action space, as well as nn output layer.

---

<div class="post-metadata">

**Author:** ![Ali\_Zargarian](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ali_zargarian/32/7814_2.png) [@Ali\_Zargarian](https://discuss.ray.io/u/Ali_Zargarian)\
**Post date:** [April 22, 2025, 11:21am UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/10 "2025-04-22T11:21:12Z")

</div>

Hello @Mike2  
thanks for sharing infos.  
How is the results from your restored trained model? satisfactory?

---

<div class="post-metadata">

**Author:** ![Mike2](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mike2/32/7898_2.png) [@Mike2](https://discuss.ray.io/u/Mike2)\
**Post date:** [April 22, 2025, 12:30pm UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/11 "2025-04-22T12:30:37Z")

</div>

Hello @Ali_Zargarian.  
The results are good.

I encountered an action prediction accuracy issue during training and after restoring the model from a checkpoint some time ago. The problem was related to data preprocessing (normalization). Once I implemented my own `Scaler` for observations preprocessing, the issue was resolved.

---

<div class="post-metadata">

**Author:** ![Ali\_Zargarian](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ali_zargarian/32/7814_2.png) [@Ali\_Zargarian](https://discuss.ray.io/u/Ali_Zargarian)\
**Post date:** [April 22, 2025, 12:51pm UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/12 "2025-04-22T12:51:00Z")

</div>

@Mike2  
Many Thanks for your rapid reply

---

<div class="post-metadata">

**Author:** ![Mike2](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mike2/32/7898_2.png) [@Mike2](https://discuss.ray.io/u/Mike2)\
**Post date:** [April 22, 2025, 1:01pm UTC](https://discuss.ray.io/t/get-policy-error-when-get-an-action-from-restored-trained-model-new-api-stack/22287/13 "2025-04-22T13:01:56Z")

</div>

Hello @Ali_Zargarian,  
I’m happy to help you, but I have a really busy schedule. Sorry for that.

Please take a look on the response from [simonsays1980](https://github.com/simonsays1980) in the original issue on the Github.

> simonsays1980: The important part is now: RLlib uses pre- and post-processing of data. For the pre-processing (converting for example to a multi-agent batch and applying observation filters (like e.g. `MeanStdFilter`) the `EnvToModulePipeline` is used and for the post-processing the `ModuleToEnvPipeline` is used.

I guess it will help you a lot to make your custom environment more predictable.

@christina, thank you for your assistance.
