# How do i compute an action from a trained RLlib PPO policy with the new API?

**URL:** <https://discuss.ray.io/t/how-do-i-compute-an-action-from-a-trained-rllib-ppo-policy-with-the-new-api/23314>\
**Category:** RLlib\
**Created:** [November 10, 2025, 4:59pm UTC](https://discuss.ray.io/t/how-do-i-compute-an-action-from-a-trained-rllib-ppo-policy-with-the-new-api/23314 "2025-11-10T16:59:43Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![jsftexas](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jsftexas/32/8355_2.png) [@jsftexas](https://discuss.ray.io/u/jsftexas)\
**Post date:** [November 10, 2025, 4:59pm UTC](https://discuss.ray.io/t/how-do-i-compute-an-action-from-a-trained-rllib-ppo-policy-with-the-new-api/23314/1 "2025-11-10T16:59:43Z")

</div>

I’m trying to get a computed action from a trained PPO model, but all of the examples I’ve found are from the old API, not the new one.

What is the new API equivalent to this?

action = my\_trained\_policy.compute\_single\_action(obs)

---

<div class="post-metadata">

**Author:** ![RunLLM](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/runllm/32/8249_2.png) [@RunLLM](https://discuss.ray.io/u/RunLLM)\
**Post date:** [November 10, 2025, 5:00pm UTC](https://discuss.ray.io/t/how-do-i-compute-an-action-from-a-trained-rllib-ppo-policy-with-the-new-api/23314/2 "2025-11-10T17:00:08Z")

</div>

In the new RLlib API stack, you should use the RLModule’s forward\_inference method instead of compute\_single\_action. Specifically, after loading your RLModule from a checkpoint, call rl\_module.forward\_inference({“obs”: torch\_obs\_batch}) to get action logits, then sample or select an action from the resulting distribution. The old compute\_single\_action API is deprecated and not available in the new stack. See the [official example here](https://github.com/ray-project/ray/blob/master/rllib/examples/inference/policy_inference_after_training.py) and [discussion in this issue](https://github.com/ray-project/ray/issues/44475).

Would you like a step-by-step code example for this process?

Sources:

- [policy\_inference\_after\_training.py](https://github.com/ray-project/ray/blob/master/rllib/examples/inference/policy_inference_after_training.py)
- [Ray GitHub Issue #44475](https://github.com/ray-project/ray/issues/44475)

Hint: Mention @RunLLM in the post for followups.

---

<div class="post-metadata">

**Author:** ![jsftexas](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jsftexas/32/8355_2.png) [@jsftexas](https://discuss.ray.io/u/jsftexas)\
**Post date:** [November 10, 2025, 6:06pm UTC](https://discuss.ray.io/t/how-do-i-compute-an-action-from-a-trained-rllib-ppo-policy-with-the-new-api/23314/3 "2025-11-10T18:06:45Z")

</div>

A step-by-step code example would be great. I have a continuous action space. So I think I need to sample from the distribution. Will the sampled values correspond to “recommended” values that optimize the reward?
