# Custom Critic (Value\_function) in PPO

**URL:** <https://discuss.ray.io/t/custom-critic-value-function-in-ppo/1204>\
**Category:** RLlib\
**Created:** [March 10, 2021, 8:29pm UTC](https://discuss.ray.io/t/custom-critic-value-function-in-ppo/1204 "2021-03-10T20:29:49Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![CodingBurmer](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/codingburmer/32/229_2.png) [@CodingBurmer](https://discuss.ray.io/u/CodingBurmer)\
**Post date:** [March 10, 2021, 8:29pm UTC](https://discuss.ray.io/t/custom-critic-value-function-in-ppo/1204/1 "2021-03-10T20:29:49Z")

</div>

Hey folks,

am right in assuming that if I want to implement a custom actor and critic network for the PPO algorithm (like Actor-Critc Type #2 in the figure), I just have to implement the Actor-Network in the  
forward() method and the Critc-Network in the value\_function() method of my custom Model?

Also, could this architecture cause problems in a MARL scenario?

Thx for every kind of help 😊

 ![ppo](https://us1.discourse-cdn.com/flex020/uploads/ray/original/1X/d1cdfe6c8665485024a140e528a29cce193efc1d.png)

```
class CustomActorCritic(TorchModelV2):
  def __init__ (self, obs_space, action_space, num_outputs, model_config, name): 
    policy_network = x
    value_network = y

  def forward(self, input_dict, state, seq_lens): 
    action = policy_network(obs)
    return action

  def value_function(self):
    q_value = value_network(obs)
    return q_value

ModelCatalog.register_custom_model("my_torch_model", CustomActorCritic)

ray.init()
trainer = ppo.PPOTrainer(env="CartPole-v0", config={
    "framework": "torch",
    "model": {
        "custom_model": "my_torch_model",
    },
})
```

---

<div class="post-metadata">

**Author:** ![Ameer\_Haj\_Ali](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ameer_haj_ali/32/279_2.png) [@Ameer\_Haj\_Ali](https://discuss.ray.io/u/Ameer_Haj_Ali)\
**Post date:** [March 10, 2021, 9:26pm UTC](https://discuss.ray.io/t/custom-critic-value-function-in-ppo/1204/2 "2021-03-10T21:26:42Z")

</div>

Thanks @CodingBurmer! We appreciate your intention to contribute.  
CC @sven1977  
If you don’t get a response for a while please tag me and I will try to help.

---

<div class="post-metadata">

**Author:** ![sven1977](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sven1977/32/53_2.png) [@sven1977](https://discuss.ray.io/u/sven1977)\
**Post date:** [March 11, 2021, 8:22am UTC](https://discuss.ray.io/t/custom-critic-value-function-in-ppo/1204/3 "2021-03-11T08:22:15Z")

</div>

Hey @CodingBurmer , that looks all correct. In the MARL case, you would have to see, whether you would want a so called “centralized critic”, which takes in observations for all (or at least some) agents, instead of just the “own” one. You can check this example script where we override the postprocess\_trajectory function to manipulate the train\_batch to add all other agents’ observations.

`ray/rllib/examples/centralized_critic.py`

---

<div class="post-metadata">

**Author:** ![CodingBurmer](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/codingburmer/32/229_2.png) [@CodingBurmer](https://discuss.ray.io/u/CodingBurmer)\
**Post date:** [March 11, 2021, 9:15am UTC](https://discuss.ray.io/t/custom-critic-value-function-in-ppo/1204/4 "2021-03-11T09:15:54Z")

</div>

Hey @sven1977 THX for your answer 🙂 I already have a working PPO with a “centralized critic”, but THX for your proposal. Right now I’m trying to add a bidirectional RNN communication layer to the PPO model.
