# Log multi agent rewards from policy\_client

**URL:** <https://discuss.ray.io/t/log-multi-agent-rewards-from-policy-client/5642>\
**Category:** RLlib\
**Created:** [April 1, 2022, 9:21am UTC](https://discuss.ray.io/t/log-multi-agent-rewards-from-policy-client/5642 "2022-04-01T09:21:41Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Blubberblub](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/blubberblub/32/1396_2.png) [@Blubberblub](https://discuss.ray.io/u/Blubberblub)\
**Post date:** [April 1, 2022, 9:21am UTC](https://discuss.ray.io/t/log-multi-agent-rewards-from-policy-client/5642/1 "2022-04-01T09:21:41Z")

</div>

How severe does this issue affect your experience of using Ray?

- High: It blocks me to complete my task.

**Situation:** I have a multi-agent env that runs inside a policy\_client(in local mode) that reports to a policy\_server. I adapted the cartpole examples to my needs and everything runs ok.  
**What works:** Everything works fine except reward loging to the server.  
**Problem:** I don’t understand how the policy\_client handles multi-agent rewards. For get\_action() i can pass a MultiAgentDict observation but log\_returns() only takes a float value as reward input. When passing a reward dict no rewards show in tensorboard.

from rllib/env/policy\_client.py

```auto
    @PublicAPI
    def log_returns(
        self,
        episode_id: str,
        reward: float,
        info: Union[EnvInfoDict, MultiAgentDict] = None,
        multiagent_done_dict: Optional[MultiAgentDict] = None,
    )

```

**Question:** How am i supposed to return a multi-agent reward dict to the server?

---

<div class="post-metadata">

**Author:** ![Blubberblub](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/blubberblub/32/1396_2.png) [@Blubberblub](https://discuss.ray.io/u/Blubberblub)\
**Post date:** [April 7, 2022, 12:09pm UTC](https://discuss.ray.io/t/log-multi-agent-rewards-from-policy-client/5642/2 "2022-04-07T12:09:18Z")

</div>

Ok the basic cartpole client/server example also doesn’t make rewards and dones(at least not correctly) visible in tensorboard…  
I had an additional look into the policy\_client source code and here is what i saw:

1. The policy client seems to have an internal environment (self.env) that wraps a RandomEnv or RandomMultiAgentEnv in an ExternalEnv/ExternalMultiAgent env, if you leave the env in the server config empty (the usual case)
2. The log\_returns function from the policy\_client then only calls the log\_returns function of this ExternalMultiAgentEnv, if you do local and not remote

Some more in detail explanation on how ExternalEnv and MultiAgentExternalEnv work from there and interact with the server would be really helpful!
