# Question about multi agent linked to the same policy

**URL:** <https://discuss.ray.io/t/question-about-multi-agent-linked-to-the-same-policy/3744>\
**Category:** RLlib\
**Created:** [October 7, 2021, 11:52am UTC](https://discuss.ray.io/t/question-about-multi-agent-linked-to-the-same-policy/3744 "2021-10-07T11:52:27Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [October 7, 2021, 12:19pm UTC](https://discuss.ray.io/t/question-about-multi-agent-linked-to-the-same-policy/3744/2 "2021-10-07T12:19:10Z")

</div>

Hi @carlorop,

The short answer is that agents that map to the same policy use the same models and buffer.

Here is a slightly more involved answer:

> [@Adding virtual agents in MARL](https://discuss.ray.io/t/adding-virtual-agents-in-marl/3708/2):
>
> Hi @Aceticia, Your idea is good. Any number of agents can share the same policy. Each will use the policy independently during execution (sampling rollouts). During training if you do not add any centralizing peices like for example the centralized critic, or an algorithm like qmix or maddpg then each transition is considered separately for each agent. During the actual loss calculations, the losses are computed in separate batches based on policy not agent. So if you have 3 agents that all m…

Does it make sense?

---

_[View the full topic](https://discuss.ray.io/t/question-about-multi-agent-linked-to-the-same-policy/3744)._
