# Distributed multi-agent training

**URL:** <https://discuss.ray.io/t/distributed-multi-agent-training/5134>\
**Category:** RLlib\
**Created:** [February 20, 2022, 1:29am UTC](https://discuss.ray.io/t/distributed-multi-agent-training/5134 "2022-02-20T01:29:32Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![limenutt](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/limenutt/32/2180_2.png) [@limenutt](https://discuss.ray.io/u/limenutt)\
**Post date:** [February 20, 2022, 1:29am UTC](https://discuss.ray.io/t/distributed-multi-agent-training/5134/1 "2022-02-20T01:29:32Z")

</div>

I have a multi-agent environment (1 env ~ 10 agents) that is wall-time-expensive, CPU-only and it’s tricky to run multiple instances of environment on one machine.  
I want to run about 10 small machines, 1 environment on each, which will give me ~100 agents to step through at the time. I want all of them to train a **single policy** (they are independent agents, do not interact with each other at all)

My intuition is that I would need to do the inference and training on one big machine with a GPU (e.g. on a head node), but open to other experiments, like doing inference locally but training centrally and syncing the policy from a head node to worker nodes every time it changes.

What would be the best way to do it with ray clusters?
