# How to run a function exactly once on each node?

**URL:** <https://discuss.ray.io/t/how-to-run-a-function-exactly-once-on-each-node/2178>\
**Category:** Ray Core\
**Created:** [May 17, 2021, 5:09pm UTC](https://discuss.ray.io/t/how-to-run-a-function-exactly-once-on-each-node/2178 "2021-05-17T17:09:01Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![dutch](https://avatars.discourse-cdn.com/v4/letter/d/85e7bf/32.png) [@dutch](https://discuss.ray.io/u/dutch)\
**Post date:** [May 17, 2021, 5:09pm UTC](https://discuss.ray.io/t/how-to-run-a-function-exactly-once-on-each-node/2178/1 "2021-05-17T17:09:01Z")

</div>

Say I have a function `copy_data` that copies data from an outside source (e.g., s3, gs, database) to a VM, and I want to run that function exactly once on each node in a cluster. Is there a good way to do that?

If I have, say, 5 nodes, and each node has 32 cores, I’ve been decorating `copy_data` with `@ray.remote(num_cpus=32)` and using ray to run it 5 times, but I thought there might be a better way.

If it helps, copying data is just one application for this functionality.

Thanks.

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [May 17, 2021, 9:03pm UTC](https://discuss.ray.io/t/how-to-run-a-function-exactly-once-on-each-node/2178/2 "2021-05-17T21:03:36Z")

</div>

Hi @dutch,

Check out this recent post for one way to address this problem.

> [@How to create @ray.remote jobs that will only run on the workers from the local node?](https://discuss.ray.io/t/how-to-create-ray-remote-jobs-that-will-only-run-on-the-workers-from-the-local-node/868/2):
>
> One possible solution is to specify the custom resources on a node that has files; ray start --resources=’{“files”:1}’ @ray.remote def f(): pass ray.get(f.options(resources={"file":0.001})).remote()

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [May 17, 2021, 9:18pm UTC](https://discuss.ray.io/t/how-to-run-a-function-exactly-once-on-each-node/2178/3 "2021-05-17T21:18:35Z")

</div>

Thanks @mannyv!! On top of that, what you can do is to use the placement group with a STRICT\_SPREAD strategy, and schedule actors on all nodes that can perform tasks you’d like to perform.

---

<div class="post-metadata">

**Author:** ![dutch](https://avatars.discourse-cdn.com/v4/letter/d/85e7bf/32.png) [@dutch](https://discuss.ray.io/u/dutch)\
**Post date:** [May 17, 2021, 10:21pm UTC](https://discuss.ray.io/t/how-to-run-a-function-exactly-once-on-each-node/2178/4 "2021-05-17T22:21:34Z")

</div>

Thanks, guys. I’m looking at the documentation. If I have a function `get_ip_addr` that I want to run once on each node, do I do something like the following:

```
bundles = [{"CPU": 16} for _ in ray.nodes()]
pg = placement_group(bundles=bundles, strategy="STRICT_SPREAD")
ray.get(pg.ready())
tasks = [
    get_ip_addr.options(placement_group=pg, placement_group_bundle_index=i).remote()
    for i in range(len(ray.nodes()))
]
ips = ray.get(tasks)
```

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [May 18, 2021, 4:08am UTC](https://discuss.ray.io/t/how-to-run-a-function-exactly-once-on-each-node/2178/5 "2021-05-18T04:08:37Z")

</div>

My recommendation is

```auto
@ray.remote(num_cpus=1)
class FunctionExecutor
def get_ip_addr(self):
    return "haha" # write your code

num_nodes = len(ray.nodes())
bundles = [{"CPU": 1} for _ in num_nodes]
pg = placement_group(bundles=bundles, strategy="STRICT_SPREAD")
ray.get(pg.ready())
executors = [FunctionExecutor.options(placement_group=pg).remote() for num_nodes]
ips = ray.get([executor.get_ip_addr.remote() for executor in executors])

```

Note that the placement group pre-reserves resources, so if you allocate 16 cpus for each bundle, it will probably take up the whole resources in every machine in your cluster (so you cannot execute other functions). If you are only looking for executing one functions with the placement group, you can use your approach
