# Letting remote function use all CPUs?

**URL:** <https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153>\
**Category:** Ray Core\
**Created:** [March 7, 2021, 12:28am UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153 "2021-03-07T00:28:56Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![choong](https://avatars.discourse-cdn.com/v4/letter/c/9f8e36/32.png) [@choong](https://discuss.ray.io/u/choong)\
**Post date:** [March 7, 2021, 12:28am UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153/1 "2021-03-07T00:28:56Z")

</div>

I’m experimenting with using Ray to offload compute from within Jupyter notebooks. Some functions will be more efficient if I let them do their own threading (e.g. for RAM or cache reasons) and that will also improve interactive latency. I’m hoping for a way to mark those as remote and have Ray schedule them so they take up all CPUs on whichever host they land on. Is there a way to do this?

I’ve checked to see if num\_cpus=-1 get special treatment (no) and considered adding a new pseudo-accelerator (kludge and doesn’t really solve the problem). Am I missing something obvious?

Thanks!

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 8, 2021, 12:20am UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153/2 "2021-03-08T00:20:16Z")

</div>

> [@choong](#):
>
> I’m experimenting with using Ray to offload compute from within Jupyter notebooks. Some functions will be more efficient if I let them do their own threading (e.g. for RAM or cache reasons) and that will also improve interactive latency. I’m hoping for a way to mark those as remote and have Ray schedule them so they take up all CPUs on whichever host they land on. Is there a way to do this?

Do you know the number of CPUs are available on each host in your cluster?

---

<div class="post-metadata">

**Author:** ![choong](https://avatars.discourse-cdn.com/v4/letter/c/9f8e36/32.png) [@choong](https://discuss.ray.io/u/choong)\
**Post date:** [March 8, 2021, 12:36am UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153/3 "2021-03-08T00:36:29Z")

</div>

Locally yes, and AWS yes (but a different #). I’m hoping to avoid the notebooks needing to know the worker host details (and updating that by hand when I switch environments) though I can see how that would be a workaround.

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 8, 2021, 9:15pm UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153/4 "2021-03-08T21:15:06Z")

</div>

Hypothetically you could use `ray.cluster_resources()` to programmatically determine the worker host details?

---

<div class="post-metadata">

**Author:** ![choong](https://avatars.discourse-cdn.com/v4/letter/c/9f8e36/32.png) [@choong](https://discuss.ray.io/u/choong)\
**Post date:** [March 8, 2021, 9:58pm UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153/5 "2021-03-08T21:58:15Z")

</div>

That won’t work because `ray.cluster_resources()` gives you the totals but not the per-host details. It turns out an easy kludge is to make a new resource type “node” that is 1.0 per node, a little ugly but gets the right behavior and could be put into a wrapper library.

```auto
@ray.remote(resources={'node': 0.001})
def foo(x):
    time.sleep(1+random.random()*0.01)
    return x

@ray.remote(resources={'node': 1})
def bar(x):
    time.sleep(1+random.random()*0.01)
    return x

```

Above foo() will bottleneck on cores while bar() won’t be scheduled until it has the node all to itself.

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 9, 2021, 8:22am UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153/6 "2021-03-09T08:22:08Z")

</div>

Another option is to do:

```auto
node_ids = {node_id for node_id in ray.cluster_resources() if node_id.startswith("node:"}
@ray.remote
def func(...):
   pass

[func.options(resources={n: 1}).remote() for n in node_ids]

```

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [March 9, 2021, 6:32pm UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153/7 "2021-03-09T18:32:32Z")

</div>

cc @simon-mo don’t you have a way to see the per-host information?

---

<div class="post-metadata">

**Author:** ![choong](https://avatars.discourse-cdn.com/v4/letter/c/9f8e36/32.png) [@choong](https://discuss.ray.io/u/choong)\
**Post date:** [March 9, 2021, 11:56pm UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153/8 "2021-03-09T23:56:09Z")

</div>

To check my understanding: This would fire off one invocation pinned to each node?

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 9, 2021, 11:59pm UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153/9 "2021-03-09T23:59:56Z")

</div>

Yeah, that is right.

---

<div class="post-metadata">

**Author:** ![choong](https://avatars.discourse-cdn.com/v4/letter/c/9f8e36/32.png) [@choong](https://discuss.ray.io/u/choong)\
**Post date:** [March 10, 2021, 12:24am UTC](https://discuss.ray.io/t/letting-remote-function-use-all-cpus/1153/10 "2021-03-10T00:24:38Z")

</div>

Ok thanks. It’s probably not spending more time polishing this but the case where I end up wanting to do this is in compressing video after generating an image sequence. Easiest path is to run ffmpeg as a subprocess, letting it use all available cores is desirable so I can see the output sooner.
