# Restricting number of actors on a given node

**URL:** <https://discuss.ray.io/t/restricting-number-of-actors-on-a-given-node/954>\
**Category:** Ray Core\
**Created:** [February 21, 2021, 1:13am UTC](https://discuss.ray.io/t/restricting-number-of-actors-on-a-given-node/954 "2021-02-21T01:13:17Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Yoav](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/yoav/32/456_2.png) [@Yoav](https://discuss.ray.io/u/Yoav)\
**Post date:** [February 21, 2021, 1:13am UTC](https://discuss.ray.io/t/restricting-number-of-actors-on-a-given-node/954/1 "2021-02-21T01:13:18Z")

</div>

Let’s say I have nodes each with K cpus and M gb of RAM.  
And I have actors who each requires m gb of RAM, such that K\*m \> M.  
That is, if I use an actor for each cpu, the node will start swapping, and I don’t want that. Is there a way to instruct ray to respect that, and not assign more actors to a node than what the memory allowed?

things i tried:

- marking an actor as requiring 1.2 cpus. _result_: ray does not allow fractional cpus \> 1.
- marking an actor as requiring a resource such as {“r”: 1.0} and specifying in the configuration file that the machine type has some number of this resource (say, 9.0). this should allow up to 9 actors to run on this node. _result_: i am not sure why, but ray said that no machine can satisfy the requirements.

i am using google cloud, if it matters.

(on a related question, ray seems to know the number of cpus provided by the node type. how does it know that? or does it have to create them first?)

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [February 21, 2021, 2:10am UTC](https://discuss.ray.io/t/restricting-number-of-actors-on-a-given-node/954/2 "2021-02-21T02:10:07Z")

</div>

There’s a way to make your actor respect memory restriction; [https://docs.ray.io/en/latest/memory-management.html#memory-aware-scheduling](https://docs.ray.io/en/latest/memory-management.html#memory-aware-scheduling)

> - marking an actor as requiring a resource such as {“r”: 1.0} and specifying in the configuration file that the machine type has some number of this resource (say, 9.0). this should allow up to 9 actors to run on this node. _result_ : i am not sure why, but ray said that no machine can satisfy the requirements.

Hmm this sounds a bit weird. Can you give me a short script I can try? Google cloud shouldn’t be related to this.

> on a related question, ray seems to know the number of cpus provided by the node type

When ray is started without --num-cpus (e.g., `ray start --address=$HEAD_NODE_ADDR`), we automatically detect the num cpus from the machine.

Lastly, note that by default the actor requires 0 cpu. E.g,

```python3
@ray.remote
class A:
    pass

# is equal to

@ray.remote(num_cpus=0)
class A:
    pass

```

---

<div class="post-metadata">

**Author:** ![Yoav](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/yoav/32/456_2.png) [@Yoav](https://discuss.ray.io/u/Yoav)\
**Post date:** [February 21, 2021, 3:02am UTC](https://discuss.ray.io/t/restricting-number-of-actors-on-a-given-node/954/3 "2021-02-21T03:02:46Z")

</div>

Here is some test code which I expect to spawn two worker machines, and it doesn’t:

```auto
import ray
import time
from ray.util import ActorPool

# my cluster will have 5 resources per node, so I exppect to allocate 2 nodes.

ray.init(address='auto')

@ray.remote(num_cpus=1.0, resources={"r": 1})
class MyActor:
    def __init__ (self):
        pass

    def work(self):
        while True:
            time.sleep(2)

actors = [MyActor.remote() for _ in range(10)]
pool = ActorPool(actors)
for x in pool.map_unordered(lambda a, v: a.work.remote(), list(range(10))):
    pass

```

Here is the relevant part from the config, the rest should be pretty much as in the example (using the ray docker image):

```auto
worker_nodes:
    tags:
      - items: ["allow-all"]
    machineType: n1-standard-64
    resources:
      r: 5
    disks:
      - boot: true
        autoDelete: true
        type: PERSISTENT
        initializeParams:
          diskSizeGb: 40
          sourceImage: projects/deeplearning-platform-release/global/images/family/tf-1-13-cpu

```

---

<div class="post-metadata">

**Author:** ![Yoav](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/yoav/32/456_2.png) [@Yoav](https://discuss.ray.io/u/Yoav)\
**Post date:** [February 21, 2021, 3:07am UTC](https://discuss.ray.io/t/restricting-number-of-actors-on-a-given-node/954/4 "2021-02-21T03:07:33Z")

</div>

this is what i see via `ray monitor`:

## ======== Autoscaler status: 2021-02-20 19:04:53.559852 ======== Node status

Healthy:  
1 ray-legacy-head-node-type  
Pending:  
(no pending nodes)  
Recent failures:  
(no failures)

## Resources

Usage:  
0.00/1.465 GiB object\_store\_memory  
0.0/2.0 CPU  
0.00/4.346 GiB memory

Demands:  
{‘r’: 1.0, ‘CPU’: 1.0}: 10+ pending tasks/actors  
2021-02-20 19:04:53,571 INFO discovery.py:873 – URL being requested: GET [https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=((status+%3D+STAGING)+OR+(status+%3D+PROVISIONING)+OR+(status+%3D+RUNNING))+AND+(labels.ray-cluster-name+%3D+test1)&alt=json](https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=%28%28status+%3D+STAGING%29+OR+%28status+%3D+PROVISIONING%29+OR+%28status+%3D+RUNNING%29%29+AND+%28labels.ray-cluster-name+%3D+test1%29&alt=json)  
2021-02-20 19:04:58,749 INFO discovery.py:873 – URL being requested: GET [https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=((labels.ray-node-type+%3D+worker))+AND+((status+%3D+STAGING)+OR+(status+%3D+PROVISIONING)+OR+(status+%3D+RUNNING))+AND+(labels.ray-cluster-name+%3D+test1)&alt=json](https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=%28%28labels.ray-node-type+%3D+worker%29%29+AND+%28%28status+%3D+STAGING%29+OR+%28status+%3D+PROVISIONING%29+OR+%28status+%3D+RUNNING%29%29+AND+%28labels.ray-cluster-name+%3D+test1%29&alt=json)  
2021-02-20 19:04:58,905 INFO discovery.py:873 – URL being requested: GET [https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=((labels.ray-node-type+%3D+worker))+AND+((status+%3D+STAGING)+OR+(status+%3D+PROVISIONING)+OR+(status+%3D+RUNNING))+AND+(labels.ray-cluster-name+%3D+test1)&alt=json](https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=%28%28labels.ray-node-type+%3D+worker%29%29+AND+%28%28status+%3D+STAGING%29+OR+%28status+%3D+PROVISIONING%29+OR+%28status+%3D+RUNNING%29%29+AND+%28labels.ray-cluster-name+%3D+test1%29&alt=json)  
2021-02-20 19:04:59,065 INFO discovery.py:873 – URL being requested: GET [https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=((labels.ray-node-type+%3D+unmanaged))+AND+((status+%3D+STAGING)+OR+(status+%3D+PROVISIONING)+OR+(status+%3D+RUNNING))+AND+(labels.ray-cluster-name+%3D+test1)&alt=json](https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=%28%28labels.ray-node-type+%3D+unmanaged%29%29+AND+%28%28status+%3D+STAGING%29+OR+%28status+%3D+PROVISIONING%29+OR+%28status+%3D+RUNNING%29%29+AND+%28labels.ray-cluster-name+%3D+test1%29&alt=json)  
2021-02-20 19:04:59,198 INFO discovery.py:873 – URL being requested: GET [https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=((status+%3D+STAGING)+OR+(status+%3D+PROVISIONING)+OR+(status+%3D+RUNNING))+AND+(labels.ray-cluster-name+%3D+test1)&alt=json](https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=%28%28status+%3D+STAGING%29+OR+%28status+%3D+PROVISIONING%29+OR+%28status+%3D+RUNNING%29%29+AND+%28labels.ray-cluster-name+%3D+test1%29&alt=json)  
2021-02-20 19:04:59,354 WARNING resource\_demand\_scheduler.py:642 – The autoscaler could not find a node type to satisfy therequest: [{‘r’: 1.0, ‘CPU’: 1.0}, {‘r’: 1.0, ‘CPU’: 1.0}, {‘r’: 1.0, ‘CPU’: 1.0}, {‘r’: 1.0, ‘CPU’: 1.0}, {‘r’: 1.0, ‘CPU’: 1.0}, {‘r’: 1.0, ‘CPU’: 1.0}, {‘r’: 1.0, ‘CPU’: 1.0}, {‘r’: 1.0, ‘CPU’: 1.0}, {‘r’: 1.0, ‘CPU’: 1.0}, {‘r’: 1.0, ‘CPU’: 1.0}]. If this request is related to placement groups the resource request will resolve itself, otherwise please specify a node type with the necessary resource [https://docs.ray.io/en/master/cluster/autoscaling.html#multiple-node-type-autoscaling](https://docs.ray.io/en/master/cluster/autoscaling.html#multiple-node-type-autoscaling).  
2021-02-20 19:04:59,370 INFO discovery.py:873 – URL being requested: GET [https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=((labels.ray-node-type+%3D+worker))+AND+((status+%3D+STAGING)+OR+(status+%3D+PROVISIONING)+OR+(status+%3D+RUNNING))+AND+(labels.ray-cluster-name+%3D+test1)&alt=json](https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=%28%28labels.ray-node-type+%3D+worker%29%29+AND+%28%28status+%3D+STAGING%29+OR+%28status+%3D+PROVISIONING%29+OR+%28status+%3D+RUNNING%29%29+AND+%28labels.ray-cluster-name+%3D+test1%29&alt=json)  
2021-02-20 19:04:59,529 INFO discovery.py:873 – URL being requested: GET [https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=((status+%3D+STAGING)+OR+(status+%3D+PROVISIONING)+OR+(status+%3D+RUNNING))+AND+(labels.ray-cluster-name+%3D+test1)&alt=json](https://compute.googleapis.com/compute/v1/projects/ai2-israel/zones/us-west1-a/instances?filter=%28%28status+%3D+STAGING%29+OR+%28status+%3D+PROVISIONING%29+OR+%28status+%3D+RUNNING%29%29+AND+%28labels.ray-cluster-name+%3D+test1%29&alt=json)  
2021-02-20 19:04:59,658 INFO autoscaler.py:305 –

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [February 21, 2021, 3:17am UTC](https://discuss.ray.io/t/restricting-number-of-actors-on-a-given-node/954/5 "2021-02-21T03:17:00Z")

</div>

Looks like your node doesn’t contain the custom resource r (based on your monitor output). (which is weird given you passed r:5). cc @Ameer_Haj_Ali @Alex o you guys know what’s the issue here?

NOTE: You can also see the comprehensive monitor status using `ray status` command

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [February 21, 2021, 8:36am UTC](https://discuss.ray.io/t/restricting-number-of-actors-on-a-given-node/954/6 "2021-02-21T08:36:20Z")

</div>

Maybe it’s some weird default overrides. As a guess, maybe you could try doing:

```auto
resources:
   r: 5
   CPU: 5

```

---

<div class="post-metadata">

**Author:** ![Ameer\_Haj\_Ali](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ameer_haj_ali/32/279_2.png) [@Ameer\_Haj\_Ali](https://discuss.ray.io/u/Ameer_Haj_Ali)\
**Post date:** [February 21, 2021, 2:07pm UTC](https://discuss.ray.io/t/restricting-number-of-actors-on-a-given-node/954/7 "2021-02-21T14:07:58Z")

</div>

Hi @Yoav , thanks for filing this issue.  
I would like to refer you to our cluster yaml configuration reference:  
[https://docs.ray.io/en/master/cluster/config.html#cluster-configuration-worker-nodes](https://docs.ray.io/en/master/cluster/config.html#cluster-configuration-worker-nodes)  
The object that you put under `worker_nodes` is of type Node Config ([Cluster YAML Configuration Options — Ray v2.0.0.dev0](https://docs.ray.io/en/master/cluster/config.html#cluster-configuration-node-config-type))  
which translates to the configuration options we pass to the cloud provider (e.g. EC2). What you are looking for is using `available_node_types` with the `resources` field:  
[Cluster YAML Configuration Options — Ray v2.0.0.dev0](https://docs.ray.io/en/master/cluster/config.html#cluster-configuration-available-node-types)

TLDR: You cannot use `resources` in `worker_nodes` because the only fields that can go under `worker_nodes` are the cloud provider configurations, instead you should use `available_node_types`.

---

<div class="post-metadata">

**Author:** ![Yoav](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/yoav/32/456_2.png) [@Yoav](https://discuss.ray.io/u/Yoav)\
**Post date:** [February 21, 2021, 5:44pm UTC](https://discuss.ray.io/t/restricting-number-of-actors-on-a-given-node/954/8 "2021-02-21T17:44:09Z")

</div>

thanks @rliaw and @Ameer_Haj_Ali !

Actually, @rliaw 's solution worked! adding `CPU: num` under resources in the worker-config solved it.

I will try also the `available_node_types` solution, though I’ll admit it is hard for me to follow the syntax via the references there. (would be great if there could be a gcp example file for the multiple-node-types option)
