# Question about resource management in Ray

**URL:** <https://discuss.ray.io/t/question-about-resource-management-in-ray/1772>\
**Category:** Ray Core\
**Created:** [April 18, 2021, 5:43pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772 "2021-04-18T17:43:35Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![javigm98](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/javigm98/32/704_2.png) [@javigm98](https://discuss.ray.io/u/javigm98)\
**Post date:** [April 18, 2021, 5:43pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/1 "2021-04-18T17:43:35Z")

</div>

Hi all! I’m using Ray for Reinforcemnet Learning via RLLib, but I want to know how does Ray manages resources. I have seen that even when saying Ray to only use one CPU (via `ray.init(num_cpus=1)`), it tends to use all the available CPU in the system (that right now is 40 CPUs). So what I want is to execute Ray (and consequently RLlib) using only one of the CPUs of my system (in order to leave the other 39 free to place other tasks on them). I tried to use the function `os.sched_setaffinity(0,{0})` in the script where I call `ray.init(num_cpus=1)` and start my training agent. This visually produces the effect that I wanted to see: only one busy CPU while executing the training, but I have still the doubt of knowing if Ray is scheduling tasks to be developed along 40 CPUs and externally these tasks are forced to be executed all in the same CPU or if Ray knows that and only creates tasks as if it was being executed in a single-CPU machine. I’d like to know also which is the param or config that Ray uses to see how much available resources has available, since I have checked that it is not possible for `num_cpus`initialization param to be used for that, because even when setting this value to be 1 all CPUs are used during ray executions.

So I thank you so much for your answers in advance, what I want to do is only to train an RLLibb PPO agent by using the lowest number possible of CPUs in my sistem (I also have two GPUs that I should use) in order to let the free for other tasks.

PD.: I’m using Ray 1.1.0 and training a PPO agent RLlib agent with TF as underlying framework

---

<div class="post-metadata">

**Author:** ![bill-anyscale](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/bill-anyscale/32/724_2.png) [@bill-anyscale](https://discuss.ray.io/u/bill-anyscale)\
**Post date:** [April 19, 2021, 4:44pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/2 "2021-04-19T16:44:47Z")

</div>

so you want to know how many resources Ray is actually using during execution? and what command to look for to see that? I’m not 100% sure I understand the question so I wanted to confirm.

---

<div class="post-metadata">

**Author:** ![javigm98](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/javigm98/32/704_2.png) [@javigm98](https://discuss.ray.io/u/javigm98)\
**Post date:** [April 19, 2021, 5:19pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/3 "2021-04-19T17:19:12Z")

</div>

Hi @bill-anyscale and thank you so much for your answer. What I want is to see how many resources Ray is managing in each moment and also to see how it is managing them (how it creates tasks according to these resources or what tasks it is scheduling onto these resources). In addition I would also like to know how can I force Ray to use only a part of the available resources in my system. I have seen that for GPU limitation usage tou can set different values to the environment variable CUDA\_VISIBLE\_DEVICES and you can select whether to use or not any of my available GPUs. So what I was looking for was to see if it was an efficient way to do the same thing but with CPUs. My final objetive is to evaluate Ray RLlib PPO training process performance (which consists of a driver and a serie of rollout workers) when executing it minimizing the number of CPUs used (in order to see if it is efficient to place all the Ray jobs in only one CPU and let the rest of the CPUs free for other process to be executed on them). So I want to know if there is any way to tell ray to use only a single (or a reduced group) of CPUs in a multi-CPU system. I have tried to set `num_cpus=1` when calling `ray.init()` but this didn’t seem to produce the desired effect, so what I tried was to force this by using the `os.sched_setaffinity()`function. This, as I said, produced the desired visual effect (when monitorizing CPU usage with htop only one CPU was shown to have tasks being executed in it) but I really don’t know if this is the efficient way to achieve my goal, because maybe Ray is scheduling tasks and threads to be excuted over 40 CPUs and later these tasks are forced to be executed only on one CPU. So my question is double: first, how can I see scheduled Ray tasks and available resources to Ray and that are taked into account when scheduling these tasks, and secondly, how I can efficiently reduce CPU utilization when using Ray for my purpose (RLLib PPO Training).

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [April 20, 2021, 6:25pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/4 "2021-04-20T18:25:25Z")

</div>

cc @sven1977 What’s the recommended way to solve this problem for rllib? Does he need to do something on the Ray layer?

---

<div class="post-metadata">

**Author:** ![bill-anyscale](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/bill-anyscale/32/724_2.png) [@bill-anyscale](https://discuss.ray.io/u/bill-anyscale)\
**Post date:** [April 21, 2021, 4:36pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/5 "2021-04-21T16:36:41Z")

</div>

@sangcho , I think this is at the Ray layer. Shouldn’t a ray.init limit fix this on a local machine? Also can’t you change the tune trials to limit the number running in parallel? cc @amogkam

---

<div class="post-metadata">

**Author:** ![amogkam](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amogkam/32/17_2.png) [@amogkam](https://discuss.ray.io/u/amogkam)\
**Post date:** [April 21, 2021, 4:54pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/6 "2021-04-21T16:54:44Z")

</div>

To answer your first question, the Ray dashboard is your best bet to see the available resources to Ray and what tasks/actors are being scheduled. In order to limit the number of CPUs that Ray uses, setting `num_cpus=1` in your `ray.init` should do the trick (@sangcho can confirm). When you say Ray is using 40 CPUs even though you limited it to 1, how are you measuring this?

---

<div class="post-metadata">

**Author:** ![javigm98](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/javigm98/32/704_2.png) [@javigm98](https://discuss.ray.io/u/javigm98)\
**Post date:** [April 21, 2021, 5:00pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/7 "2021-04-21T17:00:47Z")

</div>

Hi @amogkam and thanks for your answer. What I mean is that even when I initialize ray with num\_cpus=4 (for example) i see tasks being executed in the 40 CPUs of my system and cpu\_util\_percent (metric reported by Ray Tune) isn’t smaller than when I don’t specify these parameter in ray initialization. So, what I wanted to know was how ray limitates resources (for example if it places tasks only on a specific set of CPUs) and how it decides how many tasks create.

And abiut ray dashboard, I’m running Ray in a remote system connected via ssh, so is there any way to see this info there?

Thank you again!

---

<div class="post-metadata">

**Author:** ![amogkam](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amogkam/32/17_2.png) [@amogkam](https://discuss.ray.io/u/amogkam)\
**Post date:** [April 21, 2021, 6:26pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/8 "2021-04-21T18:26:03Z")

</div>

Can you send what your console output looks like?

---

<div class="post-metadata">

**Author:** ![javigm98](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/javigm98/32/704_2.png) [@javigm98](https://discuss.ray.io/u/javigm98)\
**Post date:** [April 21, 2021, 6:56pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/9 "2021-04-21T18:56:50Z")

</div>

Of course! But what do you want me to show? I mean what output (which information) do you want to see? When I before mentioned that I saw tasks being executed in the 40 CPUs I was referring to monitorizing CPU usage via htop.

Thanks again!

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [April 22, 2021, 4:38pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/10 "2021-04-22T16:38:08Z")

</div>

@javigm98 are you using the core ray or library?

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [April 22, 2021, 4:39pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/11 "2021-04-22T16:39:55Z")

</div>

@javigm98 actually, Ray doesn’t ensure the cpu affinity. That says, although you set 4 cpus, it doesn’t mean it will actually use 4 cpus of your machine. Also note that Ray has many other components that use CPUs, such as raylet (which can use multiple threads that can use all cpus of your machine).

`num_cpus` are more for bookeeping (the actual resource isolation is the job for the container layer).

---

<div class="post-metadata">

**Author:** ![javigm98](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/javigm98/32/704_2.png) [@javigm98](https://discuss.ray.io/u/javigm98)\
**Post date:** [April 22, 2021, 4:41pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/12 "2021-04-22T16:41:56Z")

</div>

Hi @sangcho, I was using RLLib but I think I discovered the problem. I was using Ray 1.1.0 and when I updated the version to 1.2.0 I was able to see executions in only 3 CPUs when initializing with ray(num\_cpus=3). So, only to confirm, has this been a feature changed form Ray 1.1.0 to 1.2.0?

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [April 22, 2021, 4:42pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/13 "2021-04-22T16:42:36Z")

</div>

Hmm I am confused what you mean by “executions in CPUs” now. You mean the number of processes are corresponding to the number of cpus? Or actual usage of CPUs?

---

<div class="post-metadata">

**Author:** ![javigm98](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/javigm98/32/704_2.png) [@javigm98](https://discuss.ray.io/u/javigm98)\
**Post date:** [April 22, 2021, 4:45pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/14 "2021-04-22T16:45:59Z")

</div>

Well in fact I was interested in both things. My final objective was to reduce in an effcient way CPU usage and distribute the workload among the lowest numbre of CPUs posiible. But what I say I have seen when updating to 1.2.0 is that only 3 cpus are in use

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [April 22, 2021, 4:58pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/15 "2021-04-22T16:58:05Z")

</div>

> [@javigm98](#):
>
> only 3 cpus are in use

I think I’d like to understand what you mean by this. What do you mean by only 3 cpus are in use here? How did you verify this?

---

<div class="post-metadata">

**Author:** ![javigm98](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/javigm98/32/704_2.png) [@javigm98](https://discuss.ray.io/u/javigm98)\
**Post date:** [April 22, 2021, 9:08pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/16 "2021-04-22T21:08:43Z")

</div>

Hi @sangcho and sorry because my explanations were so confusing. I’m initializing Ray with `ray.init(num_cpus=3)` and I’m using it a to train a RLlib model PPO model that consists of one driver (which is expected to use a CPU) and two workers (which are expected to use one CPU each), and they all share a GPU. So, when I execute that what I expected to see when monitirizing CPU usage with htop, for, example, was that only 3 CPUs have processes running on them. But the result that I obtain is that all CPUs are in use, as you can see in the image:

 ![Captura](https://us1.discourse-cdn.com/flex020/uploads/ray/original/1X/ab14770cb74e197bc70f81e041378f17a549ee1d.png)

So I’d like to know if this is the normal Ray behaviour, and in that case why does it make sense to restric the numbre of resources when initializing Ray. In addition, I’d like to know how Ray plans tasks to be executed (threads and processes), I mean, which volume of resources it considers that are available to plan. So for example, will it behave in the same way if for the same execution I would have initialized Ray without specifiying resource constraints?

And even more, is there any way to say Ray to execute only in a part of the resources?, I mean, for example I want Ray to only use CPUs from 0 to 2 and let the other ones completly free for other tasks and programs. Is it appropiate to do so using the `os.sched_setaffinity()` function?

I hope the explanations are clearer now, and once again sorry for the confussion. Thank you so much in advance!

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [April 22, 2021, 10:25pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/17 "2021-04-22T22:25:53Z")

</div>

Looking at your screen shot, I am seeing at least 6 instances of your training script running. Is it possible that you have past training runs that did not shut down all running at the same time?

I am looking at the “python training\_scripts/train\_model\_ppo…” processes.

---

<div class="post-metadata">

**Author:** ![javigm98](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/javigm98/32/704_2.png) [@javigm98](https://discuss.ray.io/u/javigm98)\
**Post date:** [April 22, 2021, 10:38pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/18 "2021-04-22T22:38:32Z")

</div>

Hi @mannyv. I don’t think so… Every time I initialize Ray I ensure to run a ray.shutdown() before… And the script is only being executed once…

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [April 22, 2021, 10:48pm UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/19 "2021-04-22T22:48:59Z")

</div>

They are definitely running in that screen shot you posted.

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [April 23, 2021, 12:54am UTC](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772/20 "2021-04-23T00:54:05Z")

</div>

It is normal it uses more than 3 cpus because num\_cpus=3 doesn’t mean it will only use 3 cpus actually. They are more of a “scheduling hint” rarther than the real resource isolation. So it can technically use all cpus on your machine (little by little).

[Next page](https://discuss.ray.io/t/question-about-resource-management-in-ray/1772.md?page=2)
