# Cpu allocation confusion

**URL:** <https://discuss.ray.io/t/cpu-allocation-confusion/9621>\
**Category:** Uncategorized\
**Created:** [March 4, 2023, 2:08am UTC](https://discuss.ray.io/t/cpu-allocation-confusion/9621 "2023-03-04T02:08:08Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![starkj](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/starkj/32/3266_2.png) [@starkj](https://discuss.ray.io/u/starkj)\
**Post date:** [March 4, 2023, 2:08am UTC](https://discuss.ray.io/t/cpu-allocation-confusion/9621/1 "2023-03-04T02:08:08Z")

</div>

Hi, I’m training with Tune (Ray 2.3.0) on a single machine that has 16 cpus and 1 gpu. Just getting a feel for how resource management works, I am explicitly setting  
num\_gpus = 0  
num\_gpus\_per\_worker = 0  
num\_cpus\_for\_local\_worker = 1  
num\_cpus\_per\_worker = 1  
num\_rollout\_workers = 1  
rollout\_fragment\_length = 200  
train\_batch\_size = 200 #must be = rollout\_fragment\_length \* num\_rollout\_workers \* num\_envs\_per\_worker  
sgc\_minibatch\_size = 32

Then I start the tune job with PPO and TuneConfig(num\_samples = 24) and the PBT scheduler. What I see is that Ray aggressively spins up 8 workers immediately. Why isn’t it limited to 1, as specified? If it is going to ignore my limit request, why wouldn’t it try to use all 16 cpus (or 15 workers + driver)?

The confusion continues: When I change num\_cpus\_per\_worker = 0 it runs a whopping 12 workers! Still not using all the resources, but no obvious explanation for the choice (I’m guessing that 0 tells Ray to do whatever it thinks is best).

And more confusion: When I change num\_cpus\_per\_worker = 1 again, then  
num\_rollout\_workers = 13  
train\_batch\_size = 2600  
It only fires up a single worker. I am now totally baffled. What is the rubric behind this worker/cpu allocation?

Thanks!

---

<div class="post-metadata">

**Author:** ![justinvyu](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/justinvyu/32/4309_2.png) [@justinvyu](https://discuss.ray.io/u/justinvyu)\
**Post date:** [March 6, 2023, 11:39pm UTC](https://discuss.ray.io/t/cpu-allocation-confusion/9621/2 "2023-03-06T23:39:04Z")

</div>

Hi @starkj,

Your current config with `num_cpus_per_worker=1` and `num_rollout_workers=1` will allocate 2 CPUs per RLlib trial: one actor with 1 CPU for training, and another actor for doing env rollouts with 1 CPU. So, Tune will allocate 8 \* 2 = 16 CPUs worth of remote actors.

The `num_rollout_workers` is an RLlib config that defines how many rollout workers should be spawned _per-trial_, so Tune’s scheduling of running trials concurrently is not related to that. This section of the RLlib docs may be useful to clarify these things: [Getting Started with RLlib — Ray 2.3.0](https://docs.ray.io/en/latest/rllib/rllib-training.html#specifying-resources)

To limit the concurrency on the Tune side, you can set `Tuner(tune_config=tune.TuneConfig(max_concurrent_trials=1))`. See [ray.tune.TuneConfig — Ray 2.3.0](https://docs.ray.io/en/latest/tune/api/doc/ray.tune.TuneConfig.html) for more info.

Let me know if that answers your questions!

---

<div class="post-metadata">

**Author:** ![starkj](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/starkj/32/3266_2.png) [@starkj](https://discuss.ray.io/u/starkj)\
**Post date:** [March 7, 2023, 2:53am UTC](https://discuss.ray.io/t/cpu-allocation-confusion/9621/3 "2023-03-07T02:53:14Z")

</div>

Hi @justinvyu , thanks for the reply.

It sounds like every trial is associated with 1 rollout worker (using num\_cpus\_per\_worker) plus an additional CPU, and then Ray automatically determines how many trials can be run in parallel based on the available cpus. ISo if I have num\_cpus\_per\_worker = 2 then each trial will use 3 cpus? If I set num\_cpus\_per\_worker = 0 does that force the trial to do everything on a single cpu?

was looking at it the other way round, where a specific number of workers is requested (num\_rollout\_workers), then as each one becomes available it is assigned to work on a new trial. After seeing your answer and re-reading the manual page, I still feel that the manual is delivering this perspective. A little more description & examples there would be most helpful!

A new confusion: I just tried setting num\_gpus = 1 (it was 0), and left everything else the same as at the top of my post. Now the number of concurrent trials drops from 8 to 1. As I understand, the num\_gpus param refers to what is available for the driver process (doing the learning algo), so I would think it doesn’t affect the rollout workers at all. Why did it do this?

Thanks again!

---

<div class="post-metadata">

**Author:** ![justinvyu](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/justinvyu/32/4309_2.png) [@justinvyu](https://discuss.ray.io/u/justinvyu)\
**Post date:** [March 7, 2023, 5:42pm UTC](https://discuss.ray.io/t/cpu-allocation-confusion/9621/4 "2023-03-07T17:42:01Z")

</div>

The resource allocation for actors is mainly for book-keeping purposes (for Ray to figure out how many things to schedule concurrently). Ray does not automatically provide resource isolation, so you should limit the number of CPUs used by your application logic (ex: setting the number of jobs in sklearn).

For the GPU question, your cluster only has 1 GPU, so you can only run one trial at a time if you have a resource request of X CPUs and 1 GPU. You can set fractional GPUs to get around this (in which case you should also make sure GPU memory usage is limited per actor).

See [GPU Support — Ray 2.3.0](https://docs.ray.io/en/latest/ray-core/tasks/using-ray-with-gpus.html#fractional-gpus)
