# Resource allocation for Ray Cluster running on Kubernetes

**URL:** <https://discuss.ray.io/t/resource-allocation-for-ray-cluster-running-on-kubernetes/5408>\
**Category:** Ray Clusters\
**Created:** [March 15, 2022, 3:49pm UTC](https://discuss.ray.io/t/resource-allocation-for-ray-cluster-running-on-kubernetes/5408 "2022-03-15T15:49:09Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Manoj\_Kumar\_Dobbali](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/manoj_kumar_dobbali/32/2172_2.png) [@Manoj\_Kumar\_Dobbali](https://discuss.ray.io/u/Manoj_Kumar_Dobbali)\
**Post date:** [March 15, 2022, 3:49pm UTC](https://discuss.ray.io/t/resource-allocation-for-ray-cluster-running-on-kubernetes/5408/1 "2022-03-15T15:49:09Z")

</div>

Hi,

I am using Pytorch Lightning for training and Ray for Hyper parameter tuning (not using ray\_lightning). I have a Kubernetes operator and head(m6a.2xlarge) that can spin up max 5 GPU workers(p2.8xlarge with 8 GPUs, 32 vCPUs).

Documentation says “if connected to existing cluster, you don’t specify resources.”

I don’t need to specify resources on Pytorch Lightning Trainer object too? Also, no need to set up `resource_per_trial` as well?

Currently if I do gpu = 1 in plt.Trainer(gpus=1…) and resource\_per\_trial = {“gpus”: 1} and Run 16 experiments, I am seeing 8 experiments on each worker. What I am expecting is, Each experiment running on its own worker and that each experiment using all GPUs to run faster.

What is the best way to allocate resource? I am running Multi Layer Perceptron model using Pytorch Lightning

---

<div class="post-metadata">

**Author:** ![Manoj\_Kumar\_Dobbali](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/manoj_kumar_dobbali/32/2172_2.png) [@Manoj\_Kumar\_Dobbali](https://discuss.ray.io/u/Manoj_Kumar_Dobbali)\
**Post date:** [March 15, 2022, 4:51pm UTC](https://discuss.ray.io/t/resource-allocation-for-ray-cluster-running-on-kubernetes/5408/2 "2022-03-15T16:51:20Z")

</div>

Also, I started using Ray Lightning to see if that helps in resource allocation efficiently

The documentation about setting up `gpus` on trainer is unclear

Doctoring says to setup `num_gpus` : [ray\_lightning/ray\_ddp.py at 3adb809aee8d1c6154e044902d359a456f1859ff · ray-project/ray\_lightning · GitHub](https://github.com/ray-project/ray_lightning/blob/3adb809aee8d1c6154e044902d359a456f1859ff/ray_lightning/ray_ddp.py#L88)

While readme says  
Don’t set `gpus` in the `Trainer`.  
The actual number of GPUs is determined by `num_workers`.
