# Multi GPU Usage on Multi VM|Ray cluster on multi VM instances

**URL:** <https://discuss.ray.io/t/multi-gpu-usage-on-multi-vm-ray-cluster-on-multi-vm-instances/11333>\
**Category:** Ray Clusters\
**Created:** [July 10, 2023, 2:45pm UTC](https://discuss.ray.io/t/multi-gpu-usage-on-multi-vm-ray-cluster-on-multi-vm-instances/11333 "2023-07-10T14:45:23Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Shobhit\_Agarwal](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/shobhit_agarwal/32/4732_2.png) [@Shobhit\_Agarwal](https://discuss.ray.io/u/Shobhit_Agarwal)\
**Post date:** [July 10, 2023, 2:45pm UTC](https://discuss.ray.io/t/multi-gpu-usage-on-multi-vm-ray-cluster-on-multi-vm-instances/11333/1 "2023-07-10T14:45:23Z")

</div>

Background:  
I want to try the LLM model, for example, flan-ul2 onto the two VM A10 GPUs provided by AWS, Each VM has 4 GPUs, so in my ray cluster I would have in total of 8 GPUs. Now, I want to create a ray cluster, which I already did by running the following commands:

on head node:  
ray start --head

on worker node:  
ray start --address=“:”

But now in my code where I created a Python class which I want to deploy, I want to share 6 GPUs for the task among the worker and head node, how can I proceed?

any leads can be beneficial.

- High: It blocks me from completing my task.

---

<div class="post-metadata">

**Author:** ![Jules\_Damji](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jules_damji/32/4058_2.png) [@Jules\_Damji](https://discuss.ray.io/u/Jules_Damji)\
**Post date:** [July 10, 2023, 8:56pm UTC](https://discuss.ray.io/t/multi-gpu-usage-on-multi-vm-ray-cluster-on-multi-vm-instances/11333/2 "2023-07-10T20:56:26Z")

</div>

@Shobhit_Agarwal Here is a goodhttps://docs.ray.io/en/latest/ray-air/examples/gptj\_serving.html example how you can use Ray Serve and Ray to serve an LLM model. For this model,  
we use 16GB GPUs. We allocate one GPU per replica, so 6 replicas will have 6 GPUs.

---

<div class="post-metadata">

**Author:** ![Shobhit\_Agarwal](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/shobhit_agarwal/32/4732_2.png) [@Shobhit\_Agarwal](https://discuss.ray.io/u/Shobhit_Agarwal)\
**Post date:** [July 11, 2023, 3:00am UTC](https://discuss.ray.io/t/multi-gpu-usage-on-multi-vm-ray-cluster-on-multi-vm-instances/11333/3 "2023-07-11T03:00:55Z")

</div>

@Jules_Damji really appreciate the quick response. But the thing is if i set num\_replicas=6 and num\_gpus=1, that means i am making 6 copies of it and each copy is utilising 1 GPU, please correct me if i am wrong.

The problem is I can’t be using single GPU for the LLM, I need at least 5/6 GPUs to serve the flan ul2 model since it is huge. So after creating the cluster, in my deployment class, I am setting num\_gpus=6, num\_replicas=1, but I am getting an error saying that, no resource can accommodate num\_gpus=6, any leads can be helpful.

---

<div class="post-metadata">

**Author:** ![Shobhit\_Agarwal](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/shobhit_agarwal/32/4732_2.png) [@Shobhit\_Agarwal](https://discuss.ray.io/u/Shobhit_Agarwal)\
**Post date:** [July 14, 2023, 5:37am UTC](https://discuss.ray.io/t/multi-gpu-usage-on-multi-vm-ray-cluster-on-multi-vm-instances/11333/4 "2023-07-14T05:37:17Z")

</div>

any leads would be really helpful.

---

<div class="post-metadata">

**Author:** ![Shobhit\_Agarwal](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/shobhit_agarwal/32/4732_2.png) [@Shobhit\_Agarwal](https://discuss.ray.io/u/Shobhit_Agarwal)\
**Post date:** [August 2, 2023, 10:46am UTC](https://discuss.ray.io/t/multi-gpu-usage-on-multi-vm-ray-cluster-on-multi-vm-instances/11333/5 "2023-08-02T10:46:15Z")

</div>

@Jules_Damji, I have a scenario, where I create a ray cluster with 2 VMs, each having 4 GPUs, how can I distribute my ray serve that utilizes 4 GPUs from the first instance and 1 GPU from another instance? is there a workaround for this?

I created a cluster  
ray start --head on head node,  
and ray start --address= on worker node

and assigned num\_gpus=5 in @serve.deployment class, but still I am getting the below error message:  
no available node types can fulfill resource request {‘gpu’: 5.0},

even when I see resources available: {“gpu”: 8.0}

I hope there should be a workaround for this.

---

<div class="post-metadata">

**Author:** ![knowledgeseeker](https://avatars.discourse-cdn.com/v4/letter/k/eb9ed0/32.png) [@knowledgeseeker](https://discuss.ray.io/u/knowledgeseeker)\
**Post date:** [January 17, 2025, 8:57pm UTC](https://discuss.ray.io/t/multi-gpu-usage-on-multi-vm-ray-cluster-on-multi-vm-instances/11333/6 "2025-01-17T20:57:23Z")

</div>

hello @Shobhit_Agarwal did you get any solution around it ?
