# Automaticly choose the most free GPU

**URL:** https://discuss.ray.io/t/automaticly-choose-the-most-free-gpu/11780
**Category:** Ray Core
**Created:** [August 14, 2023, 11:25am UTC](https://discuss.ray.io/t/automaticly-choose-the-most-free-gpu/11780 "2023-08-14T11:25:44Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![Ilnur786](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ilnur786/32/4919_2.png) [@Ilnur786](https://discuss.ray.io/u/Ilnur786)
#### Post date: [August 14, 2023, 11:25am UTC](https://discuss.ray.io/t/automaticly-choose-the-most-free-gpu/11780/1 "2023-08-14T11:25:44Z")

</div>

I have 2 gpu on the machine and how to choose the most free GPU for each run? I wrapped predict func with @ray.remote(num\_gpus=1, num\_cpus=8) decorator, wrote func, which shows the most free GPU and set it through os.environ[‘CUDA\_VISIBLE\_DEVICES’] = str(gpu\_id). When the most free GPU is changed and a new instance of model loading on another GPU, ray releases model instances in the previous GPU. How to solve this problem?

Update: as a default, ray chooses GPU with 0 id, even if was sat CUDA\_VISIBLE\_DEVICES=0,1 and @ray.remote(num\_gpus=2, num\_cpus=8, max\_calls=1)

---

<div class="post-metadata">

### Author: ![XIE](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/xie/32/4315_2.png) [@XIE](https://discuss.ray.io/u/XIE)
#### Post date: [August 15, 2023, 5:33am UTC](https://discuss.ray.io/t/automaticly-choose-the-most-free-gpu/11780/2 "2023-08-15T05:33:40Z")

</div>

cc: @yic could you take a quick look?

---

<div class="post-metadata">

### Author: ![Ilnur786](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ilnur786/32/4919_2.png) [@Ilnur786](https://discuss.ray.io/u/Ilnur786)
#### Post date: [August 15, 2023, 4:18pm UTC](https://discuss.ray.io/t/automaticly-choose-the-most-free-gpu/11780/3 "2023-08-15T16:18:27Z")

</div>

I should say that I’m doing this on aws ec2 machine and the model is one of the hugging face transformers

---

<div class="post-metadata">

### Author: ![yic](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/yic/32/437_2.png) [@yic](https://discuss.ray.io/u/yic)
#### Post date: [August 17, 2023, 10:28pm UTC](https://discuss.ray.io/t/automaticly-choose-the-most-free-gpu/11780/4 "2023-08-17T22:28:14Z")

</div>

@Ilnur786 could you give me a script to show what do you mean by most free GPU? IIUC, Ray only treat GPU as logic resource and doesn’t check ‘most free’ GPU.

I’ll be nice if you can have a script showing what’s going wrong and what’s expected.

---

<div class="post-metadata">

### Author: ![Ilnur786](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ilnur786/32/4919_2.png) [@Ilnur786](https://discuss.ray.io/u/Ilnur786)
#### Post date: [August 21, 2023, 4:08pm UTC](https://discuss.ray.io/t/automaticly-choose-the-most-free-gpu/11780/5 "2023-08-21T16:08:50Z")

</div>

Hello. I’m getting info about gpus free space by this func, which return List[int]:

```auto
def get_gpu_memory():
    command = "nvidia-smi --query-gpu=memory.free --format=csv"
    memory_free_info = sp.check_output(command.split()).decode('ascii').split('\n')[:-1][1:]
    memory_free_values = [int(x.split()[0]) for i, x in enumerate(memory_free_info)]
    return memory_free_values

```

then, I set the GPU index in os.environ[‘CUDA\_VISIBLE\_DEVICES’] = gpu\_id. It worked, but when some tasks are already was running on gpu:0 and the next job should be run on gpu:1 (because it was the most free GPU at this moment), ray released resources from gpu:0 which lead killing the tasks on it.  
I chose ray, because struggled from that I wasn’t able to release resources after huggingface transformer, but, unfortunately, ray doesn’t give the opportunity to notice gpu index exactly.  
I solved the problem with can’t releasing resources after task finishing by running the task in another process by multiprocessing and fortunately, hugging face model gives opportunity to choose gpu index  
P.S. I tried to use ray and give gpu index to model, but this schema wasn’t work

---

<div class="post-metadata">

### Author: ![Ilnur786](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ilnur786/32/4919_2.png) [@Ilnur786](https://discuss.ray.io/u/Ilnur786)
#### Post date: [August 29, 2023, 10:47am UTC](https://discuss.ray.io/t/automaticly-choose-the-most-free-gpu/11780/6 "2023-08-29T10:47:13Z")

</div>

Some updates for future visitors: I changed the model to a pure torch one and met the same issue. It can be because of either the task management system (dramatiq in my case) or amazon ec2 machine. In most other cases, I think there should be an opportunity to release resources with standard methods: move the model and other tensors to the CPU, delete variables, and clean torch cache. If not, use ray or run the model calculation in another process.
