# Data Parallelism with Ray Tune

**URL:** <https://discuss.ray.io/t/data-parallelism-with-ray-tune/12636>\
**Category:** Uncategorized\
**Created:** [October 28, 2023, 5:07pm UTC](https://discuss.ray.io/t/data-parallelism-with-ray-tune/12636 "2023-10-28T17:07:17Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![amztc34283](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amztc34283/32/5250_2.png) [@amztc34283](https://discuss.ray.io/u/amztc34283)\
**Post date:** [October 28, 2023, 5:07pm UTC](https://discuss.ray.io/t/data-parallelism-with-ray-tune/12636/1 "2023-10-28T17:07:17Z")

</div>

I am running Ray Tune with gpu=1 per trial, but each trial is only using one GPU although I specified net = torch.nn.DataParallel(net) as opposed to splitting the data across all 8 GPUs, do I need to set gpu=8 in the resource configuration to allow data parallelism across GPUs?

This is the resource config:  
tuner = tune.Tuner(  
tune.with\_resources(  
…,  
resources={“cpu”: 2, “gpu”: 1} # per trial by default  
)  
)

I think my goal is to split the computation of each trial across the GPUs (using DataParallel) while running multiple trials with multi-processing in parallel, but I have not figured out the best way to do it. For example, if I run my tuner with resources={“cpu”: 2, “gpu”: 1}, each trial will run on its own GPU but this would not benefit from the data parallelism you could get from DataParallel. In addition, if I were to run with resources={“cpu”: 2, “gpu”: 8}, only one trial could benefit from the data parallelism.

---

<div class="post-metadata">

**Author:** ![matthewdeng](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/matthewdeng/32/1446_2.png) [@matthewdeng](https://discuss.ray.io/u/matthewdeng)\
**Post date:** [October 31, 2023, 10:06pm UTC](https://discuss.ray.io/t/data-parallelism-with-ray-tune/12636/2 "2023-10-31T22:06:00Z")

</div>

What is your desired behavior here? How many GPUs do you want each trial to use?

---

<div class="post-metadata">

**Author:** ![amztc34283](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amztc34283/32/5250_2.png) [@amztc34283](https://discuss.ray.io/u/amztc34283)\
**Post date:** [November 1, 2023, 4:48am UTC](https://discuss.ray.io/t/data-parallelism-with-ray-tune/12636/3 "2023-11-01T04:48:49Z")

</div>

There are 8 GPUs available, and I want to maximize the parallelism available without limiting each trial to the number of physical gpus assigned to each worker. As they are shared resources, I want to avoid hotspotting a single GPU.

The desired behavior is to enable each trial to provision all 8 GPUs available so that the batch can be sharded across the GPUs while running more than one trials in parallel.

Currently this is impossible because setting “gpu”: 1 causes hotspotting a single gpu while “gpu”: 8 causes abuse of resources (leaving a lot of resources unused).

---

<div class="post-metadata">

**Author:** ![matthewdeng](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/matthewdeng/32/1446_2.png) [@matthewdeng](https://discuss.ray.io/u/matthewdeng)\
**Post date:** [November 1, 2023, 9:17pm UTC](https://discuss.ray.io/t/data-parallelism-with-ray-tune/12636/4 "2023-11-01T21:17:44Z")

</div>

Can you explain more about your use case? Typically you want to use data parallelism if you are bound by the memory of a single GPU. If you have N trials, it would not be beneficial for each of them to be running in a data parallel fashion across all 8 GPUs.

---

<div class="post-metadata">

**Author:** ![amztc34283](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amztc34283/32/5250_2.png) [@amztc34283](https://discuss.ray.io/u/amztc34283)\
**Post date:** [November 2, 2023, 4:12pm UTC](https://discuss.ray.io/t/data-parallelism-with-ray-tune/12636/5 "2023-11-02T16:12:01Z")

</div>

As I am running in a shared environment where lots of people could be using the same GPUs, training the whole batch in a single GPU could potentially cause OOM error.

To avoid the OOM error, I would like to shard the batch across GPUs.

As you mentioned this is probably not a good solution performance-wise but a workaround would be nice.
