# Under-utilization of gpus at end of experiment

**URL:** <https://discuss.ray.io/t/under-utilization-of-gpus-at-end-of-experiment/8728>\
**Category:** Ray Tune\
**Created:** [December 20, 2022, 6:24pm UTC](https://discuss.ray.io/t/under-utilization-of-gpus-at-end-of-experiment/8728 "2022-12-20T18:24:01Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mike1](https://avatars.discourse-cdn.com/v4/letter/m/e19adc/32.png) [@Mike1](https://discuss.ray.io/u/Mike1)\
**Post date:** [December 20, 2022, 6:24pm UTC](https://discuss.ray.io/t/under-utilization-of-gpus-at-end-of-experiment/8728/1 "2022-12-20T18:24:01Z")

</div>

I am using Tune for hyperparameter optimization for Pytorch Lightning (not using Ray Lightning). I have 8 GPUs and I am assigning 1 of them to each trial. As a result, at the end of the experiment, when I have less than 8 trials left, some GPUs remain idle. I would like the idle resources reallocated to the remaining trials. Is there a way to automatically make sure all resources are used at all times? It’s worth noting that my lightning trainer has accelerator=“auto”. Thank you!

---

<div class="post-metadata">

**Author:** ![amogkam](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amogkam/32/17_2.png) [@amogkam](https://discuss.ray.io/u/amogkam)\
**Post date:** [December 20, 2022, 7:07pm UTC](https://discuss.ray.io/t/under-utilization-of-gpus-at-end-of-experiment/8728/2 "2022-12-20T19:07:54Z")

</div>

Hey @Mike1, yes this is possible with the `ResourceChangingScheduler` API: [Trial Schedulers (tune.schedulers) — Ray 2.2.0](https://docs.ray.io/en/latest/tune/api_docs/schedulers.html#resourcechangingscheduler). This allows resources to be re-allocated during the run of the experiment and solves the exact issue that you are describing!

There is an example using this with XGBoost here: [XGBoost Dynamic Resources Example — Ray 2.2.0](https://docs.ray.io/en/latest/tune/examples/includes/xgboost_dynamic_resources_example.html), but the same idea can be applied to PyTorch Lightning.

This is still experimental though, so please let us know if you run into any issues!

---

<div class="post-metadata">

**Author:** ![Mike1](https://avatars.discourse-cdn.com/v4/letter/m/e19adc/32.png) [@Mike1](https://discuss.ray.io/u/Mike1)\
**Post date:** [December 21, 2022, 5:41pm UTC](https://discuss.ray.io/t/under-utilization-of-gpus-at-end-of-experiment/8728/3 "2022-12-21T17:41:49Z")

</div>

Thank you for the answer. I tried it and unfortunately, this small modification tends to produce many errors (about 70% of the trials end in error, and the types of errors are inconsistent). I then tried a different solution: run the trials analogically- giving all the resources to one trial, one at a time. However, I found out that assigning more than 1 gpu to a trial (via resources\_per\_trial in tune.with\_resources) makes the trial not run at all: Tune reports the trial as running, but nothing happens- not a single iteration is run, and tune is stuck in an infinite loop. Any advice about how to make that work? Thank you!

---

<div class="post-metadata">

**Author:** ![amogkam](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amogkam/32/17_2.png) [@amogkam](https://discuss.ray.io/u/amogkam)\
**Post date:** [December 21, 2022, 11:46pm UTC](https://discuss.ray.io/t/under-utilization-of-gpus-at-end-of-experiment/8728/4 "2022-12-21T23:46:05Z")

</div>

Hey @Mike1, would you be able to share what your code looks like, as well as what is being printed out to `stdout` and the errors that you are seeing?
