# Optimizing Ray Tune for Large-Scale Hyperparameter Search with High Resource Utilization

**URL:** <https://discuss.ray.io/t/optimizing-ray-tune-for-large-scale-hyperparameter-search-with-high-resource-utilization/21137>\
**Category:** Uncategorized\
**Created:** [December 18, 2024, 11:36am UTC](https://discuss.ray.io/t/optimizing-ray-tune-for-large-scale-hyperparameter-search-with-high-resource-utilization/21137 "2024-12-18T11:36:07Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![romie090](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/romie090/32/7424_2.png) [@romie090](https://discuss.ray.io/u/romie090)\
**Post date:** [December 18, 2024, 11:36am UTC](https://discuss.ray.io/t/optimizing-ray-tune-for-large-scale-hyperparameter-search-with-high-resource-utilization/21137/1 "2024-12-18T11:36:08Z")

</div>

Hi everyone,

I’ve been experimenting with Ray Tune to perform hyperparameter optimization for a deep learning model. While it works well for small-scale tasks, I’m running into some issues when scaling up to larger workloads that involve many trials and significant resource consumption.

Here’s my current setup:

- **Framework** : PyTorch
- **Cluster** : 8-node setup, each with 16 cores, 64GB RAM, and 1 GPU per node
- **Search Algorithm** : Optuna + ASHA scheduler
- **Task** : Training a model for ~50,000 trials with varying learning rates, batch sizes, and hidden layer dimensions

### Challenges I’m Facing:

1. **Resource Utilization** : While I expect high resource utilization, I see CPU/GPU usage fluctuating, and sometimes nodes sit idle. I’ve ensured my cluster configuration in Ray is correct, but there still seems to be inefficiencies.
2. **Trial Scheduling Overhead** : With a large number of concurrent trials, the scheduling overhead increases significantly. Are there specific configurations or scheduler parameters I should tweak to minimize this?
3. **Checkpointing** : I’m saving checkpoints for each trial, but this becomes resource-intensive and sometimes causes delays. Any suggestions on optimizing checkpoint frequency or handling storage efficiently?

### Goals:

I want to ensure:

- Maximum resource utilization across nodes (CPU/GPU)
- Efficient scheduling to reduce trial overhead
- Scalable checkpointing for large-scale experiments

> [@Specify trial resources when using search algorithm to tune hyper-parameters](https://discuss.ray.io/t/specify-trial-resources-when-using-search-algorithm-to-tune-hyper-parameters/7468):
>
> How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. Trying to use ray.tune.Tuner, ray.tune.search.optuna.OptunaSearch, ray.tune.schedulers.ASHAScheduler using Ray 2 to find the best hyper-parameters for a PPO policy that maximizes mean reward while also performing early termination of bad trials. Code snippet below highlights the current process, but that generates the following error: ray.tune.error.TuneError: No trial resources are avail…

Has anyone faced similar issues or found effective strategies for optimizing Ray Tune under high workloads? I’d love to hear your experiences or advice on best practices, configuration tweaks, or alternative [alteryx](https://www.igmguru.com/data-science-bi/alteryx-training) tools that complement Ray Tune.

Thanks in advance for your insights!
