# Optimizing Ray Tune for Large-Scale Hyperparameter Search with High Resource Utilization

**URL:** <https://discuss.ray.io/t/optimizing-ray-tune-for-large-scale-hyperparameter-search-with-high-resource-utilization/21136>\
**Category:** Uncategorized\
**Created:** [December 18, 2024, 11:32am UTC](https://discuss.ray.io/t/optimizing-ray-tune-for-large-scale-hyperparameter-search-with-high-resource-utilization/21136 "2024-12-18T11:32:49Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![romie090](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/romie090/32/7424_2.png) [@romie090](https://discuss.ray.io/u/romie090)\
**Post date:** [December 18, 2024, 11:32am UTC](https://discuss.ray.io/t/optimizing-ray-tune-for-large-scale-hyperparameter-search-with-high-resource-utilization/21136/1 "2024-12-18T11:32:49Z")

</div>

Hi everyone,

I’ve been experimenting with Ray Tune to perform hyperparameter optimization for a deep learning model. While it works well for small-scale tasks, I’m running into some issues when scaling up to larger workloads that involve many trials and significant resource consumption.

Here’s my current setup:

- **Framework** : PyTorch
- **Cluster** : 8-node setup, each with 16 cores, 64GB RAM, and 1 GPU per node
- **Search Algorithm** : Optuna + ASHA scheduler
- **Task** : Training a model for ~50,000 trials with varying learning rates, batch sizes, and hidden layer dimensions

### Challenges I’m Facing:

1. **Resource Utilization** : While I expect high resource utilization, I see CPU/GPU usage fluctuating, and sometimes nodes sit idle. I’ve ensured my cluster configuration in Ray is correct, but there still seems to be inefficiencies.
2. **Trial Scheduling Overhead** : With a large number of concurrent trials, the scheduling overhead increases significantly. Are there specific configurations or scheduler parameters I should tweak to minimize this?
3. **Checkpointing** : I’m saving checkpoints for each trial, but this becomes resource-intensive and sometimes causes delays. Any suggestions on optimizing checkpoint frequency or handling storage efficiently?

### Goals:

I want to ensure:

- Maximum resource utilization across nodes (CPU/GPU)
- Efficient scheduling to reduce trial overhead
- Scalable checkpointing for large-scale experiments

> [@TuneReportCallback is unable to read PyTorch Lightning metrics during Tuner.fit(...)](https://discuss.ray.io/t/tunereportcallback-is-unable-to-read-pytorch-lightning-metrics-during-tuner-fit/9063):
>
> How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. I’m running Ray Tune 2.2.0 with a PyTorch Lightning module and found that tune.report(…) inside TuneReportCallback is unable to relay metrics back to the Ray session. Diving deeper, I found that the Ray session is disabled during the training/validation steps of the PyTorch Lightning module. This is the error I have been receiving: Session not detected. You should not be calling report …

Has anyone faced similar issues or found effective strategies for optimizing Ray Tune under high workloads? I’d love to hear your experiences or advice on best practices, configuration tweaks, or alternative [alteryx certification](https://www.igmguru.com/data-science-bi/alteryx-training) tools that complement Ray Tune.

Thanks in advance for your insights!

Regards  
Romie
