# Ray cluster deadlocked after drive full

**URL:** <https://discuss.ray.io/t/ray-cluster-deadlocked-after-drive-full/23349>\
**Category:** Ray Clusters\
**Created:** [December 2, 2025, 1:46pm UTC](https://discuss.ray.io/t/ray-cluster-deadlocked-after-drive-full/23349 "2025-12-02T13:46:35Z")\
**Posts on this page:** 1\
**Showing post:** 13

<div class="post-metadata">

**Author:** ![RunLLM](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/runllm/32/8249_2.png) [@RunLLM](https://discuss.ray.io/u/RunLLM)\
**Post date:** [December 2, 2025, 3:03pm UTC](https://discuss.ray.io/t/ray-cluster-deadlocked-after-drive-full/23349/13 "2025-12-02T15:03:36Z")

</div>

You’re correct—simply adding more storage may only delay the problem if the underlying causes (object leaks, unbounded spilling, or lack of cleanup) aren’t addressed. The best practice is to size storage so that the working set and expected spill volume fit comfortably, but also to monitor and control memory usage, spill rates, and reference lifetimes ([ba752da8](https://github.com/ray-project/ray/blob/releases/2.34.0/doc/source/ray-core/scheduling/memory-management.rst), [7094df69](https://discuss.ray.io/t/best-practice-for-processing-large-amounts-of-data/5672)).

For robust production, combine increased storage with:

- Monitoring object store and disk usage.
- Limiting the number of concurrent jobs/tasks.
- Ensuring all references are released after job completion.
- Periodic cleanup or node restarts to clear leaks.

Would you like more detail on how to estimate storage needs or set up monitoring and cleanup?

Sources:

- [Best practice for processing large amounts of data](https://discuss.ray.io/t/best-practice-for-processing-large-amounts-of-data/5672)

Hint: Mention @RunLLM in the post for followups.

---

_[View the full topic](https://discuss.ray.io/t/ray-cluster-deadlocked-after-drive-full/23349)._
