# Ray Blocking Spark Jobs

**URL:** <https://discuss.ray.io/t/ray-blocking-spark-jobs/21941>\
**Category:** Ray Clusters\
**Created:** [March 2, 2025, 10:52am UTC](https://discuss.ray.io/t/ray-blocking-spark-jobs/21941 "2025-03-02T10:52:15Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![naushad](https://avatars.discourse-cdn.com/v4/letter/n/eada6e/32.png) [@naushad](https://discuss.ray.io/u/naushad)\
**Post date:** [March 2, 2025, 10:52am UTC](https://discuss.ray.io/t/ray-blocking-spark-jobs/21941/1 "2025-03-02T10:52:15Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- Medium: It contributes to significant difficulty to complete my task, but I can work around it.

I have setup ray cluster on Databricks runtime 15.4 LTS without specifying min\_worker\_nodes and max\_worker\_nodes to dynamically allocate resources. But it’s blocking other spark tasks/jobs which wait indefinitely to start.

I was not facing this issue while running it on 12.2 LTS.

ray version: 2.33.0

---

<div class="post-metadata">

**Author:** ![christina](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/christina/32/7542_2.png) [@christina](https://discuss.ray.io/u/christina)\
**Post date:** [March 3, 2025, 8:26pm UTC](https://discuss.ray.io/t/ray-blocking-spark-jobs/21941/2 "2025-03-03T20:26:21Z")

</div>

Hi! Are there any error messages happening while you’re waiting for the Spark jobs to start? I’m guessing this might be due to how Ray is using resources, so that means there might not be any resources for Spark to start if Ray is using them all. Have you tried manually setting `min_worker_nodes` and `max_worker_nodes` to different values to see if that unblocks it?

I’m not too familiar with 15.4 LTS, did they change resource allocation or anything when upgrading from 12.2?

---

<div class="post-metadata">

**Author:** ![naushad](https://avatars.discourse-cdn.com/v4/letter/n/eada6e/32.png) [@naushad](https://discuss.ray.io/u/naushad)\
**Post date:** [March 4, 2025, 7:20am UTC](https://discuss.ray.io/t/ray-blocking-spark-jobs/21941/3 "2025-03-04T07:20:21Z")

</div>

There is no error. Spark jobs are stuck in waiting state indefinitely. I have tried setting min\_worker\_nodes and max\_worker\_nodes which works fine as ray scales up/down depending upon load leaving resources for spark jobs to execute.

When I don’t specify number of worker nodes, ray is starting with all available worker nodes available in databricks spark cluster to meet minimum number of worker nodes

---

<div class="post-metadata">

**Author:** ![lsjhome](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lsjhome/32/7765_2.png) [@lsjhome](https://discuss.ray.io/u/lsjhome)\
**Post date:** [March 11, 2025, 1:51am UTC](https://discuss.ray.io/t/ray-blocking-spark-jobs/21941/4 "2025-03-11T01:51:43Z")

</div>

If you allocate all available workers, the Databricks Spark driver appears to have no resources left, causing it to get stuck.

To avoid this, I use the following approach:

ray\_worker\_count = max(1, int(total\_workers \* 0.75))

For instance, when my driver is assigned 3 workers, I set `max_worker_nodes` to 2, which allows it to function properly.
