# Ray Clusters

**URL:** https://discuss.ray.io/c/ray-clusters/13.md?page=1

[Latest](https://discuss.ray.io/latest.md) · [Categories](https://discuss.ray.io/categories.md) · [Tags](https://discuss.ray.io/tags.md)

**Page:** 2

---

## [vLLM + Ray multi-node tensor-parallel deployment completely blocked by pending placement groups and raylet heartbeat failures](https://discuss.ray.io/t/vllm-ray-multi-node-tensor-parallel-deployment-completely-blocked-by-pending-placement-groups-and-raylet-heartbeat-failures/22941)

<div class="topic-metadata">

**Author:** [@NullUser](https://discuss.ray.io/u/NullUser)\
**Replies:** 0\
**Last updated:** [August 5, 2025, 5:06am UTC](https://discuss.ray.io/t/vllm-ray-multi-node-tensor-parallel-deployment-completely-blocked-by-pending-placement-groups-and-raylet-heartbeat-failures/22941 "2025-08-05T05:06:24Z")

</div>

1. Severity of the issue: High: Completely blocks me. 2. Environment: Ray version: 2.x vLLM version: 0.9.2 Python version: 3.9 OS / Container base: Linux (CentOS-based UBI8 in Kubernetes) Cloud / Infrastructure: AWS…

---

## [Global cluster resource limit or resource limit for a group of workers](https://discuss.ray.io/t/global-cluster-resource-limit-or-resource-limit-for-a-group-of-workers/22927)

<div class="topic-metadata">

**Author:** [@Liquidmasl](https://discuss.ray.io/u/Liquidmasl)\
**Replies:** 0\
**Last updated:** [August 1, 2025, 2:02pm UTC](https://discuss.ray.io/t/global-cluster-resource-limit-or-resource-limit-for-a-group-of-workers/22927 "2025-08-01T14:02:43Z")

</div>

1. Severity of the issue: (select one) Medium: Significantly affects my productivity but can find a workaround. High: Completely blocks me. 2. Environment: Ray version: 2.48.0 Python version: 3.12.11 OS: linux Oth…

---

## [A way to show GCP logs in the dashboard?](https://discuss.ray.io/t/a-way-to-show-gcp-logs-in-the-dashboard/22910)

<div class="topic-metadata">

**Author:** [@svartalf](https://discuss.ray.io/u/svartalf)\
**Replies:** 1\
**Last updated:** [July 31, 2025, 7:47pm UTC](https://discuss.ray.io/t/a-way-to-show-gcp-logs-in-the-dashboard/22910 "2025-07-31T19:47:48Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Unable to access files from disk filesystem inside methods run using ray multiprocessing](https://discuss.ray.io/t/unable-to-access-files-from-disk-filesystem-inside-methods-run-using-ray-multiprocessing/22915)

<div class="topic-metadata">

**Author:** [@Pranav\_Thampi](https://discuss.ray.io/u/Pranav_Thampi)\
**Replies:** 0\
**Last updated:** [July 31, 2025, 10:25am UTC](https://discuss.ray.io/t/unable-to-access-files-from-disk-filesystem-inside-methods-run-using-ray-multiprocessing/22915 "2025-07-31T10:25:26Z")

</div>

1. Severity of the issue: Medium: Significantly affects my productivity but can find a workaround. 2. Environment: Ray version: ray\[default,tune,client,data\]==2.37 Python version: python 3.11 OS: linux Cloud/Infrastr…

---

## [Grafana Dashboard shows No Data for GPU metrics](https://discuss.ray.io/t/grafana-dashboard-shows-no-data-for-gpu-metrics/22833)

<div class="topic-metadata">

**Author:** [@bpistone](https://discuss.ray.io/u/bpistone)\
**Replies:** 0\
**Last updated:** [July 12, 2025, 11:43pm UTC](https://discuss.ray.io/t/grafana-dashboard-shows-no-data-for-gpu-metrics/22833 "2025-07-12T23:43:08Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Graceful shutdown of FastAPI Background task from Ray Serve on KubeRay](https://discuss.ray.io/t/graceful-shutdown-of-fastapi-background-task-from-ray-serve-on-kuberay/22807)

<div class="topic-metadata">

**Author:** [@duc\_manh\_le](https://discuss.ray.io/u/duc_manh_le)\
**Replies:** 0\
**Last updated:** [July 7, 2025, 1:29pm UTC](https://discuss.ray.io/t/graceful-shutdown-of-fastapi-background-task-from-ray-serve-on-kuberay/22807 "2025-07-07T13:29:42Z")

</div>

I have an Ray Serve application utilizing FastAPI Background task. When running the application locally, when interrupted by ctrl + c the script gracefully waiting for the fastapi background task to finish before termina…

---

## [Ray Job creating through existing ray cluster](https://discuss.ray.io/t/ray-job-creating-through-existing-ray-cluster/22750)

<div class="topic-metadata">

**Author:** [@canopus](https://discuss.ray.io/u/canopus)\
**Replies:** 3\
**Last updated:** [July 3, 2025, 8:32pm UTC](https://discuss.ray.io/t/ray-job-creating-through-existing-ray-cluster/22750 "2025-07-03T20:32:20Z")

</div>

Hi i am trying deploy rayjob on eks cluster through existing ray cluster i am getting error while using clusterselector you cannot mention suspend on rayjob yaml what is the fix

---

## [Ray 2.20 now pulls a grpcio package that doesn't match it's own requirement](https://discuss.ray.io/t/ray-2-20-now-pulls-a-grpcio-package-that-doesnt-match-its-own-requirement/22608)

<div class="topic-metadata">

**Author:** [@Ran\_Cao](https://discuss.ray.io/u/Ran_Cao)\
**Replies:** 0\
**Last updated:** [June 5, 2025, 7:14pm UTC](https://discuss.ray.io/t/ray-2-20-now-pulls-a-grpcio-package-that-doesnt-match-its-own-requirement/22608 "2025-06-05T19:14:07Z")

</div>

1. Severity of the issue: (select one) High: Completely blocks me. 2. Environment: Ray version: 2.20 Python version: 3.10 OS: linux Cloud/Infrastructure: aws Other libs/tools (if relevant): 3. What happened vs. wha…

---

## [How to use 'ray up' with kuberay?](https://discuss.ray.io/t/how-to-use-ray-up-with-kuberay/22533)

<div class="topic-metadata">

**Author:** [@vladidobro](https://discuss.ray.io/u/vladidobro)\
**Replies:** 2\
**Last updated:** [May 28, 2025, 9:19am UTC](https://discuss.ray.io/t/how-to-use-ray-up-with-kuberay/22533 "2025-05-28T09:19:34Z")

</div>

Hi, it seems like it should be possible to use ‘ray up’ with provider=kuberay, but i haven’t been able to make it work, nor find any documentation. The cluster config examples directory on github is also empty for kuber…

---

## [\[Serve\] RayServe Pods Stuck in Unready State Causing API Outages](https://discuss.ray.io/t/serve-rayserve-pods-stuck-in-unready-state-causing-api-outages/22566)

<div class="topic-metadata">

**Author:** [@sambbhavgarg](https://discuss.ray.io/u/sambbhavgarg)\
**Replies:** 0\
**Last updated:** [May 27, 2025, 7:25am UTC](https://discuss.ray.io/t/serve-rayserve-pods-stuck-in-unready-state-causing-api-outages/22566 "2025-05-27T07:25:40Z")

</div>

What happened + What you expected to happen Predict APIs on one of our endpoints were consistently returning errors, leading to outages in one of our production workflows. The issue later extended to other API routes too…

---

## [Remote ray cluster not spilling to disk](https://discuss.ray.io/t/remote-ray-cluster-not-spilling-to-disk/21217)

<div class="topic-metadata">

**Author:** [@Lichtz](https://discuss.ray.io/u/Lichtz)\
**Replies:** 2\
**Last updated:** [May 14, 2025, 7:14am UTC](https://discuss.ray.io/t/remote-ray-cluster-not-spilling-to-disk/21217 "2025-05-14T07:14:53Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. Hi all, We have been experimenting with large datasets using Modin on Ray (\> 50-80 GB CSVs of 20+ columns) and …

---

## [AssertionError: Session name does not match persisted value](https://discuss.ray.io/t/assertionerror-session-name-does-not-match-persisted-value/13067)

<div class="topic-metadata">

**Author:** [@alikleit](https://discuss.ray.io/u/alikleit)\
**Replies:** 3\
**Last updated:** [April 27, 2025, 5:02pm UTC](https://discuss.ray.io/t/assertionerror-session-name-does-not-match-persisted-value/13067 "2025-04-27T17:02:17Z")

</div>

Posted this on Slack but got no replies. I’m getting the following error when starting ray ray start --head: AssertionError: Session name session\_2023-12-06\_10-13-50\_173232\_14279 does not match persisted value b'sessio…

---

## [Remote worker nodes only alive for 30 seconds](https://discuss.ray.io/t/remote-worker-nodes-only-alive-for-30-seconds/8394)

<div class="topic-metadata">

**Author:** [@bananajoe182](https://discuss.ray.io/u/bananajoe182)\
**Replies:** 7\
**Last updated:** [April 24, 2025, 8:34am UTC](https://discuss.ray.io/t/remote-worker-nodes-only-alive-for-30-seconds/8394 "2025-04-24T08:34:19Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. Hi, I’m trying to setup a small on premise cluster of high end GPU machines. I managed to connect the worker no…

---

## [\[Core\] Task Status Check Failure in Ray Data Job with Preempted Workers](https://discuss.ray.io/t/core-task-status-check-failure-in-ray-data-job-with-preempted-workers/22343)

<div class="topic-metadata">

**Author:** [@dragongu](https://discuss.ray.io/u/dragongu)\
**Replies:** 2\
**Last updated:** [April 23, 2025, 2:43am UTC](https://discuss.ray.io/t/core-task-status-check-failure-in-ray-data-job-with-preempted-workers/22343 "2025-04-23T02:43:22Z")

</div>

When running a simple Ray Data job where worker pods are frequently preempted, the following error occurs: task\_manager.cc:1412: Check failed: it-\>second.GetStatus() == rpc::TaskStatus::PENDING\_NODE\_ASSIGNMENT task ID =…

---

## [Ray crashes on Raspberry Pi 5 (64-bit) due to unsupported jemalloc page size (64K) — unhandled runtime failure on ARM64](https://discuss.ray.io/t/ray-crashes-on-raspberry-pi-5-64-bit-due-to-unsupported-jemalloc-page-size-64k-unhandled-runtime-failure-on-arm64/22314)

<div class="topic-metadata">

**Author:** [@nincns](https://discuss.ray.io/u/nincns)\
**Replies:** 0\
**Last updated:** [April 14, 2025, 11:44pm UTC](https://discuss.ray.io/t/ray-crashes-on-raspberry-pi-5-64-bit-due-to-unsupported-jemalloc-page-size-64k-unhandled-runtime-failure-on-arm64/22314 "2025-04-14T23:44:15Z")

</div>

Title: Ray crashes on Raspberry Pi 5 (64-bit) due to unsupported jemalloc page size (64K) — unhandled runtime failure on ARM64 1. Severity of the issue: High: Completely blocks me. 2. Environment: Ray version: 2.…

---

## [Ray head stuck on ssh when implementing Cloudwatch](https://discuss.ray.io/t/ray-head-stuck-on-ssh-when-implementing-cloudwatch/22127)

<div class="topic-metadata">

**Author:** [@doyen](https://discuss.ray.io/u/doyen)\
**Replies:** 1\
**Last updated:** [March 28, 2025, 1:49am UTC](https://discuss.ray.io/t/ray-head-stuck-on-ssh-when-implementing-cloudwatch/22127 "2025-03-28T01:49:39Z")

</div>

We are running Ray on AWS EC2(Linux-Ubuntu) and attempting to integrate CloudWatch into our setup. However, after updating our YAML template and running ray up, the Ray head node is unable to launch any worker instances.…

---

## [Runtime Environment Caching with Ray Serve and Persistent Volumes](https://discuss.ray.io/t/runtime-environment-caching-with-ray-serve-and-persistent-volumes/22139)

<div class="topic-metadata">

**Author:** [@Mayank\_Garg](https://discuss.ray.io/u/Mayank_Garg)\
**Replies:** 1\
**Last updated:** [March 27, 2025, 9:00pm UTC](https://discuss.ray.io/t/runtime-environment-caching-with-ray-serve-and-persistent-volumes/22139 "2025-03-27T21:00:55Z")

</div>

Hello Ray Team, I’m working with Ray Serve on Kubernetes and using serveConfigV2 to deploy two applications, image\_classifier and text\_generator. Here’s a snippet of my serveConfigV2: serveConfigV2: | applications:…

---

## [Starting Ray on RKE2 Does Not work](https://discuss.ray.io/t/starting-ray-on-rke2-does-not-work/22149)

<div class="topic-metadata">

**Author:** [@DaGeRe](https://discuss.ray.io/u/DaGeRe)\
**Replies:** 0\
**Last updated:** [March 26, 2025, 9:34am UTC](https://discuss.ray.io/t/starting-ray-on-rke2-does-not-work/22149 "2025-03-26T09:34:01Z")

</div>

Hi, I’m currently trying to get ray working on rke2. Therefore, I executed the following steps: helm install kuberay-operator kuberay/kuberay-operator --version 1.3.0 kubectl get pods # wait until kuberay-operator-5c7…

---

## [serveConfig with import path](https://discuss.ray.io/t/serveconfig-with-import-path/22080)

<div class="topic-metadata">

**Author:** [@puppadas](https://discuss.ray.io/u/puppadas)\
**Replies:** 0\
**Last updated:** [March 18, 2025, 4:18pm UTC](https://discuss.ray.io/t/serveconfig-with-import-path/22080 "2025-03-18T16:18:40Z")

</div>

1. Severity of the issue: (select one) Medium: Significantly affects my productivity but can find a workaround. 2. Environment: Ray version: latest 2.43 Python version: 3.11 OS: Ubuntu Cloud/Infrastructure: On pre…

---

## [Initializing ray in multi-node environment with NCCL](https://discuss.ray.io/t/initializing-ray-in-multi-node-environment-with-nccl/21813)

<div class="topic-metadata">

**Author:** [@j93hahn](https://discuss.ray.io/u/j93hahn)\
**Replies:** 1\
**Last updated:** [March 13, 2025, 7:29pm UTC](https://discuss.ray.io/t/initializing-ray-in-multi-node-environment-with-nccl/21813 "2025-03-13T19:29:16Z")

</div>

How severe does this issue affect your experience of using Ray? Medium: It contributes to significant difficulty to complete my task, but I can work around it. Hi, I have a multi-node cluster, that has been initializ…

---

## [How to Use an Existing Public IP and Subnet for Ray Cluster on Azure](https://discuss.ray.io/t/how-to-use-an-existing-public-ip-and-subnet-for-ray-cluster-on-azure/21952)

<div class="topic-metadata">

**Author:** [@Jayasankar\_KK](https://discuss.ray.io/u/Jayasankar_KK)\
**Replies:** 2\
**Last updated:** [March 12, 2025, 3:50pm UTC](https://discuss.ray.io/t/how-to-use-an-existing-public-ip-and-subnet-for-ray-cluster-on-azure/21952 "2025-03-12T15:50:52Z")

</div>

Hi everyone, I’m setting up a Ray cluster on Azure, and I want to use an existing subnet, virtual network, and public IP instead of letting Ray create new ones. Here’s what I’ve done so far in my YAML configuration: y…

---

## [Setting up docker as a virtual Ray cluster](https://discuss.ray.io/t/setting-up-docker-as-a-virtual-ray-cluster/21949)

<div class="topic-metadata">

**Author:** [@Shahbaz\_Chaudhary](https://discuss.ray.io/u/Shahbaz_Chaudhary)\
**Replies:** 1\
**Last updated:** [March 11, 2025, 5:43pm UTC](https://discuss.ray.io/t/setting-up-docker-as-a-virtual-ray-cluster/21949 "2025-03-11T17:43:18Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. Hi, I’m trying to set up a virtual Ray cluster to learn and demonstrate various features of Ray. I could do this…

---

## [Ray.init() hangs on macOS M4](https://discuss.ray.io/t/ray-init-hangs-on-macos-m4/22006)

<div class="topic-metadata">

**Author:** [@j93hahn](https://discuss.ray.io/u/j93hahn)\
**Replies:** 3\
**Last updated:** [March 11, 2025, 2:54am UTC](https://discuss.ray.io/t/ray-init-hangs-on-macos-m4/22006 "2025-03-11T02:54:14Z")

</div>

I simply am doing: ray.init() on my MacOS M4 with ray version 2.42.1, and it hangs indefinitely. Is there any workaround? I have a need to do lots of simple, local testing with different ray functions and it’s quite bl…

---

## [Ray Blocking Spark Jobs](https://discuss.ray.io/t/ray-blocking-spark-jobs/21941)

<div class="topic-metadata">

**Author:** [@naushad](https://discuss.ray.io/u/naushad)\
**Replies:** 3\
**Last updated:** [March 11, 2025, 1:51am UTC](https://discuss.ray.io/t/ray-blocking-spark-jobs/21941 "2025-03-11T01:51:43Z")

</div>

How severe does this issue affect your experience of using Ray? Medium: It contributes to significant difficulty to complete my task, but I can work around it. I have setup ray cluster on Databricks runtime 15.4 LTS w…

---

## [Strange errors running Ray on M1 Mac using podman](https://discuss.ray.io/t/strange-errors-running-ray-on-m1-mac-using-podman/14713)

<div class="topic-metadata">

**Author:** [@blublinsky](https://discuss.ray.io/u/blublinsky)\
**Replies:** 7\
**Last updated:** [March 11, 2025, 12:31am UTC](https://discuss.ray.io/t/strange-errors-running-ray-on-m1-mac-using-podman/14713 "2025-03-11T00:31:50Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. We are trying to run Ray on a Kind cluster on M1 Mac with Podman. We are using KubeRay to create the cluster and…

---

## [Ray \<-\> Ray Operator compatibility](https://discuss.ray.io/t/ray-ray-operator-compatibility/21916)

<div class="topic-metadata">

**Author:** [@rnorden](https://discuss.ray.io/u/rnorden)\
**Replies:** 1\
**Last updated:** [March 10, 2025, 10:33pm UTC](https://discuss.ray.io/t/ray-ray-operator-compatibility/21916 "2025-03-10T22:33:40Z")

</div>

Hi, I am using kuberay-operator-1.2.2 which uses ray-project/ray 2.9.0 I have seen this compatibility matrix: upgrade-guide.html#kuberay-ray-compatibility I want to train on a custom environment written in Python, …

---

## [Ray cluster-launcher not starting up properly](https://discuss.ray.io/t/ray-cluster-launcher-not-starting-up-properly/21693)

<div class="topic-metadata">

**Author:** [@mt-clemente](https://discuss.ray.io/u/mt-clemente)\
**Replies:** 3\
**Last updated:** [March 6, 2025, 4:45pm UTC](https://discuss.ray.io/t/ray-cluster-launcher-not-starting-up-properly/21693 "2025-03-06T16:45:46Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. Hi, I am trying to set up a ray cluster but it seems to be hanging at some point in the process. The ray up clus…

---

## [Question: How to set SSH port for nodes in auto\_scaler YAML?](https://discuss.ray.io/t/question-how-to-set-ssh-port-for-nodes-in-auto-scaler-yaml/1972)

<div class="topic-metadata">

**Author:** [@xmzzyo](https://discuss.ray.io/u/xmzzyo)\
**Replies:** 1\
**Last updated:** [May 1, 2021, 11:43am UTC](https://discuss.ray.io/t/question-how-to-set-ssh-port-for-nodes-in-auto-scaler-yaml/1972 "2021-05-01T11:43:23Z")

</div>

Hi all. The SSH port for my server is 1022 rather than the default port (22). However, when I set the YAML file as follows: provider: type: local head\_ip: 10.0.0.40 worker\_ips: \[10.0.0.39, 10.0.0.41\] Th…

---

## [Workers crashes after few seconds automatically](https://discuss.ray.io/t/workers-crashes-after-few-seconds-automatically/13444)

<div class="topic-metadata">

**Author:** [@Abhishek\_Jaiswal](https://discuss.ray.io/u/Abhishek_Jaiswal)\
**Replies:** 1\
**Last updated:** [March 5, 2025, 12:57pm UTC](https://discuss.ray.io/t/workers-crashes-after-few-seconds-automatically/13444 "2025-03-05T12:57:31Z")

</div>

My workers nodes and head are running on different vm , i am able to see my workers on dashboard , but after few seconds workers gets killed automatically 1The node with node id: 3428fdc8ee0b2accb9f106e3f63207b3722a65c9…

---

## [How to Use an Existing Public IP and Subnet for Ray Cluster on Azure?](https://discuss.ray.io/t/how-to-use-an-existing-public-ip-and-subnet-for-ray-cluster-on-azure/21951)

<div class="topic-metadata">

**Author:** [@Jayasankar\_KK](https://discuss.ray.io/u/Jayasankar_KK)\
**Replies:** 2\
**Last updated:** [March 4, 2025, 5:54pm UTC](https://discuss.ray.io/t/how-to-use-an-existing-public-ip-and-subnet-for-ray-cluster-on-azure/21951 "2025-03-04T17:54:45Z")

</div>

Hi everyone, I’m setting up a Ray cluster on Azure, and I want to use an existing subnet, virtual network, and public IP instead of letting Ray create new ones. Here’s what I’ve done so far in my YAML configuration: y…

[Previous page](https://discuss.ray.io/c/ray-clusters/13.md)

[Next page](https://discuss.ray.io/c/ray-clusters/13.md?page=2)
