# Kubernetes

**URL:** https://discuss.ray.io/c/ray-clusters/ray-kubernetes/11.md

[Latest](https://discuss.ray.io/latest.md) · [Categories](https://discuss.ray.io/categories.md) · [Tags](https://discuss.ray.io/tags.md)

---

## [About the Kubernetes category](https://discuss.ray.io/t/about-the-kubernetes-category/621)

<div class="topic-metadata">

**Author:** [@rliaw](https://discuss.ray.io/u/rliaw)\
**Replies:** 0

</div>

The Ray team is working hard to make Ray work really well on Kubernetes! Ray provides a Kubernetes Operator for managing autoscaling Ray clusters. Using the operator provides similar functionality to deploying a Ray clu…

---

## [How does Ray actor work?](https://discuss.ray.io/t/how-does-ray-actor-work/22206)

<div class="topic-metadata">

**Author:** [@Hongbo-Miao](https://discuss.ray.io/u/Hongbo-Miao)\
**Replies:** 2\
**Last updated:** [July 16, 2026, 8:04pm UTC](https://discuss.ray.io/t/how-does-ray-actor-work/22206 "2026-07-16T20:04:59Z")

</div>

We have a Ray cluster running, but today, for about 1.5 hours, the cluster was up, yet no Ray actors were running. During this time, users were unable to submit any Ray jobs. Here is a simplified version of our deplo…

---

## [Is it possible to support multi batchScheduler for kuberay](https://discuss.ray.io/t/is-it-possible-to-support-multi-batchscheduler-for-kuberay/23452)

<div class="topic-metadata">

**Author:** [@jacklove2run](https://discuss.ray.io/u/jacklove2run)\
**Replies:** 4\
**Last updated:** [June 1, 2026, 12:12pm UTC](https://discuss.ray.io/t/is-it-possible-to-support-multi-batchscheduler-for-kuberay/23452 "2026-06-01T12:12:42Z")

</div>

Hi everyone, we are using kuberay to submit RayJob to K8s cluster. The batchScheduler was set to false which means using the default-scheduler for batch schedulering. Now we want to set up volcano scheduler for our clust…

---

## [Ray Serve LLM on KubeRay with CPU](https://discuss.ray.io/t/ray-serve-llm-on-kuberay-with-cpu/23524)

<div class="topic-metadata">

**Author:** [@CaoimheCahill](https://discuss.ray.io/u/CaoimheCahill)\
**Replies:** 1\
**Last updated:** [April 30, 2026, 9:57am UTC](https://discuss.ray.io/t/ray-serve-llm-on-kuberay-with-cpu/23524 "2026-04-30T09:57:16Z")

</div>

1. Severity of the issue: None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Completely blocks m…

---

## [Why does the Ray job driver process contain autoscaler logs even when autoscaling is disabled?](https://discuss.ray.io/t/why-does-the-ray-job-driver-process-contain-autoscaler-logs-even-when-autoscaling-is-disabled/23423)

<div class="topic-metadata">

**Author:** [@psnilesh](https://discuss.ray.io/u/psnilesh)\
**Replies:** 1\
**Last updated:** [December 29, 2025, 8:28pm UTC](https://discuss.ray.io/t/why-does-the-ray-job-driver-process-contain-autoscaler-logs-even-when-autoscaling-is-disabled/23423 "2025-12-29T20:28:39Z")

</div>

I have a KubeRay cluster that does not have autoscaling enabled. I have not set enableInTreeAutoscaling or autoscalerOptions in my RayCluster resource. However, I still see below autoscaler related logs when I do ray job…

---

## [Dependency loading appears to race worker startup on Ray running on GKE with KubeRay and uv, leading to missing modules and .so errors.](https://discuss.ray.io/t/dependency-loading-appears-to-race-worker-startup-on-ray-running-on-gke-with-kuberay-and-uv-leading-to-missing-modules-and-so-errors/23416)

<div class="topic-metadata">

**Author:** [@Warlink](https://discuss.ray.io/u/Warlink)\
**Replies:** 0\
**Last updated:** [December 27, 2025, 9:38pm UTC](https://discuss.ray.io/t/dependency-loading-appears-to-race-worker-startup-on-ray-running-on-gke-with-kuberay-and-uv-leading-to-missing-modules-and-so-errors/23416 "2025-12-27T21:38:04Z")

</div>

Hi all, I am running into what looks like a dependency race condition when running Ray on GKE using KubeRay and would appreciate any insight. When training a small toy model from the Ray documentation, workers appear t…

---

## [Ray Clusters with Bazel](https://discuss.ray.io/t/ray-clusters-with-bazel/23384)

<div class="topic-metadata">

**Author:** [@aniketsinghrawat](https://discuss.ray.io/u/aniketsinghrawat)\
**Replies:** 2\
**Last updated:** [December 17, 2025, 10:37pm UTC](https://discuss.ray.io/t/ray-clusters-with-bazel/23384 "2025-12-17T22:37:17Z")

</div>

I am sure this is a solved problem, but I want to confirm what I intend to do is correct. I am building my python project using Bazel and I want to run some distributed workload on Ray Clusters on GKE. The issue is when…

---

## [Domestic GPU recognition and adaptation](https://discuss.ray.io/t/domestic-gpu-recognition-and-adaptation/23345)

<div class="topic-metadata">

**Author:** [@Hxinyue](https://discuss.ray.io/u/Hxinyue)\
**Replies:** 2\
**Last updated:** [December 1, 2025, 11:59am UTC](https://discuss.ray.io/t/domestic-gpu-recognition-and-adaptation/23345 "2025-12-01T11:59:23Z")

</div>

For domestic Gpus such as the K100 of Sugon DCU, when the ray cluster is started through k8s, the autoscaler of Ray-head cannot automatically recognize it as a GPU. Therefore, resources=“{“DCU”:1}” is configured in the r…

---

## [Deploying RayCluster: Readiness and Liveness Probes for the Head Node Continuously Failing](https://discuss.ray.io/t/deploying-raycluster-readiness-and-liveness-probes-for-the-head-node-continuously-failing/23049)

<div class="topic-metadata">

**Author:** [@MiniSho](https://discuss.ray.io/u/MiniSho)\
**Replies:** 0\
**Last updated:** [August 27, 2025, 3:26am UTC](https://discuss.ray.io/t/deploying-raycluster-readiness-and-liveness-probes-for-the-head-node-continuously-failing/23049 "2025-08-27T03:26:49Z")

</div>

\[Ray Usage Environment\] Production \[Ray Version and Libraries\] ver 3.0.0 dev \[Environment\] CentOS 7, Kubernetes (k8s) \[Issue Reproduction\] After initializing KubeRay, I attempted to start the Ray cluster. Since our …

---

## [Graceful shutdown of FastAPI Background task from Ray Serve on KubeRay](https://discuss.ray.io/t/graceful-shutdown-of-fastapi-background-task-from-ray-serve-on-kuberay/22807)

<div class="topic-metadata">

**Author:** [@duc\_manh\_le](https://discuss.ray.io/u/duc_manh_le)\
**Replies:** 0\
**Last updated:** [July 7, 2025, 1:29pm UTC](https://discuss.ray.io/t/graceful-shutdown-of-fastapi-background-task-from-ray-serve-on-kuberay/22807 "2025-07-07T13:29:42Z")

</div>

I have an Ray Serve application utilizing FastAPI Background task. When running the application locally, when interrupted by ctrl + c the script gracefully waiting for the fastapi background task to finish before termina…

---

## [Ray Job creating through existing ray cluster](https://discuss.ray.io/t/ray-job-creating-through-existing-ray-cluster/22750)

<div class="topic-metadata">

**Author:** [@canopus](https://discuss.ray.io/u/canopus)\
**Replies:** 3\
**Last updated:** [July 3, 2025, 8:32pm UTC](https://discuss.ray.io/t/ray-job-creating-through-existing-ray-cluster/22750 "2025-07-03T20:32:20Z")

</div>

Hi i am trying deploy rayjob on eks cluster through existing ray cluster i am getting error while using clusterselector you cannot mention suspend on rayjob yaml what is the fix

---

## [How to use 'ray up' with kuberay?](https://discuss.ray.io/t/how-to-use-ray-up-with-kuberay/22533)

<div class="topic-metadata">

**Author:** [@vladidobro](https://discuss.ray.io/u/vladidobro)\
**Replies:** 2\
**Last updated:** [May 28, 2025, 9:19am UTC](https://discuss.ray.io/t/how-to-use-ray-up-with-kuberay/22533 "2025-05-28T09:19:34Z")

</div>

Hi, it seems like it should be possible to use ‘ray up’ with provider=kuberay, but i haven’t been able to make it work, nor find any documentation. The cluster config examples directory on github is also empty for kuber…

---

## [\[Serve\] RayServe Pods Stuck in Unready State Causing API Outages](https://discuss.ray.io/t/serve-rayserve-pods-stuck-in-unready-state-causing-api-outages/22566)

<div class="topic-metadata">

**Author:** [@sambbhavgarg](https://discuss.ray.io/u/sambbhavgarg)\
**Replies:** 0\
**Last updated:** [May 27, 2025, 7:25am UTC](https://discuss.ray.io/t/serve-rayserve-pods-stuck-in-unready-state-causing-api-outages/22566 "2025-05-27T07:25:40Z")

</div>

What happened + What you expected to happen Predict APIs on one of our endpoints were consistently returning errors, leading to outages in one of our production workflows. The issue later extended to other API routes too…

---

## [AssertionError: Session name does not match persisted value](https://discuss.ray.io/t/assertionerror-session-name-does-not-match-persisted-value/13067)

<div class="topic-metadata">

**Author:** [@alikleit](https://discuss.ray.io/u/alikleit)\
**Replies:** 3\
**Last updated:** [April 27, 2025, 5:02pm UTC](https://discuss.ray.io/t/assertionerror-session-name-does-not-match-persisted-value/13067 "2025-04-27T17:02:17Z")

</div>

Posted this on Slack but got no replies. I’m getting the following error when starting ray ray start --head: AssertionError: Session name session\_2023-12-06\_10-13-50\_173232\_14279 does not match persisted value b'sessio…

---

## [Runtime Environment Caching with Ray Serve and Persistent Volumes](https://discuss.ray.io/t/runtime-environment-caching-with-ray-serve-and-persistent-volumes/22139)

<div class="topic-metadata">

**Author:** [@Mayank\_Garg](https://discuss.ray.io/u/Mayank_Garg)\
**Replies:** 1\
**Last updated:** [March 27, 2025, 9:00pm UTC](https://discuss.ray.io/t/runtime-environment-caching-with-ray-serve-and-persistent-volumes/22139 "2025-03-27T21:00:55Z")

</div>

Hello Ray Team, I’m working with Ray Serve on Kubernetes and using serveConfigV2 to deploy two applications, image\_classifier and text\_generator. Here’s a snippet of my serveConfigV2: serveConfigV2: | applications:…

---

## [Starting Ray on RKE2 Does Not work](https://discuss.ray.io/t/starting-ray-on-rke2-does-not-work/22149)

<div class="topic-metadata">

**Author:** [@DaGeRe](https://discuss.ray.io/u/DaGeRe)\
**Replies:** 0\
**Last updated:** [March 26, 2025, 9:34am UTC](https://discuss.ray.io/t/starting-ray-on-rke2-does-not-work/22149 "2025-03-26T09:34:01Z")

</div>

Hi, I’m currently trying to get ray working on rke2. Therefore, I executed the following steps: helm install kuberay-operator kuberay/kuberay-operator --version 1.3.0 kubectl get pods # wait until kuberay-operator-5c7…

---

## [Ray \<-\> Ray Operator compatibility](https://discuss.ray.io/t/ray-ray-operator-compatibility/21916)

<div class="topic-metadata">

**Author:** [@rnorden](https://discuss.ray.io/u/rnorden)\
**Replies:** 1\
**Last updated:** [March 10, 2025, 10:33pm UTC](https://discuss.ray.io/t/ray-ray-operator-compatibility/21916 "2025-03-10T22:33:40Z")

</div>

Hi, I am using kuberay-operator-1.2.2 which uses ray-project/ray 2.9.0 I have seen this compatibility matrix: upgrade-guide.html#kuberay-ray-compatibility I want to train on a custom environment written in Python, …

---

## [How to set Ray head node in high availability mode using KubeRay Helm chart?](https://discuss.ray.io/t/how-to-set-ray-head-node-in-high-availability-mode-using-kuberay-helm-chart/21914)

<div class="topic-metadata">

**Author:** [@Hongbo-Miao](https://discuss.ray.io/u/Hongbo-Miao)\
**Replies:** 0\
**Last updated:** [February 26, 2025, 7:16am UTC](https://discuss.ray.io/t/how-to-set-ray-head-node-in-high-availability-mode-using-kuberay-helm-chart/21914 "2025-02-26T07:16:07Z")

</div>

Originally asked at How to set Ray head node in high availability mode using KubeRay Helm chart? - Stack Overflow Here is a copy: I am trying to set up high availability (HA) for Ray head node. Currently, if Ray head …

---

## [KubeRay clusters fail to start when workers memory limit \>=4GiB](https://discuss.ray.io/t/kuberay-clusters-fail-to-start-when-workers-memory-limit-4gib/21057)

<div class="topic-metadata">

**Author:** [@Noobfire](https://discuss.ray.io/u/Noobfire)\
**Replies:** 2\
**Last updated:** [December 13, 2024, 8:00am UTC](https://discuss.ray.io/t/kuberay-clusters-fail-to-start-when-workers-memory-limit-4gib/21057 "2024-12-13T08:00:51Z")

</div>

Hello, I’m using KubeRay on an internal k8s cluster with really beefy machines (hundreds of cores, 2.3TiB RAM). Setting up clusters in general through Helm Charts and using the KubeRay operator works without a problem: …

---

## [\[Autoscaler\]\[K8s\] Is it possible to configure the autoscaler to minimize resource usage?](https://discuss.ray.io/t/autoscaler-k8s-is-it-possible-to-configure-the-autoscaler-to-minimize-resource-usage/21012)

<div class="topic-metadata">

**Author:** [@aelskens](https://discuss.ray.io/u/aelskens)\
**Replies:** 0\
**Last updated:** [December 10, 2024, 3:28pm UTC](https://discuss.ray.io/t/autoscaler-k8s-is-it-possible-to-configure-the-autoscaler-to-minimize-resource-usage/21012 "2024-12-10T15:28:21Z")

</div>

Hi, I have deployed three worker groups with the following differences in their resource configurations: resources.limits.memory=2G, resources.requests.memory=2G, and maxReplicas=15. resources.limits.memory=5G, resour…

---

## [Timed out while waiting for GCS to become available](https://discuss.ray.io/t/timed-out-while-waiting-for-gcs-to-become-available/20613)

<div class="topic-metadata">

**Author:** [@Mtaibeer](https://discuss.ray.io/u/Mtaibeer)\
**Replies:** 5\
**Last updated:** [November 18, 2024, 8:11am UTC](https://discuss.ray.io/t/timed-out-while-waiting-for-gcs-to-become-available/20613 "2024-11-18T08:11:54Z")

</div>

When I execute ray-cluster in k8s, my worker pod keeps initializing, the log shows： Timed out while waiting for GCS to become available log gcs\_server.err in /tmp/ray/session\_latest/logs/ is empty. The k8s network is …

---

## [Ray-worker pod is waiting to start](https://discuss.ray.io/t/ray-worker-pod-is-waiting-to-start/20490)

<div class="topic-metadata">

**Author:** [@Mtaibeer](https://discuss.ray.io/u/Mtaibeer)\
**Replies:** 5\
**Last updated:** [November 11, 2024, 2:03pm UTC](https://discuss.ray.io/t/ray-worker-pod-is-waiting-to-start/20490 "2024-11-11T14:03:10Z")

</div>

I follow the doc to start my cluster with the default yaml configuration kuberay/ray-operator/config/samples/ray-cluster.complete.yaml at master · ray-project/kuberay, worker nodes always display container “ray-worker” …

---

## [Don't we provide a way to build ray images from source code?](https://discuss.ray.io/t/dont-we-provide-a-way-to-build-ray-images-from-source-code/20312)

<div class="topic-metadata">

**Author:** [@dragongu](https://discuss.ray.io/u/dragongu)\
**Replies:** 1\
**Last updated:** [November 5, 2024, 6:37pm UTC](https://discuss.ray.io/t/dont-we-provide-a-way-to-build-ray-images-from-source-code/20312 "2024-11-05T18:37:41Z")

</div>

I don’t seem to see a way to build a ray image from the source code, the currently provided build-docker.sh is still a wheel download from the Internet

---

## [Cannot create directory '/mnt/cluster\_storage'](https://discuss.ray.io/t/cannot-create-directory-mnt-cluster-storage/12382)

<div class="topic-metadata">

**Author:** [@swapkh91](https://discuss.ray.io/u/swapkh91)\
**Replies:** 1\
**Last updated:** [October 23, 2024, 9:47pm UTC](https://discuss.ray.io/t/cannot-create-directory-mnt-cluster-storage/12382 "2024-10-23T21:47:07Z")

</div>

I tried to run the XGBoost training example as described here on a Ray Cluster created using ray-ml:2.7.0-py310-cpu on GKE. When I submit the job by running the script in the link, I get PermissionError: \[Errno 13\] Cann…

---

## [Kuberay operator upgrade from v1.0.0 to v1.2.2](https://discuss.ray.io/t/kuberay-operator-upgrade-from-v1-0-0-to-v1-2-2/16019)

<div class="topic-metadata">

**Author:** [@Sowmith\_Renumakala](https://discuss.ray.io/u/Sowmith_Renumakala)\
**Replies:** 1\
**Last updated:** [October 18, 2024, 10:30am UTC](https://discuss.ray.io/t/kuberay-operator-upgrade-from-v1-0-0-to-v1-2-2/16019 "2024-10-18T10:30:04Z")

</div>

Hello We have kuberay operator which is deployed at namespace level at version v1.0.0. We want to upgrade to latest version of v1.2.2. Cant find enough documentation on If there are any breaking changes of directly up…

---

## [Ray Service not able to load code outside current app directory](https://discuss.ray.io/t/ray-service-not-able-to-load-code-outside-current-app-directory/16031)

<div class="topic-metadata">

**Author:** [@Parag\_Patil](https://discuss.ray.io/u/Parag_Patil)\
**Replies:** 1\
**Last updated:** [October 18, 2024, 10:42am UTC](https://discuss.ray.io/t/ray-service-not-able-to-load-code-outside-current-app-directory/16031 "2024-10-18T10:42:57Z")

</div>

I have folder structure src util app/app.py I have copied this code inside docker image and when I’m runnning ray service(app.py) I’m getting “src module not found error”. Incase if I mentioned zip file in working dir…

---

## [What is \`configmaps/status\` subresource and why is it needed?](https://discuss.ray.io/t/what-is-configmaps-status-subresource-and-why-is-it-needed/15796)

<div class="topic-metadata">

**Author:** [@lindhe](https://discuss.ray.io/u/lindhe)\
**Replies:** 0\
**Last updated:** [September 13, 2024, 2:18pm UTC](https://discuss.ray.io/t/what-is-configmaps-status-subresource-and-why-is-it-needed/15796 "2024-09-13T14:18:39Z")

</div>

Hi! I’m a Kubernetes admin and we have a team that wants to deploy the KubeRay operator in our cluster. When trying to do that, they were unable to install the leader election Role, because it includes get, update and p…

---

## [Set dfloat arg for KubeRay vLLM example](https://discuss.ray.io/t/set-dfloat-arg-for-kuberay-vllm-example/15771)

<div class="topic-metadata">

**Author:** [@marvin.steinke](https://discuss.ray.io/u/marvin.steinke)\
**Replies:** 0\
**Last updated:** [September 11, 2024, 8:49am UTC](https://discuss.ray.io/t/set-dfloat-arg-for-kuberay-vllm-example/15771 "2024-09-11T08:49:34Z")

</div>

Im trying to run this KubeRay Serve example. However, with the default settings(ray-service.vllm.yaml), this error occurs upon deployment: ValueError: Bfloat16 is only supported on GPUs with compute capability of at lea…

---

## [\[Cluster\]\[Autoscaler-v2\]-Autoscaler v2 does not honor minReplicas/replicas count of the worker nodes and constantly terminates after idletimeout](https://discuss.ray.io/t/cluster-autoscaler-v2-autoscaler-v2-does-not-honor-minreplicas-replicas-count-of-the-worker-nodes-and-constantly-terminates-after-idletimeout/15763)

<div class="topic-metadata">

**Author:** [@vbalumuri](https://discuss.ray.io/u/vbalumuri)\
**Replies:** 0\
**Last updated:** [September 10, 2024, 2:55pm UTC](https://discuss.ray.io/t/cluster-autoscaler-v2-autoscaler-v2-does-not-honor-minreplicas-replicas-count-of-the-worker-nodes-and-constantly-terminates-after-idletimeout/15763 "2024-09-10T14:55:25Z")

</div>

Github issue link: \[Cluster\]\[Autoscaler-v2\]-Autoscaler v2 does not honor minReplicas/replicas count of the worker nodes and constantly terminates after idletimeout · Issue #47578 · ray-project/ray · GitHub What happened…

---

## [\[Cluster\] Multiple programs running on one ray cluster](https://discuss.ray.io/t/cluster-multiple-programs-running-on-one-ray-cluster/566)

<div class="topic-metadata">

**Author:** [@delioda79](https://discuss.ray.io/u/delioda79)\
**Replies:** 9\
**Last updated:** [September 9, 2024, 2:01pm UTC](https://discuss.ray.io/t/cluster-multiple-programs-running-on-one-ray-cluster/566 "2024-09-09T14:01:08Z")

</div>

if I have say 10 projects I would like to use ray with, if I deploy all of them in my cluster, they will all share the same runtime, but would this influence the performance somehow? I mean I’m planning to use ray, if it…

[Next page](https://discuss.ray.io/c/ray-clusters/ray-kubernetes/11.md?page=1)
