# Kubernetes cluster only creates head node

**URL:** <https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366>\
**Category:** Ray Clusters\
**Created:** [June 1, 2022, 10:22pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366 "2022-06-01T22:22:45Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![cloudhaxor](https://avatars.discourse-cdn.com/v4/letter/c/82dd89/32.png) [@cloudhaxor](https://discuss.ray.io/u/cloudhaxor)\
**Post date:** [June 1, 2022, 10:22pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/1 "2022-06-01T22:22:45Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- High: It blocks me to complete my task.

Hi, I followed the guide on [Deploying on Kubernetes — Ray 1.12.1](https://docs.ray.io/en/latest/cluster/kubernetes.html) for deployment on Kubernetes. Everything goes well with no issues. When executing the command `kubectl -n ray get pods` it only shows the head node in “Running” status. I am using an unmodified version of the config yaml file. I also replicated the commands from top to bottom on that page to make sure its not an issue. Any ideas?

Thanks

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [June 1, 2022, 10:29pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/2 "2022-06-01T22:29:17Z")

</div>

Hey, can you confirm that your kubernetes cluster has enough available resources to run the ray cluster?

Could you also share the logs from your operator? (`kubectl -n ray logs ray-operator`)

---

<div class="post-metadata">

**Author:** ![cloudhaxor](https://avatars.discourse-cdn.com/v4/letter/c/82dd89/32.png) [@cloudhaxor](https://discuss.ray.io/u/cloudhaxor)\
**Post date:** [June 1, 2022, 10:46pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/3 "2022-06-01T22:46:47Z")

</div>

> [@Alex](#):
>
> `kubectl -n ray logs ray-operator`

I cannot share all of the logs of that file but the only error I see is

```auto
grpc._channel._InactiveRpcError: <_InactiveRpcError of RPC that terminated with:
	status = StatusCode.UNAVAILABLE
	details = "failed to connect to all addresses"

```

The resources on all of the nodes are more than enough to run Ray (RTX Titan/ modern xeon processor)

The network is heavily firewalled so I am unsure if that may be the issue (blocked ports?)

Everything from the install page shows “Running”

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [June 2, 2022, 5:35pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/4 "2022-06-02T17:35:17Z")

</div>

> [@cloudhaxor](#):
>
> The network is heavily firewalled so I am unsure if that may be the issue (blocked ports?)

That sounds like a likely/possible cause. Have you had a chance to check out the port configuration page?

[https://docs.ray.io/en/latest/ray-core/configure.html#ray-ports](https://docs.ray.io/en/latest/ray-core/configure.html#ray-ports)

---

<div class="post-metadata">

**Author:** ![cloudhaxor](https://avatars.discourse-cdn.com/v4/letter/c/82dd89/32.png) [@cloudhaxor](https://discuss.ray.io/u/cloudhaxor)\
**Post date:** [June 2, 2022, 8:33pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/5 "2022-06-02T20:33:40Z")

</div>

So this did work, but only for the static non-autoscaling version of Ray. I have it up and running without Ray operator. Any recommendations for the Ray operator version?

Also side question:  
Should the #of replicas be equal to the number of nodes?

I ask because the replicas launch on the same node unless there a large number of replicas (i.e set them to 70 and they disperse unequally)

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [June 6, 2022, 5:37pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/6 "2022-06-06T17:37:25Z")

</div>

What happens on the version with the ray operator? Are there any error messages?

---

<div class="post-metadata">

**Author:** ![cloudhaxor](https://avatars.discourse-cdn.com/v4/letter/c/82dd89/32.png) [@cloudhaxor](https://discuss.ray.io/u/cloudhaxor)\
**Post date:** [June 6, 2022, 6:56pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/7 "2022-06-06T18:56:20Z")

</div>

No errors on that.

I have switched to KubeRay per some recommendations through similar issues in GitHub. Now, with KubeRay everything deploys but the worker seems to be stuck in the PodInitizalization stage with no errors. Any ideas?

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [June 6, 2022, 9:38pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/8 "2022-06-06T21:38:27Z")

</div>

can you `kubectl describe pod ` on the pod that stuck? the events there should provide some clues.

---

<div class="post-metadata">

**Author:** ![cloudhaxor](https://avatars.discourse-cdn.com/v4/letter/c/82dd89/32.png) [@cloudhaxor](https://discuss.ray.io/u/cloudhaxor)\
**Post date:** [June 7, 2022, 4:01pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/9 "2022-06-07T16:01:02Z")

</div>

Yes, so this led me down a rabbit hole yesterday which has given me a lot more success! But, now I have another issue and I’m unsure how to move forward.

Turns out there was a coredns issue with the kubernetes cluster. It was unable to resolve nameservers and so busybox would loop infinitely. I fixed this issue and nameservers can be resolved. Now, when ray container spins up it attempts to connect to raycluster-complete-head-svc:6379. Somehow, it is unable to resolve the underlying IP address, and server coredns errors popup like  
[ERROR] plugin/errors: 2 raycluster-complete-head-svc. A: read udp → i/o timeout.  
I’ve tried troubleshooting this online, but it’s not working.  
One thing I noticed is that if I hard code the IP address of the service within the ray start command, the pods spin up and actually connect to the head. So the issue is most definitely related to CoreDNS and Ray start command.

---

<div class="post-metadata">

**Author:** ![cloudhaxor](https://avatars.discourse-cdn.com/v4/letter/c/82dd89/32.png) [@cloudhaxor](https://discuss.ray.io/u/cloudhaxor)\
**Post date:** [June 7, 2022, 8:00pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/10 "2022-06-07T20:00:49Z")

</div>

So I have solved the issue:

firewall-cmd --add-masquerade --permanent

and restarting the firewall fixed the issue.  
Ray now works as intended.

---

<div class="post-metadata">

**Author:** ![Dmitri](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/dmitri/32/657_2.png) [@Dmitri](https://discuss.ray.io/u/Dmitri)\
**Post date:** [June 7, 2022, 9:32pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/11 "2022-06-07T21:32:53Z")

</div>

Glad you were able to work this out – all the errors you’ve described are consistent with firewalls in your K8s cluster. We’ll keep in mind to clearly document what kind of network communication Ray-on-K8s components are doing.

---

<div class="post-metadata">

**Author:** ![Dmitri](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/dmitri/32/657_2.png) [@Dmitri](https://discuss.ray.io/u/Dmitri)\
**Post date:** [June 7, 2022, 9:38pm UTC](https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/12 "2022-06-07T21:38:56Z")

</div>

> <https://github.com/ray-project/ray/issues/25565>
>
> \### Description
> 
> Related to https://github.com/ray-project/ray/issues/24491.
> 
> …The docs should have clear explanations of the network details of how Ray components interact.
> In particular, these explanations and how they relate to K8s cluster setup should be clear for K8s users.
> See https://discuss.ray.io/t/kubernetes-cluster-only-creates-head-node/6366/9
> 
> \### Link
> 
> \_No response\_
