# Automatically restart head node on kubernetes

**URL:** <https://discuss.ray.io/t/automatically-restart-head-node-on-kubernetes/2570>\
**Category:** Kubernetes\
**Created:** [June 18, 2021, 2:42pm UTC](https://discuss.ray.io/t/automatically-restart-head-node-on-kubernetes/2570 "2021-06-18T14:42:53Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![simenandresen](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/simenandresen/32/1066_2.png) [@simenandresen](https://discuss.ray.io/u/simenandresen)\
**Post date:** [June 18, 2021, 2:42pm UTC](https://discuss.ray.io/t/automatically-restart-head-node-on-kubernetes/2570/1 "2021-06-18T14:42:53Z")

</div>

I’m running ray (1.4.0) on kubernetes (in GKE) and have problem with our head-node pod being removed. I’m not sure why the pod is removed in the first place, however, with kubernetes one can typically set the restartPolicy=Always to make it restarts after an outage.

I tried `setting restartPolicy=Always` in our config.yaml (similar to the [kubernetes example](https://github.com/ray-project/ray/blob/master/python/ray/autoscaler/kubernetes/example-full.yaml#L167)), but when running `ray.init` after the head node restarts, I just get the following error:

`RuntimeError: Unable to connect to Redis at 10.43.250.14:6379 after 12 retries. Check that 10.43.250.14:6379 is reachable from this machine. If it is not, your firewall may be blocking this port. If the problem is a flaky connection, try setting the environment variable `RAY\_START\_REDIS\_WAIT\_RETRIES` to increase the number of attempts to ping the Redis server.`

Is setting `restartPolicy=Always` supposed to work, or is there another way to make sure the head node stays up?

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [June 22, 2021, 4:24pm UTC](https://discuss.ray.io/t/automatically-restart-head-node-on-kubernetes/2570/2 "2021-06-22T16:24:36Z")

</div>

> [@simenandresen](#):
>
> I tried `setting restartPolicy=Always` in our config.yaml (similar to the [kubernetes example](https://github.com/ray-project/ray/blob/master/python/ray/autoscaler/kubernetes/example-full.yaml#L167)), but when running `ray.init` after the head node restarts, I just get the following error:

Hey @simenandresen, thanks a bunch for making this issue!

@tgaddair @Dmitri do you know if this is resolved with the Kopf integration?

---

<div class="post-metadata">

**Author:** ![Dmitri](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/dmitri/32/657_2.png) [@Dmitri](https://discuss.ray.io/u/Dmitri)\
**Post date:** [June 22, 2021, 5:16pm UTC](https://discuss.ray.io/t/automatically-restart-head-node-on-kubernetes/2570/3 "2021-06-22T17:16:09Z")

</div>

The K8s operator is now the recommended tool to launch Ray on K8s. The operator handles head restarts correctly.

[https://docs.ray.io/en/master/cluster/kubernetes.html](https://docs.ray.io/en/master/cluster/kubernetes.html)  
[https://docs.ray.io/en/master/cluster/kubernetes-advanced.html#restart-behavior](https://docs.ray.io/en/master/cluster/kubernetes-advanced.html#restart-behavior)

Note that a head restart will reboot all ray processes/state – you might want to protect the head pod with a pod disruption budget / pod priority.

---

<div class="post-metadata">

**Author:** ![simenandresen](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/simenandresen/32/1066_2.png) [@simenandresen](https://discuss.ray.io/u/simenandresen)\
**Post date:** [June 24, 2021, 6:03am UTC](https://discuss.ray.io/t/automatically-restart-head-node-on-kubernetes/2570/4 "2021-06-24T06:03:28Z")

</div>

Thanks @Dmitri, I’ll try out setting up ray using the ray kubernetes operator
