# Start cluster with multiple head node

**URL:** <https://discuss.ray.io/t/start-cluster-with-multiple-head-node/9425>\
**Category:** Ray Core\
**Created:** [February 20, 2023, 6:47am UTC](https://discuss.ray.io/t/start-cluster-with-multiple-head-node/9425 "2023-02-20T06:47:51Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [February 20, 2023, 6:47am UTC](https://discuss.ray.io/t/start-cluster-with-multiple-head-node/9425/1 "2023-02-20T06:47:51Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- High: It blocks me to complete my task.

I am running ray cluster with 100 around nodes. In this environment, there is high probability that while cluster is running, due to some issue head node is down. With current scenario, if head node is down, complete cluster is useless. I am using cluster.yaml file for cluster creation and all the nodes are in local network.

**To resolve this issue, I was thinking to have two/three (multiple) head node, where if any one is down, another can handle incoming job requests. Is there anyway which can help me to approach this?**

OR

**Is there anyway, If head node down, all worked node should continue their job, and meanwhile we can attach head node again to cluster?**

---

<div class="post-metadata">

**Author:** ![tarjintor](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/tarjintor/32/3552_2.png) [@tarjintor](https://discuss.ray.io/u/tarjintor)\
**Post date:** [February 20, 2023, 9:15am UTC](https://discuss.ray.io/t/start-cluster-with-multiple-head-node/9425/2 "2023-02-20T09:15:47Z")

</div>

I think this is very important feature,the HA of head node seems not finished now.  
I did some search and find a git issue : [[RFC] GCS High availability · Issue #20498 · ray-project/ray (github.com)](https://github.com/ray-project/ray/issues/20498)  
maybe this still need a lot of work?

---

<div class="post-metadata">

**Author:** ![yic](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/yic/32/437_2.png) [@yic](https://discuss.ray.io/u/yic)\
**Post date:** [February 21, 2023, 6:03pm UTC](https://discuss.ray.io/t/start-cluster-with-multiple-head-node/9425/3 "2023-02-21T18:03:11Z")

</div>

GCS FT is supported in KubeRay. You can also just bring a redis cluster and if the head node is down, just restart it.

[https://docs.ray.io/en/latest/serve/production-guide/fault-tolerance.html#step-2-add-redis-info-to-rayservice](https://docs.ray.io/en/latest/serve/production-guide/fault-tolerance.html#step-2-add-redis-info-to-rayservice)

Btw, do you mind showing the error logs why the head node crashed?

The HA GCS is more complicated than FT GCS which will be a long term project.

---

<div class="post-metadata">

**Author:** ![shyampatel](https://avatars.discourse-cdn.com/v4/letter/s/cab0a1/32.png) [@shyampatel](https://discuss.ray.io/u/shyampatel)\
**Post date:** [February 22, 2023, 6:07am UTC](https://discuss.ray.io/t/start-cluster-with-multiple-head-node/9425/4 "2023-02-22T06:07:24Z")

</div>

@yic Thanks for your reply.

> [@yic](#):
>
> do you mind showing the error logs why the head node crashed?

It’s not about head node crashing. For our case, sometimes, system running head node itself was down due to environmental issues.

---

<div class="post-metadata">

**Author:** ![yic](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/yic/32/437_2.png) [@yic](https://discuss.ray.io/u/yic)\
**Post date:** [February 22, 2023, 8:57pm UTC](https://discuss.ray.io/t/start-cluster-with-multiple-head-node/9425/5 "2023-02-22T20:57:49Z")

</div>

Got it. I think FT GCS is what you need then.
