# Intial setup for ray on a HPC

**URL:** <https://discuss.ray.io/t/intial-setup-for-ray-on-a-hpc/13465>\
**Category:** Ray Serve\
**Created:** [January 18, 2024, 8:01pm UTC](https://discuss.ray.io/t/intial-setup-for-ray-on-a-hpc/13465 "2024-01-18T20:01:08Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![JayTea](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jaytea/32/5627_2.png) [@JayTea](https://discuss.ray.io/u/JayTea)\
**Post date:** [January 18, 2024, 8:01pm UTC](https://discuss.ray.io/t/intial-setup-for-ray-on-a-hpc/13465/1 "2024-01-18T20:01:08Z")

</div>

Hi Y’all,

Sorry a head of time, if I am breaking some forum rules.  
I am trying to test out the capabilities of RAY on a bar metal HPC using RHEL 8~9.

Ultimate goal is to get a ray cluster setup and have a head node running on one node and a worker node that runs a VLLM backed model on another node.

However before doing that, I am having trouble setting up the initial ray start.

1. I ran `ray start --head` and successfully started the Ray head node
2. I ran `ray start --address=HEAD_NODE_IP:6379 --num-cpus=32 --num-gpus=1`

However, shortly after it prints out Local node IP: ip\_address, and after sometime (~2 mins) it prints out an error message saying **RuntimeError( RuntimeError: Failed to connect to GCS.**

Somethings to note,

1. I am currently allocating resources using slurm as this is going to be a proof of concept to push it more system wide.
2. Ray is installed using conda for both nodes
3. I am able to ping the head node from the worker node and vice versa.
4. Running ray status on the head node works

---

<div class="post-metadata">

**Author:** ![max\_ronda](https://avatars.discourse-cdn.com/v4/letter/m/f0a364/32.png) [@max\_ronda](https://discuss.ray.io/u/max_ronda)\
**Post date:** [January 18, 2024, 10:45pm UTC](https://discuss.ray.io/t/intial-setup-for-ray-on-a-hpc/13465/2 "2024-01-18T22:45:11Z")

</div>

Hey @JayTea , I think you might be missing a TMPDIR . Try providing --temp-dir and also --redis-password

e.g.

On head node, open a terminal and type:

`ray start --head --port=6379 --num-cpus=<total-cpus> --redis-password=2173274697--temp-dir=/data1/gridtmp/ray_temp_dir --dashboard-host '0.0.0.0' --block`

On worker nodes, open a terminal and type

`ray start --address=HEAD_NODE_IP_ADDRESS:6379 --num-cpus=32 --redis-password=2173274697 --temp-dir=/data1/gridtmp/ray_temp_dir --block`

and connect to it via:

```auto
ray.init(
    address=f"ray://{head_node_ip_address}:10001",
    log_to_driver=False,
    ignore_reinit_error=True
)

```

See if that works.

---

<div class="post-metadata">

**Author:** ![JayTea](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jaytea/32/5627_2.png) [@JayTea](https://discuss.ray.io/u/JayTea)\
**Post date:** [January 18, 2024, 11:26pm UTC](https://discuss.ray.io/t/intial-setup-for-ray-on-a-hpc/13465/3 "2024-01-18T23:26:43Z")

</div>

Hi @max_ronda,

Thanks for reaching out.

Sadly it did not work.

I got the same error **RuntimeError: Failed to connect to GCS.** when executing the worker node command, and therefore when I try to run the ray.init I got **ConnectionError: ray client connection timeout**

Also when running `ray start --address=HEAD_NODE_IP_ADDRESS:6379 --num-cpus=32 --redis-password=2173274697 --temp-dir=/data1/gridtmp/ray_temp_dir --block` I got the warning

> `--temp-dir=/data1/gridtmp/ray_temp_dir` option will be ignored. `--head` is a required flag to use `--temp-dir`. temp\_dir is only configurable from a head node. All the worker nodes will use the same temp\_dir as a head node.

**Update**  
Interestingly,  
When I run the head node and worker node on two different slurm instances that are on the same node, the commands run correctly and I can see when I run `ray status` that a node was added to the head node.

---

<div class="post-metadata">

**Author:** ![JayTea](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jaytea/32/5627_2.png) [@JayTea](https://discuss.ray.io/u/JayTea)\
**Post date:** [January 19, 2024, 5:34pm UTC](https://discuss.ray.io/t/intial-setup-for-ray-on-a-hpc/13465/4 "2024-01-19T17:34:12Z")

</div>

Found out it was actually a firewall issue and port 6379 was not open.  
To test I used `nmap head_node_ip -p 6369` within the worker node and saw it was not open.

---

<div class="post-metadata">

**Author:** ![Avi](https://avatars.discourse-cdn.com/v4/letter/a/b9bd4f/32.png) [@Avi](https://discuss.ray.io/u/Avi)\
**Post date:** [January 20, 2024, 5:57pm UTC](https://discuss.ray.io/t/intial-setup-for-ray-on-a-hpc/13465/5 "2024-01-20T17:57:55Z")

</div>

**Indeed, most of networks might not updated with the CIDR range with 6379.**

**Alternatively you can refer the following the recent ray job submission client approach this also helps you to submit the jobs:**

Sample code

import ray  
from ray.job\_submission import JobSubmissionClient  
import time

# Ray cluster information

ray\_head\_ip = “kuberay-head-svc.kuberay.svc.cluster.local”  
ray\_head\_port = 8265  
ray\_address = f"http://{ray\_head\_ip}:{ray\_head\_port}"

while True:  
# Submit Ray job using JobSubmissionClient  
client = JobSubmissionClient(ray\_address)  
job\_id = client.submit\_job(  
entrypoint=“python run.py”,  
runtime\_env={  
“working\_dir”: “./”  
},  
entrypoint\_num\_cpus = 1,  
)

```
print(client. __dict__ )
print(f"Ray job submitted with job_id: {job_id}")
# Wait for a while to let the jobs run
time.sleep(10)

job_status = client.get_job_status(job_id)
get_job_logs = client.get_job_logs(job_id)
get_job_info = client.get_job_info(job_id)
async for lines in client.tail_job_logs(job_id):
    print(lines, end="")

# Shutdown Ray
ray.shutdown()

```
