# Start Ray cluster with error but working

**URL:** <https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884>\
**Category:** Ray Clusters\
**Created:** [April 22, 2022, 1:41am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884 "2022-04-22T01:41:22Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![xyzyx](https://avatars.discourse-cdn.com/v4/letter/x/bcef8e/32.png) [@xyzyx](https://discuss.ray.io/u/xyzyx)\
**Post date:** [April 22, 2022, 1:41am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/1 "2022-04-22T01:41:22Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- Medium: It contributes to significant difficulty to complete my task, but I can work around it.  
I start ray cluster using a slurm script. There are some errors when I start cluster but my program can run. The error output in one node shows below:

```auto
e[2me[33m(raylet, ip=10.6.12.47)e[0m Traceback (most recent call last):
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/ray/dashboard/agent.py", line 391, in <module>
e[2me[33m(raylet, ip=10.6.12.47)e[0m loop.run_until_complete(agent.run())
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/software/apps/AI/apps/DeepLearning/PyTorch/cccp/pytorch_1.8-rocm_4.0.1-fastmoe/lib/python3.6/asyncio/base_events.py", line 484, in run_until_complete
e[2me[33m(raylet, ip=10.6.12.47)e[0m return future.result()
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/ray/dashboard/agent.py", line 178, in run
e[2me[33m(raylet, ip=10.6.12.47)e[0m modules = self._load_modules()
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/ray/dashboard/agent.py", line 120, in _load_modules
e[2me[33m(raylet, ip=10.6.12.47)e[0m c = cls(self)
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/ray/dashboard/modules/reporter/reporter_agent.py", line 163, in __init__
e[2me[33m(raylet, ip=10.6.12.47)e[0m dashboard_agent.metrics_export_port)
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/ray/_private/metrics_agent.py", line 79, in __init__
e[2me[33m(raylet, ip=10.6.12.47)e[0m address=metrics_export_address)))
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/ray/_private/prometheus_exporter.py", line 333, in new_stats_exporter
e[2me[33m(raylet, ip=10.6.12.47)e[0m options=option, gatherer=option.registry, collector=collector)
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/ray/_private/prometheus_exporter.py", line 265, in __init__
e[2me[33m(raylet, ip=10.6.12.47)e[0m self.serve_http()
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/ray/_private/prometheus_exporter.py", line 320, in serve_http
e[2me[33m(raylet, ip=10.6.12.47)e[0m port=self.options.port, addr=str(self.options.address))
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/prometheus_client/exposition.py", line 168, in start_wsgi_server
e[2me[33m(raylet, ip=10.6.12.47)e[0m TmpServer.address_family, addr = _get_best_family(addr, port)
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/prometheus_client/exposition.py", line 157, in _get_best_family
e[2me[33m(raylet, ip=10.6.12.47)e[0m infos = socket.getaddrinfo(address, port)
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/software/apps/AI/apps/DeepLearning/PyTorch/cccp/pytorch_1.8-rocm_4.0.1-fastmoe/lib/python3.6/socket.py", line 745, in getaddrinfo
e[2me[33m(raylet, ip=10.6.12.47)e[0m for res in _socket.getaddrinfo(host, port, family, type, proto, flags):
e[2me[33m(raylet, ip=10.6.12.47)e[0m socket.gaierror: [Errno -2] Name or service not known
e[2me[33m(raylet, ip=10.6.12.47)e[0m 
e[2me[33m(raylet, ip=10.6.12.47)e[0m During handling of the above exception, another exception occurred:
e[2me[33m(raylet, ip=10.6.12.47)e[0m 
e[2me[33m(raylet, ip=10.6.12.47)e[0m Traceback (most recent call last):
e[2me[33m(raylet, ip=10.6.12.47)e[0m File "/public/home/lifei/xinzk/envs/ray_base/site-packages/ray/dashboard/agent.py", line 407, in <module>
e[2me[33m(raylet, ip=10.6.12.47)e[0m gcs_publisher = GcsPublisher(args.gcs_address)
e[2me[33m(raylet, ip=10.6.12.47)e[0m TypeError: __init__ () takes 1 positional argument but 2 were given

```

My slurm script is following:

```auto
#!/bin/bash
#SBATCH -p normal
#SBATCH --gres=dcu:4
#SBATCH --exclusive

module unload compiler/rocm/2.9
module load apps/ray/hpcx-2.4.1-gcc-7.3.1-rocm4.0.1

redis_password=$(uuidgen)
export redis_password

nodes=$(scontrol show hostnames $SLURM_JOB_NODELIST) # Getting the node names
nodes_array=( $nodes )

node_1=${nodes_array[0]} 
ip=$(srun --nodes=1 --ntasks=1 -w $node_1 hostname --ip-address) # making redis-address
port=6379
ip_head=$ip:$port
export ip_head
echo "IP Head: $ip_head"

echo "STARTING HEAD at $node_1"
srun --nodes=1 --ntasks=1 -w $node_1 start-head.sh $ip $redis_password &
sleep 30
worker_num=$(($SLURM_JOB_NUM_NODES - 1)) #number of nodes other than the head node
for (( i=1; i <= ${worker_num}; i++ ))
do
  node_i=${nodes_array[$i]}
  echo "STARTING WORKER $i at $node_i"
  srun --nodes=1 --ntasks=1 -w $node_i start-worker.sh $ip_head $redis_password &
  sleep 5
done

which python3
python3 -u ps.py -c $1 -b 16

```

start-head.sh:

```auto
#!/bin/bash
echo "starting ray head node"
# Launch the head node
ray start --head --node-ip-address=$1 --port=6379 --redis-password=$2 --num-gpus=4
sleep infinity

```

start-worker.sh

```auto
#!/bin/bash
echo "starting ray worker node"
ray start --address $1 --redis-password=$2 --num-gpus=4
sleep infinity

```

Is there something wrong when I run the script?

---

<div class="post-metadata">

**Author:** ![Ameer\_Haj\_Ali](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ameer_haj_ali/32/279_2.png) [@Ameer\_Haj\_Ali](https://discuss.ray.io/u/Ameer_Haj_Ali)\
**Post date:** [April 25, 2022, 11:45am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/2 "2022-04-25T11:45:52Z")

</div>

@Alex can you please help?

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [April 25, 2022, 2:27pm UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/3 "2022-04-25T14:27:39Z")

</div>

@xyzyx it looks like your script is trying to start a head node with an external redis server. Is that intentional? (If so, how are you verifying redis is healthy?)

If not, you may want your head start command to not include mentions of redis/addresses/ports

```auto
ray start --head --node-ip-address=$1 --num-gpus=4

```

---

<div class="post-metadata">

**Author:** ![xyzyx](https://avatars.discourse-cdn.com/v4/letter/x/bcef8e/32.png) [@xyzyx](https://discuss.ray.io/u/xyzyx)\
**Post date:** [April 26, 2022, 2:18am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/4 "2022-04-26T02:18:11Z")

</div>

Thanks! 😃  
I actually do not want to start with an external redis server. So I do not need to specify `--redis-password`?

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [April 27, 2022, 3:00pm UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/5 "2022-04-27T15:00:30Z")

</div>

yep that’s correct. in fact, Ray no longer has a hard dependency on redis and won’t use redis by default now.

---

<div class="post-metadata">

**Author:** ![xyzyx](https://avatars.discourse-cdn.com/v4/letter/x/bcef8e/32.png) [@xyzyx](https://discuss.ray.io/u/xyzyx)\
**Post date:** [April 28, 2022, 4:38am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/6 "2022-04-28T04:38:11Z")

</div>

I’m not including mention of Redis but the error is still here. The command I run is `ray start --block --address=$ip_head`

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [April 29, 2022, 3:45pm UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/7 "2022-04-29T15:45:11Z")

</div>

Do you mind verifying the version of Ray that you’re using (on both the head and worker nodes?)

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [April 29, 2022, 3:46pm UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/8 "2022-04-29T15:46:07Z")

</div>

heads up @mwtian (who knows more than me)

---

<div class="post-metadata">

**Author:** ![GoingMyWay](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/goingmyway/32/1102_2.png) [@GoingMyWay](https://discuss.ray.io/u/GoingMyWay)\
**Post date:** [June 26, 2022, 8:42am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/9 "2022-06-26T08:42:27Z")

</div>

Same issue on k8s cluster. Have you solved this problem?

---

<div class="post-metadata">

**Author:** ![ckw017](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ckw017/32/1683_2.png) [@ckw017](https://discuss.ray.io/u/ckw017)\
**Post date:** [June 27, 2022, 9:46pm UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/10 "2022-06-27T21:46:05Z")

</div>

What Ray version are you on? And can you share more details about your setup process

---

<div class="post-metadata">

**Author:** ![Dmitri](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/dmitri/32/657_2.png) [@Dmitri](https://discuss.ray.io/u/Dmitri)\
**Post date:** [June 28, 2022, 12:47am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/11 "2022-06-28T00:47:45Z")

</div>

@GoingMyWay please do provide a detailed reproduction on K8s if possible.

---

<div class="post-metadata">

**Author:** ![xyzyx](https://avatars.discourse-cdn.com/v4/letter/x/bcef8e/32.png) [@xyzyx](https://discuss.ray.io/u/xyzyx)\
**Post date:** [June 28, 2022, 12:55am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/12 "2022-06-28T00:55:10Z")

</div>

I used 1.11.0. 😃  
I use the ray on a slurm cluster and I startup using a modified script from [here](https://docs.ray.io/en/latest/cluster/examples/slurm-template.html).

---

<div class="post-metadata">

**Author:** ![Dmitri](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/dmitri/32/657_2.png) [@Dmitri](https://discuss.ray.io/u/Dmitri)\
**Post date:** [June 28, 2022, 1:43am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/13 "2022-06-28T01:43:55Z")

</div>

Re: slurm @tupui might be able to help.

---

<div class="post-metadata">

**Author:** ![tupui](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/tupui/32/2509_2.png) [@tupui](https://discuss.ray.io/u/tupui)\
**Post date:** [June 28, 2022, 1:44pm UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/14 "2022-06-28T13:44:06Z")

</div>

I did not observe such issue on my cluster. @xyzyx you are saying that your program is running, but since you have an exception, is it running in parallel on all nodes or just on the head node? Also could you try using the latest version of ray?

---

<div class="post-metadata">

**Author:** ![xyzyx](https://avatars.discourse-cdn.com/v4/letter/x/bcef8e/32.png) [@xyzyx](https://discuss.ray.io/u/xyzyx)\
**Post date:** [June 29, 2022, 12:49am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/15 "2022-06-29T00:49:51Z")

</div>

My program is running fine but outputs these error messages. It is running in parallel on all nodes.  
I will try the latest version of ray later.

---

<div class="post-metadata">

**Author:** ![GoingMyWay](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/goingmyway/32/1102_2.png) [@GoingMyWay](https://discuss.ray.io/u/GoingMyWay)\
**Post date:** [July 4, 2022, 10:47am UTC](https://discuss.ray.io/t/start-ray-cluster-with-error-but-working/5884/16 "2022-07-04T10:47:10Z")

</div>

Hi @Dmitri, please see this comment: [Ray k8s cluster, cannot run new task when previous task failed - #6 by GoingMyWay](https://discuss.ray.io/t/ray-k8s-cluster-cannot-run-new-task-when-previous-task-failed/6459/6)
