# Ray on slurm - Problems with initialization

**URL:** <https://discuss.ray.io/t/ray-on-slurm-problems-with-initialization/6361>\
**Category:** Ray Clusters\
**Created:** [June 1, 2022, 3:45pm UTC](https://discuss.ray.io/t/ray-on-slurm-problems-with-initialization/6361 "2022-06-01T15:45:49Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Pierre\_houdouin](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/pierre_houdouin/32/2657_2.png) [@Pierre\_houdouin](https://discuss.ray.io/u/Pierre_houdouin)\
**Post date:** [June 1, 2022, 3:45pm UTC](https://discuss.ray.io/t/ray-on-slurm-problems-with-initialization/6361/1 "2022-06-01T15:45:49Z")

</div>

Hello everyone,

I write this post because since I use slurm, I have not been able to use ray correctly.  
Whenever I use the commands :

- ray.init
- trainer = A3CTrainer(env = “my\_env”) (I have registered my env on tune)  
, the program crashes with the following message :

**core\_worker.cc:137: Failed to register worker 01000000ffffffffffffffffffffffffffffffffffffffffffffffff to Raylet. IOError: [RayletClient] Unable to register worker with raylet. No such file or directory**

The program works fine on my computer, the problem appeared with the use of Slurm. I only ask slurm for one gpu.

Thank you for reading me and maybe answering.  
Have a great day

---

<div class="post-metadata">

**Author:** ![cade](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/cade/32/3837_2.png) [@cade](https://discuss.ray.io/u/cade)\
**Post date:** [June 17, 2022, 10:19pm UTC](https://discuss.ray.io/t/ray-on-slurm-problems-with-initialization/6361/2 "2022-06-17T22:19:03Z")

</div>

I see further details on the same question on [SO](https://stackoverflow.com/questions/72464756/ray-on-slurm-problems-with-initialization), copying here for visibility:

```nohighlight
import ray
from ray.rllib.agents.a3c import A3CTrainer
import tensorflow as tf
from MM1c_queue_env import my_env #my_env is already registered in tune

ray.shutdown()
ray.init(ignore_reinit_error=True)
trainer = A3CTrainer(env = "my_env")

print("success")

```

To launch the program with slurm, I use the following program :

```auto
#!/bin/bash

#SBATCH --job-name=rl_for_insensitive_policies
#SBATCH --time=0:05:00 
#SBATCH --ntasks=1
#SBATCH --gres=gpu:1
#SBATCH --partition=gpu

module load anaconda3/2020.02/gcc-9.2.0
python test.py

```

---

<div class="post-metadata">

**Author:** ![cade](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/cade/32/3837_2.png) [@cade](https://discuss.ray.io/u/cade)\
**Post date:** [June 17, 2022, 10:20pm UTC](https://discuss.ray.io/t/ray-on-slurm-problems-with-initialization/6361/3 "2022-06-17T22:20:23Z")

</div>

@Pierre_houdouin can you share how you’re starting the Ray worker nodes? For example, are you following [Starting the Ray worker nodes](https://docs.ray.io/en/latest/cluster/slurm.html#id6) in the docs?

---

<div class="post-metadata">

**Author:** ![mgerstgrasser](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mgerstgrasser/32/2728_2.png) [@mgerstgrasser](https://discuss.ray.io/u/mgerstgrasser)\
**Post date:** [September 7, 2022, 9:32pm UTC](https://discuss.ray.io/t/ray-on-slurm-problems-with-initialization/6361/4 "2022-09-07T21:32:26Z")

</div>

Just to share that I had the same problem, and for me it seems to be resolved by setting `num_cpus=1` (or whatever number of cores I request from slurm) in `ray.init` solved the problem, as per one of the answers on SO.

---

<div class="post-metadata">

**Author:** ![mgerstgrasser](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mgerstgrasser/32/2728_2.png) [@mgerstgrasser](https://discuss.ray.io/u/mgerstgrasser)\
**Post date:** [October 17, 2022, 4:31pm UTC](https://discuss.ray.io/t/ray-on-slurm-problems-with-initialization/6361/5 "2022-10-17T16:31:29Z")

</div>

@cade Actually, it turns out this still happens for me! But only if I request only a single CPU core from SLURM (i.e. set `-n 1` in `sbatch`). Two or more cores are fine, most of the time even if Ray then uses more worker processes than I requested cores. This happens even if I run ray in local mode! I am told by our cluster support that `-n 1` does nothing other than pin all the processes to a single physical core.

One thing I should add is that I am not running a whole Ray cluster on top of slurm, I just want to run one single rllib experiment per slurm job (but multiple such slurm jobs in parallel).

---

<div class="post-metadata">

**Author:** ![cupe](https://avatars.discourse-cdn.com/v4/letter/c/4af34b/32.png) [@cupe](https://discuss.ray.io/u/cupe)\
**Post date:** [December 25, 2022, 6:29am UTC](https://discuss.ray.io/t/ray-on-slurm-problems-with-initialization/6361/6 "2022-12-25T06:29:31Z")

</div>

Just a note: this _may_ be due to a large number of OpenBLAS threads when using a large number of CPU cores. See comment with workaround here: [[\<Ray component: Core] Failed to register worker . Slurm - srun - · Issue #30012 · ray-project/ray · GitHub](https://github.com/ray-project/ray/issues/30012#issuecomment-1364633366)

---

<div class="post-metadata">

**Author:** ![cade](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/cade/32/3837_2.png) [@cade](https://discuss.ray.io/u/cade)\
**Post date:** [December 29, 2022, 1:51am UTC](https://discuss.ray.io/t/ray-on-slurm-problems-with-initialization/6361/7 "2022-12-29T01:51:17Z")

</div>

Could you share the output of `ulimit -a`? I wonder if the fd limit is too low
