# When I was doing DDPPO training, cluster initialization failed

**URL:** <https://discuss.ray.io/t/when-i-was-doing-ddppo-training-cluster-initialization-failed/9471>\
**Category:** RLlib\
**Created:** [February 23, 2023, 8:08am UTC](https://discuss.ray.io/t/when-i-was-doing-ddppo-training-cluster-initialization-failed/9471 "2023-02-23T08:08:53Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![zhangzhang](https://avatars.discourse-cdn.com/v4/letter/z/b77776/32.png) [@zhangzhang](https://discuss.ray.io/u/zhangzhang)\
**Post date:** [February 23, 2023, 8:08am UTC](https://discuss.ray.io/t/when-i-was-doing-ddppo-training-cluster-initialization-failed/9471/1 "2023-02-23T08:08:53Z")

</div>

ERROR worker.py:756 – Exception raised in creation task: The actor died because of an error raised in its creation task, ray::DDPPO. **init** () (pid=126499, ip=10.19.xxx.xx, repr=DDPPO)  
(DDPPO pid=126499) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/algorithms/ddppo/ddppo.py”, line 179, in **init**  
(DDPPO pid=126499) super(). **init** (  
(DDPPO pid=126499) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/algorithms/algorithm.py”, line 308, in **init**  
(DDPPO pid=126499) super(). **init** (config=config, logger\_creator=logger\_creator, \*\*kwargs)  
(DDPPO pid=126499) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/tune/trainable/trainable.py”, line 157, in **init**  
(DDPPO pid=126499) self.setup(copy.deepcopy(self.config))  
(DDPPO pid=126499) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/algorithms/ddppo/ddppo.py”, line 264, in setup  
(DDPPO pid=126499) ray.get(  
(DDPPO pid=126499) ray.exceptions.RayTaskError(RuntimeError): ray::RolloutWorker.setup\_torch\_data\_parallel() (pid=3420304, ip=10.20.xx.xx, repr=\<ray.rllib.evaluation.rollout\_worker.RolloutWorker object at 0x7f17b66a6640\>)  
(DDPPO pid=126499) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/evaluation/rollout\_worker.py”, line 1680, in setup\_torch\_data\_parallel  
(DDPPO pid=126499) torch.distributed.init\_process\_group(  
(DDPPO pid=126499) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/torch/distributed/distributed\_c10d.py”, line 602, in init\_process\_group  
(DDPPO pid=126499) default\_pg = \_new\_process\_group\_helper(  
(DDPPO pid=126499) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/torch/distributed/distributed\_c10d.py”, line 703, in \_new\_process\_group\_helper  
(DDPPO pid=126499) pg = ProcessGroupGloo(prefix\_store, rank, world\_size, timeout=timeout)  
(DDPPO pid=126499) RuntimeError: [/opt/conda/conda-bld/pytorch\_1659484683044/work/third\_party/gloo/gloo/transport/tcp/pair.cc:799] connect [127.0.1.1]:4505: Connection refused

My config:  
“num\_workers”: 4, #2,  
“num\_envs\_per\_worker”: 1, # 5  
“num\_cpus\_per\_worker”: 2, # 16  
“framework”: “torch”,  
“no\_done\_at\_end”: True,  
“sample\_async”: False,  
“placement\_strategy”: “SPREAD”,  
“keep\_local\_weights\_in\_sync”:True,  
“num\_gpus\_per\_worker”: 0.4, # 0.5  
“rollout\_fragment\_length”: 200,  
“torch\_distributed\_backend”:“gloo”,

I trained with the docker environment on both servers.

when I use nccl

(DDPPO pid=128397) 2023-02-23 08:23:22,790 ERROR algorithm.py:2173 – Error in training or evaluation attempt! Trying to recover.  
(DDPPO pid=128397) Traceback (most recent call last):  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/algorithms/algorithm.py”, line 2373, in \_run\_one\_training\_iteration  
(DDPPO pid=128397) results = self.training\_step()  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/util/tracing/tracing\_helper.py”, line 466, in \_resume\_span  
(DDPPO pid=128397) return method(self, \*\_args, \*\*\_kwargs)  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/algorithms/ddppo/ddppo.py”, line 290, in training\_step  
(DDPPO pid=128397) sample\_and\_update\_results = self.\_ddppo\_worker\_manager.get\_ready()  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/execution/parallel\_requests.py”, line 173, in get\_ready  
(DDPPO pid=128397) objs = ray.get(ready\_requests)  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/\_private/client\_mode\_hook.py”, line 105, in wrapper  
(DDPPO pid=128397) return func(\*args, \*\*kwargs)  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/\_private/worker.py”, line 2275, in get  
(DDPPO pid=128397) raise value.as\_instanceof\_cause()  
(DDPPO pid=128397) ray.exceptions.RayTaskError(ValueError): ray::RolloutWorker.apply() (pid=128448, ip=10.19.xxx.xx, repr=\<ray.rllib.evaluation.rollout\_worker.RolloutWorker object at 0x7f7ef1967790\>)  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/torch/distributed/distributed\_c10d.py”, line 1320, in all\_reduce  
(DDPPO pid=128397) work = default\_pg.allreduce([tensor], opts)  
(DDPPO pid=128397) RuntimeError: NCCL error in: /opt/conda/conda-bld/pytorch\_1659484683044/work/torch/csrc/distributed/c10d/ProcessGroupNCCL.cpp:1191, unhandled system error, NCCL version 2.10.3  
(DDPPO pid=128397) ncclSystemError: System call (e.g. socket, malloc) or external library call failed or device error. It can be also caused by unexpected exit of a remote peer, you can check NCCL warnings for failure reason and see if there is connection closure by a peer.  
(DDPPO pid=128397)  
(DDPPO pid=128397) The above exception was the direct cause of the following exception:  
(DDPPO pid=128397)  
(DDPPO pid=128397) ray::RolloutWorker.apply() (pid=128448, ip=10.19.xxx.xx, repr=\<ray.rllib.evaluation.rollout\_worker.RolloutWorker object at 0x7f7ef1967790\>)  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/evaluation/rollout\_worker.py”, line 1669, in apply  
(DDPPO pid=128397) return func(self, \*args, \*\*kwargs)  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/algorithms/ddppo/ddppo.py”, line 351, in \_sample\_and\_train\_torch\_distributed  
(DDPPO pid=128397) info = do\_minibatch\_sgd(  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/utils/sgd.py”, line 129, in do\_minibatch\_sgd  
(DDPPO pid=128397) local\_worker.learn\_on\_batch(  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/evaluation/rollout\_worker.py”, line 919, in learn\_on\_batch  
(DDPPO pid=128397) info\_out[pid] = policy.learn\_on\_batch(batch)  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/utils/threading.py”, line 24, in wrapper  
(DDPPO pid=128397) return func(self, \*a, \*\*k)  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/policy/torch\_policy\_v2.py”, line 606, in learn\_on\_batch  
(DDPPO pid=128397) grads, fetches = self.compute\_gradients(postprocessed\_batch)  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/utils/threading.py”, line 24, in wrapper  
(DDPPO pid=128397) return func(self, \*a, \*\*k)  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/policy/torch\_policy\_v2.py”, line 789, in compute\_gradients  
(DDPPO pid=128397) tower\_outputs = self.\_multi\_gpu\_parallel\_grad\_calc([postprocessed\_batch])  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/policy/torch\_policy\_v2.py”, line 1179, in \_multi\_gpu\_parallel\_grad\_calc  
(DDPPO pid=128397) raise last\_result[0] from last\_result[1]  
(DDPPO pid=128397) ValueError: NCCL error in: /opt/conda/conda-bld/pytorch\_1659484683044/work/torch/csrc/distributed/c10d/ProcessGroupNCCL.cpp:1191, unhandled system error, NCCL version 2.10.3  
(DDPPO pid=128397) ncclSystemError: System call (e.g. socket, malloc) or external library call failed or device error. It can be also caused by unexpected exit of a remote peer, you can check NCCL warnings for failure reason and see if there is connection closure by a peer.  
(DDPPO pid=128397) tracebackTraceback (most recent call last):  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/ray/rllib/policy/torch\_policy\_v2.py”, line 1137, in \_worker  
(DDPPO pid=128397) torch.distributed.all\_reduce(  
(DDPPO pid=128397) File “/opt/conda/envs/rl\_decision/lib/python3.8/site-packages/torch/distributed/distributed\_c10d.py”, line 1320, in all\_reduce  
(DDPPO pid=128397) work = default\_pg.allreduce([tensor], opts)  
(DDPPO pid=128397) RuntimeError: NCCL error in: /opt/conda/conda-bld/pytorch\_1659484683044/work/torch/csrc/distributed/c10d/ProcessGroupNCCL.cpp:1191, unhandled system error, NCCL version 2.10.3  
(DDPPO pid=128397) ncclSystemError: System call (e.g. socket, malloc) or external library call failed or device error. It can be also caused by unexpected exit of a remote peer, you can check NCCL warnings for failure reason and see if there is connection closure by a peer.  
(DDPPO pid=128397)  
(DDPPO pid=128397) In tower 0 on device cuda:0

---

<div class="post-metadata">

**Author:** ![arturn](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/arturn/32/2096_2.png) [@arturn](https://discuss.ray.io/u/arturn)\
**Post date:** [March 22, 2023, 1:53am UTC](https://discuss.ray.io/t/when-i-was-doing-ddppo-training-cluster-initialization-failed/9471/2 "2023-03-22T01:53:40Z")

</div>

Hi @zhangzhang,

Thanks for raising this. Can you report this issue to the torch-distributed folks?  
We don’t “mess” with torch ddp, so this is unlikely to be an RLlib related error.  
Have you found anything so far?
