# \[Serve\] The \`ray start --head --node-ip-address ip\` is not working correctly in Docker. And it's not clear which ports to open

**URL:** <https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214>\
**Category:** Ray Serve\
**Created:** [December 20, 2023, 2:08pm UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214 "2023-12-20T14:08:01Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![psydok](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/psydok/32/3518_2.png) [@psydok](https://discuss.ray.io/u/psydok)\
**Post date:** [December 20, 2023, 2:08pm UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214/1 "2023-12-20T14:08:01Z")

</div>

I’m trying to connect nodes deployed via docker to the master node. I am having a number of problems, which locally I was able to solve by setting `network_mode: host` to containers. But on the servers there is a firewall running and now I don’t understand how I can solve my problem. I can’t find any logs as to why the node connection was broken about 10 seconds after connection. Also, I can’t get the main node up with `--node-ip-address x.x.x.x.x` specifying. I get the error “RuntimeError: Failed to start GCS. Last 0 lines of error files:”

`x.x.x.x` - external ip

```docker
# docker-compose

version: "3.7"

services:
  node:
    image: <my-image>
    env_file:
      - .env
    environments:
      - RAY_num_heartbeats_timeout=300
      - RAY_CONFIG_CREATING_NODE=--head --metrics-export-port 9088 --dashboard-agent-listen-port 8266 --dashboard-agent-grpc-port 9266 --min-worker-port 10002 --max-worker-port 10010 --dashboard-host=0.0.0.0 --port 6378 --redis-shard-ports 6099 --dashboard-grpc-port 9265 --num-cpus=5 --node-ip-address x.x.x.x
    volumes:
      - ./data:/data
    runtime: nvidia
    restart: unless-stopped
    privileged: true
    # network_mode: "host"
    ports:
      - 8265:8265
      - 9122:9122
      - 6378:6378
      - 8099:8099
      - 9099:9099
      - 9265:9265
      - 6099:6099
      - 9266:9266
      - 8266:8266
      - 9088:9088
      - 10001-10010:10001-10010

```

Then with the same docker-compose I connect another node by changing only `RAY_CONFIG_CREATING_NODE="--address x.x.x.x:6378 --dashboard-grpc-port 9265 --metrics-export-port 9088 --dashboard-agent-listen-port 8266 --dashboard-agent-grpc-port 9266 --min-worker-port 10002 --max-worker-port 10010 --num-cpus=5 --node-ip-address y.y.y.y"`

After about 15 seconds, the connection between the nodes breaks.  
`|The node with node id: <id> and address: y.y.y.y and node name: y.y.y.y has been marked dead because the detector has missed too many heartbeats from it. This can happen when a |(1) raylet crashes unexpectedly (OOM, preempted node, etc.) | |---|---| |2|(2) raylet has lagging heartbeats due to slow network or busy workload.|`

```bash
docker compose exec node serve start --proxy-location EveryNode \
        --http-host 0.0.0.0 --http-port 8099 --grpc-port 9099 \
        --grpc-servicer-functions dto.test_pb2_grpc.add_TestServicer_to_server

2023-12-20 19:14:33,573 INFO worker.py:1489 -- Connecting to existing Ray cluster at address: x.x.x.x:6378...
2023-12-20 19:14:33,591 INFO worker.py:1664 -- Connected to Ray cluster. View the dashboard at http://192.168.208.2:8265 
[2023-12-20 19:14:42,602 E 188 264] core_worker_process.cc:216: Failed to get the system config from raylet because it is dead. Worker will terminate. Status: GrpcUnavailable: RPC Error message: failed to connect to all addresses; last error: UNKNOWN: ipv4:y.y.y.y:40985: Failed to connect to remote host: Connection refused; RPC Error details: .Please see `raylet.out` for more details.

```

I’m guessing it’s a port problem. Which ports should be opened? How do I remove the randomness and configure the exact port values?

---

<div class="post-metadata">

**Author:** ![Sihan\_Wang](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sihan_wang/32/2570_2.png) [@Sihan\_Wang](https://discuss.ray.io/u/Sihan_Wang)\
**Post date:** [December 28, 2023, 9:26pm UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214/2 "2023-12-28T21:26:18Z")

</div>

Hi @psydok ,

Can you give a try to expose the port 6379?

Btw you can specify the port number in the start up cli by having “–port xxx”. (ref: [Cluster Management CLI — Ray 2.9.0](https://docs.ray.io/en/latest/cluster/cli.html#ray-start))

---

<div class="post-metadata">

**Author:** ![psydok](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/psydok/32/3518_2.png) [@psydok](https://discuss.ray.io/u/psydok)\
**Post date:** [December 29, 2023, 5:11am UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214/3 "2023-12-29T05:11:12Z")

</div>

Thanks for the reply!  
Yes, the main node is open on port 6378 (as you said via --port 6378). I can’t use 6379 because the server is busy on that port…

---

<div class="post-metadata">

**Author:** ![zhanghx0905](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/zhanghx0905/32/5565_2.png) [@zhanghx0905](https://discuss.ray.io/u/zhanghx0905)\
**Post date:** [January 4, 2024, 10:35am UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214/4 "2024-01-04T10:35:51Z")

</div>

Hello, I’m attempting to manage a GPU cluster using ray. I’ve encountered the exact same issue as you have.  
My master node and workers are deployed on different servers, with the worker nodes being deployed within Docker containers. The worker disconnects a few seconds after the worker nodes connect to the master node. The phenomenon I’m experiencing is exactly as you’ve described.  
Do you now know how to resolve this now?

---

<div class="post-metadata">

**Author:** ![psydok](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/psydok/32/3518_2.png) [@psydok](https://discuss.ray.io/u/psydok)\
**Post date:** [January 9, 2024, 10:00am UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214/5 "2024-01-09T10:00:37Z")

</div>

I haven’t decided. But I think to try to add a rule in iptables for requests sent from docker-network.

---

<div class="post-metadata">

**Author:** ![Abhishek\_Jaiswal](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/abhishek_jaiswal/32/5121_2.png) [@Abhishek\_Jaiswal](https://discuss.ray.io/u/Abhishek_Jaiswal)\
**Post date:** [January 17, 2024, 9:55pm UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214/6 "2024-01-17T21:55:45Z")

</div>

Did you solve this problem? I’m facing the same problem

---

<div class="post-metadata">

**Author:** ![tej](https://avatars.discourse-cdn.com/v4/letter/t/48db29/32.png) [@tej](https://discuss.ray.io/u/tej)\
**Post date:** [April 19, 2024, 10:58pm UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214/7 "2024-04-19T22:58:19Z")

</div>

I also am facing this problem… is there an issue tracking this on the Ray GitHub?

Edit: I opened up ports from `10000:10099`, since on a docker-less setup with 3 workers, the controller took over ports `10000->10054` or so. This solved my issue for small ray cluster of 3. (Also increased `--max-worker-port 10099` ). Maybe the heartbeat evaluation is happening through another port which is limited by the small `10000:10010` range

---

<div class="post-metadata">

**Author:** ![Morphlng](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/morphlng/32/5785_2.png) [@Morphlng](https://discuss.ray.io/u/Morphlng)\
**Post date:** [October 25, 2025, 11:40am UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214/8 "2025-10-25T11:40:50Z")

</div>

@tej Hi! We are also trying to build a local ray cluster through docker swarm. Even though I published all ports ranging in `--min-worker-port` to `--max-worker-port`, we are still unable to train with the cluster, it will periodically report OwnerDiedError (worker cannot fetch Object).

I’m wondering if you could share some more detail about the setup of your cluster?

---

<div class="post-metadata">

**Author:** ![RunLLM](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/runllm/32/8249_2.png) [@RunLLM](https://discuss.ray.io/u/RunLLM)\
**Post date:** [October 25, 2025, 11:41am UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214/9 "2025-10-25T11:41:22Z")

</div>

The OwnerDiedError in Ray clusters running in Docker (including Docker Swarm) is often caused by networking or volume misconfiguration, not just by port exposure. Even if you publish all ports in the --min-worker-port to --max-worker-port range, issues can persist if, for example, the /tmp/ray directory is shared between head and worker containers, which can cause conflicts and actor failures. This was confirmed as the root cause in a similar case, where removing the shared /tmp/ray volume between head and worker nodes resolved the OwnerDiedError and actor death issues. Additionally, ensure that each container has its own /tmp/ray directory and that all required Ray ports (not just worker ports) are properly mapped and not blocked by firewalls or Docker networking rules.

For more details on this issue and its resolution, see the discussion and solution in [this GitHub issue](https://github.com/ray-project/ray/issues/48994). Would you like a step-by-step breakdown of recommended Docker Compose or Swarm settings for a stable Ray cluster?

Sources:

- [GitHub issue: Actor is dead because all references to the actor were removed including lineage ref count.](https://github.com/ray-project/ray/issues/48994)

Hint: Mention @RunLLM in the post for followups.
