# Set different HTTP port for different deployments

**URL:** <https://discuss.ray.io/t/set-different-http-port-for-different-deployments/21746>\
**Category:** Uncategorized\
**Created:** [February 11, 2025, 4:28am UTC](https://discuss.ray.io/t/set-different-http-port-for-different-deployments/21746 "2025-02-11T04:28:07Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![kuku](https://avatars.discourse-cdn.com/v4/letter/k/76d3ee/32.png) [@kuku](https://discuss.ray.io/u/kuku)\
**Post date:** [February 11, 2025, 4:28am UTC](https://discuss.ray.io/t/set-different-http-port-for-different-deployments/21746/1 "2025-02-11T04:28:07Z")

</div>

I’m using Ray Serve and have two separate deployments.

- In the first deployment, I start Ray Serve with:

`serve.start(http_options={"host": "0.0.0.0", "port": os.environ.get("HTTP_PORT")})`

- In the second deployment, which has a different application name, I use:

`serve.start(http_options={"host": "0.0.0.0", "port": os.environ.get("HTTP_PORT_2")})`

However, in the second deployment, `HTTP_PORT_2` is being ignored.

How can I set a different port for the second deployment?

---

<div class="post-metadata">

**Author:** ![christina](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/christina/32/7542_2.png) [@christina](https://discuss.ray.io/u/christina)\
**Post date:** [February 11, 2025, 8:28pm UTC](https://discuss.ray.io/t/set-different-http-port-for-different-deployments/21746/2 "2025-02-11T20:28:31Z")

</div>

Hi kuku! Welcome to the Ray community 😊

Ray Serve only allows **one HTTP server per Ray cluster**. When you call `serve.start()` a second time with a different port, it does not create a new HTTP server—it simply connects to the existing Ray Serve instance, which is already using the first HTTP port.

So there’s a few different ways you can get around resolving this.

#### Option 1: Run Each Deployment on a Separate Ray Cluster

Since HTTP configuration is cluster-scoped, you need to run each application in a separate Ray cluster to have different HTTP ports. Example:

```auto
# Start first Ray cluster
ray start --head --port=6379

# Deploy first application
serve.start(http_options={"host": "0.0.0.0", "port": os.environ.get("HTTP_PORT")})

# Start second Ray cluster on a different port
ray start --head --port=6380

# Deploy second application
serve.start(http_options={"host": "0.0.0.0", "port": os.environ.get("HTTP_PORT_2")})

```

Each deployment runs independently on its own Ray cluster, allowing different ports.

#### Option 2: Use Ray Serve Multi-Application Support

If running multiple clusters is not feasible, you can deploy multiple applications on the same Serve instance using **Serve Deployments**.

1. Define multiple apps in a Serve config YAML
2. Deploy it (you can read our docs to see how to do this specifically, I will link it below)  
Instead of different ports, each application gets a different route name (e.g., `/app1` and `/app2`).

#### Option 3: Reverse Proxy

If you must use the same Ray cluster, but different external ports, you can use a reverse proxy like NGINX to map requests to different Serve applications.

Here’s some of the docs:  
**Docs:**

- [Deploy Multiple Applications — Ray 2.52.0](https://docs.ray.io/en/latest/serve/multi-app.html#deploy-multiple-applications)
- [Development Workflow — Ray 2.52.0](https://docs.ray.io/en/latest/serve/advanced-guides/dev-workflow.html#local-development-with-http-requests)
- [Configure Ray Serve deployments — Ray 2.52.0](https://docs.ray.io/en/latest/serve/configure-serve-deployment.html#configure-ray-serve-deployments)
- [Deploy on VM — Ray 2.52.0](https://docs.ray.io/en/latest/serve/advanced-guides/deploy-vm.html#using-a-remote-cluster)

---

<div class="post-metadata">

**Author:** ![kuku](https://avatars.discourse-cdn.com/v4/letter/k/76d3ee/32.png) [@kuku](https://discuss.ray.io/u/kuku)\
**Post date:** [February 11, 2025, 11:29pm UTC](https://discuss.ray.io/t/set-different-http-port-for-different-deployments/21746/3 "2025-02-11T23:29:33Z")

</div>

Thanks, Christina.

I’m currently using option 2, where I set a different application name and `route_prefix` in `serve.run()`.

My current setup:

1. In **`file1.py`** , I have:

2. In **`file2.py`** , the structure is similar to `file1.py`.

3. In the terminal, I run:

I tend to confuse with ray.init(), serve.start and serve.run(), what is the better deployment workflow for my case?

---

<div class="post-metadata">

**Author:** ![christina](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/christina/32/7542_2.png) [@christina](https://discuss.ray.io/u/christina)\
**Post date:** [February 12, 2025, 4:11am UTC](https://discuss.ray.io/t/set-different-http-port-for-different-deployments/21746/4 "2025-02-12T04:11:19Z")

</div>

🙂 I can try to explain what the different functions do.

1. **ray.init()**: This is basically letting your script know it needs to connect to an existing Ray cluster. If you don’t provide specific details, it’ll try to start a local Ray cluster. This is necessary before you can use any Ray functionalities, including Ray Serve.
2. **serve.start()**: This kicks off Ray Serve in your cluster. It reads your HTTP options (but in your case, since it’s multi-app mode, it cares more about route prefixes). You only need to call this once per cluster session.
3. **serve.run()**: This is where you actually set your deployments live, using any configurations you’ve set up, like your app names and routes. If you set `blocking=True` , the function will block the terminal, which is useful for development and debugging as it streams logs to the console. However, for running multiple applications or scripts, you might want to run it in a non-blocking mode or in the background.

There’s a few deployment workflows too.

- **Single Application** : If you are running a single application, you can use `serve.run()` with `blocking=True` to keep the terminal open for logs and debugging.
- **Multiple Applications** : Since you have multiple scripts (`file1.py` and `file2.py` ), you should consider running `serve.run()` in a non-blocking mode or in the background. This can be done by using `&` in the terminal to run the command in the background or by setting `blocking=False` if you are using a script. (By running the scripts in the background, you can manage multiple applications more effectively.) Ensure that each application has a unique `route_prefix` to avoid conflicts.

Essentially, you can try starting each Python script using a non-blocking approach if they need to run concurrently. If you don’t want to use blocking=True, you could devise a way to keep the process running after deployment without blocking the terminal, with proper process management.

---

<div class="post-metadata">

**Author:** ![ABINA\_SRINIVASAN](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/abina_srinivasan/32/8416_2.png) [@ABINA\_SRINIVASAN](https://discuss.ray.io/u/ABINA_SRINIVASAN)\
**Post date:** [December 22, 2025, 6:59am UTC](https://discuss.ray.io/t/set-different-http-port-for-different-deployments/21746/5 "2025-12-22T06:59:54Z")

</div>

I am building a microservices-based architecture using **Ray Serve** , where each microservice has its own deployment YAML and a predefined port configuration.

However, I am facing the following challenges:

- When using `serve deploy` or `serve run`, each microservice appears to start its **own Ray cluster** , instead of deploying onto a shared cluster.

- Deploying multiple microservices on the **same port** results in the previous deployment being overwritten.

- Even when different ports are specified in individual YAML files, the services are not coexisting as expected on the same cluster.

My goal is to:

- Deploy **multiple Ray Serve applications (microservices)** on a **single shared Ray cluster**

- Expose each microservice on a **different port**

- Manage all services centrally without them deleting or replacing each other

Would **KubeRay** be the correct approach to achieve this?  
Specifically:

- Can KubeRay manage a single Ray cluster and deploy multiple Ray Serve applications into it?

- Is it possible to expose each Ray Serve application on a different port using Kubernetes Services or Ingress?

- What is the recommended architecture or best practice for deploying multiple Ray Serve microservices in a single Kubernetes cluster?

Any guidance, architectural patterns, or references would be greatly helpful to us

---

<div class="post-metadata">

**Author:** ![ABINA\_SRINIVASAN](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/abina_srinivasan/32/8416_2.png) [@ABINA\_SRINIVASAN](https://discuss.ray.io/u/ABINA_SRINIVASAN)\
**Post date:** [January 6, 2026, 10:57am UTC](https://discuss.ray.io/t/set-different-http-port-for-different-deployments/21746/6 "2026-01-06T10:57:06Z")

</div>

can anyone give me guidence for this?
