# Ray Serve

**URL:** https://discuss.ray.io/c/ray-serve/6.md

[Latest](https://discuss.ray.io/latest.md) · [Categories](https://discuss.ray.io/categories.md) · [Tags](https://discuss.ray.io/tags.md)

---

## [About the Ray Serve category](https://discuss.ray.io/t/about-the-ray-serve-category/48)

<div class="topic-metadata">

**Author:** [@bill-anyscale](https://discuss.ray.io/u/bill-anyscale)\
**Replies:** 0\
**Last updated:** [November 17, 2020, 12:02am UTC](https://discuss.ray.io/t/about-the-ray-serve-category/48 "2020-11-17T00:02:52Z")

</div>

Topics include model serving and inference. Use Serve to deploy and scale machine learning models with built-in support for APIs, batching, and multi-GPU inference. Ray Serve is a scalable model serving library for buil…

---

## [Maturity and plans of asynchronous inference](https://discuss.ray.io/t/maturity-and-plans-of-asynchronous-inference/23600)

<div class="topic-metadata">

**Author:** [@manuel.ramblr](https://discuss.ray.io/u/manuel.ramblr)\
**Replies:** 8\
**Last updated:** [September 3, 2026, 7:11am UTC](https://discuss.ray.io/t/maturity-and-plans-of-asynchronous-inference/23600 "2026-09-03T07:11:33Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Replica Ranker design](https://discuss.ray.io/t/replica-ranker-design/23593)

<div class="topic-metadata">

**Author:** [@glingleNxn](https://discuss.ray.io/u/glingleNxn)\
**Replies:** 1\
**Last updated:** [July 28, 2026, 1:10am UTC](https://discuss.ray.io/t/replica-ranker-design/23593 "2026-07-28T01:10:07Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Ray is creating hundreds of logs files under /tmp/ray/session\_latest/logs/ causing disk space issue and I/O Spikes](https://discuss.ray.io/t/ray-is-creating-hundreds-of-logs-files-under-tmp-ray-session-latest-logs-causing-disk-space-issue-and-i-o-spikes/12058)

<div class="topic-metadata">

**Author:** [@Abhi\_Sharma](https://discuss.ray.io/u/Abhi_Sharma)\
**Replies:** 11\
**Last updated:** [July 27, 2026, 11:20pm UTC](https://discuss.ray.io/t/ray-is-creating-hundreds-of-logs-files-under-tmp-ray-session-latest-logs-causing-disk-space-issue-and-i-o-spikes/12058 "2026-07-27T23:20:32Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. Ray v2.6.3, Ray is creating hundreds of files under /tmp/ray/session\_latest/logs/ worker-\*.err, worker-\*.out and…

---

## [Programmatic lightweight update from rest call](https://discuss.ray.io/t/programmatic-lightweight-update-from-rest-call/23395)

<div class="topic-metadata">

**Author:** [@glingleNxn](https://discuss.ray.io/u/glingleNxn)\
**Replies:** 2\
**Last updated:** [July 27, 2026, 11:07pm UTC](https://discuss.ray.io/t/programmatic-lightweight-update-from-rest-call/23395 "2026-07-27T23:07:11Z")

</div>

1. Severity of the issue: (select one) Medium: Significantly affects my productivity but can find a workaround. High: Completely blocks me. 2. Environment: Ray version: 2.48.0 Python version: 3.9 OS: macos, linux C…

---

## [Ray server for use case with limited memory resources](https://discuss.ray.io/t/ray-server-for-use-case-with-limited-memory-resources/23586)

<div class="topic-metadata">

**Author:** [@mountim](https://discuss.ray.io/u/mountim)\
**Replies:** 4\
**Last updated:** [July 27, 2026, 10:46am UTC](https://discuss.ray.io/t/ray-server-for-use-case-with-limited-memory-resources/23586 "2026-07-27T10:46:16Z")

</div>

1. Severity of the issue: (select one) High: Completely blocks me. 2. Environment: Ray version: Latest (currently 2.56.1) OS: Linux Hello, I am exploring Ray serve to see if it fits the next use case. I hope someb…

---

## [Memory not released to default levels: \`ray::IDLE\` Processes Not Released\*\*](https://discuss.ray.io/t/memory-not-released-to-default-levels-ray-idle-processes-not-released/23295)

<div class="topic-metadata">

**Author:** [@dmtryzarubin](https://discuss.ray.io/u/dmtryzarubin)\
**Replies:** 50\
**Last updated:** [June 3, 2026, 12:33pm UTC](https://discuss.ray.io/t/memory-not-released-to-default-levels-ray-idle-processes-not-released/23295 "2026-06-03T12:33:08Z")

</div>

1. Severity of the issue: (select one) Medium: Significantly affects my productivity but can find a workaround. 2. Environment: Ray version: 2.49.1 Python version: 3.10.18 OS: rayproject/ray:2.49.1-py310-cpu image C…

---

## [Setup api key to call LLM via rayserve](https://discuss.ray.io/t/setup-api-key-to-call-llm-via-rayserve/23319)

<div class="topic-metadata">

**Author:** [@AlaEddine](https://discuss.ray.io/u/AlaEddine)\
**Replies:** 15\
**Last updated:** [June 2, 2026, 11:41am UTC](https://discuss.ray.io/t/setup-api-key-to-call-llm-via-rayserve/23319 "2026-06-02T11:41:58Z")

</div>

Dear Ray community, I deployed kubeRay on kubernetes , and im serve my local LLMs , i want to add API Key when requesting my LLM, via rayserve service how i can setup api key ,is there an environment variable must be a…

---

## [How do I run unit tests for Ray Serve Pull Request?](https://discuss.ray.io/t/how-do-i-run-unit-tests-for-ray-serve-pull-request/23547)

<div class="topic-metadata">

**Author:** [@wanadzhar913](https://discuss.ray.io/u/wanadzhar913)\
**Replies:** 3\
**Last updated:** [May 13, 2026, 6:14pm UTC](https://discuss.ray.io/t/how-do-i-run-unit-tests-for-ray-serve-pull-request/23547 "2026-05-13T18:14:06Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [How to route traffic to LiteLLM models using Serving LLMs](https://discuss.ray.io/t/how-to-route-traffic-to-litellm-models-using-serving-llms/22493)

<div class="topic-metadata">

**Author:** [@James\_Wong](https://discuss.ray.io/u/James_Wong)\
**Replies:** 8\
**Last updated:** [May 3, 2026, 10:29pm UTC](https://discuss.ray.io/t/how-to-route-traffic-to-litellm-models-using-serving-llms/22493 "2026-05-03T22:29:43Z")

</div>

We’re currently serving vLLM models via the ‘Serving LLMs’, Serving LLMs — Ray 2.46.0 Is there a way or an example to use Ray Serve to pass through LiteLLM models? We’re looking at the ‘Deploy Compositions of Models’ a…

---

## [Ray Serve LLM on CPU with KubeRay](https://discuss.ray.io/t/ray-serve-llm-on-cpu-with-kuberay/23523)

<div class="topic-metadata">

**Author:** [@CaoimheCahill](https://discuss.ray.io/u/CaoimheCahill)\
**Replies:** 1\
**Last updated:** [April 30, 2026, 9:51am UTC](https://discuss.ray.io/t/ray-serve-llm-on-cpu-with-kuberay/23523 "2026-04-30T09:51:40Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Actor task fail running under Serve: is it normal to have this depth?](https://discuss.ray.io/t/actor-task-fail-running-under-serve-is-it-normal-to-have-this-depth/23522)

<div class="topic-metadata">

**Author:** [@scorchio](https://discuss.ray.io/u/scorchio)\
**Replies:** 1\
**Last updated:** [April 30, 2026, 9:40am UTC](https://discuss.ray.io/t/actor-task-fail-running-under-serve-is-it-normal-to-have-this-depth/23522 "2026-04-30T09:40:00Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. 2. Environment: Ray version: 2.51.1 Python version: 3.10 OS: python:3.10-slim Docker image running on Ubuntu 24.04 Cloud/Infrastru…

---

## [HAProxy Config customization](https://discuss.ray.io/t/haproxy-config-customization/23516)

<div class="topic-metadata">

**Author:** [@glingleNxn](https://discuss.ray.io/u/glingleNxn)\
**Replies:** 4\
**Last updated:** [April 27, 2026, 9:13pm UTC](https://discuss.ray.io/t/haproxy-config-customization/23516 "2026-04-27T21:13:24Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Load models from Docker volume without creating copies](https://discuss.ray.io/t/load-models-from-docker-volume-without-creating-copies/23481)

<div class="topic-metadata">

**Author:** [@mgb](https://discuss.ray.io/u/mgb)\
**Replies:** 1\
**Last updated:** [February 18, 2026, 2:39pm UTC](https://discuss.ray.io/t/load-models-from-docker-volume-without-creating-copies/23481 "2026-02-18T14:39:17Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Downloading models from custom sources when using LLMConfig](https://discuss.ray.io/t/downloading-models-from-custom-sources-when-using-llmconfig/23473)

<div class="topic-metadata">

**Author:** [@psnilesh](https://discuss.ray.io/u/psnilesh)\
**Replies:** 5\
**Last updated:** [February 12, 2026, 2:08am UTC](https://discuss.ray.io/t/downloading-models-from-custom-sources-when-using-llmconfig/23473 "2026-02-12T02:08:10Z")

</div>

I’m following the quickstart example given at Quickstart examples — Ray 2.53.0 for deploying Ray serve applications. For one of my deployments, I have a unique requirement to download the model on the fly from a proprie…

---

## [Optimal redis cache size for ray gcs backup](https://discuss.ray.io/t/optimal-redis-cache-size-for-ray-gcs-backup/23457)

<div class="topic-metadata">

**Author:** [@bhartendu\_kumar](https://discuss.ray.io/u/bhartendu_kumar)\
**Replies:** 0\
**Last updated:** [January 20, 2026, 3:24pm UTC](https://discuss.ray.io/t/optimal-redis-cache-size-for-ray-gcs-backup/23457 "2026-01-20T15:24:57Z")

</div>

I am planning to use redis aws elasticache as backup for GCS, Serving production traffic with RayServe and kuberay as operator. What should be the optimal AWS elasticache instance size, considering high scale?

---

## [Example docker compose to run RayServe app](https://discuss.ray.io/t/example-docker-compose-to-run-rayserve-app/23409)

<div class="topic-metadata">

**Author:** [@glingleNxn](https://discuss.ray.io/u/glingleNxn)\
**Replies:** 1\
**Last updated:** [December 23, 2025, 8:23pm UTC](https://discuss.ray.io/t/example-docker-compose-to-run-rayserve-app/23409 "2025-12-23T20:23:50Z")

</div>

1. Severity of the issue: (select one) Medium: Significantly affects my productivity but can find a workaround. High: Completely blocks me. 2. Environment: Ray version: 2.48 Python version: 3.9 OS: osx/linux Cloud…

---

## [Deploying Multiple Ray Serve Microservices on a Single Cluster with Separate Ports](https://discuss.ray.io/t/deploying-multiple-ray-serve-microservices-on-a-single-cluster-with-separate-ports/23403)

<div class="topic-metadata">

**Author:** [@ABINA\_SRINIVASAN](https://discuss.ray.io/u/ABINA_SRINIVASAN)\
**Replies:** 1\
**Last updated:** [December 22, 2025, 7:04am UTC](https://discuss.ray.io/t/deploying-multiple-ray-serve-microservices-on-a-single-cluster-with-separate-ports/23403 "2025-12-22T07:04:56Z")

</div>

I am building a microservices-based architecture using Ray Serve, where each microservice has its own deployment YAML and a predefined port configuration. However, I am facing the following challenges: When using serve…

---

## [About Ray DAG API for serve.deployment at Ray 2.44.1](https://discuss.ray.io/t/about-ray-dag-api-for-serve-deployment-at-ray-2-44-1/23366)

<div class="topic-metadata">

**Author:** [@czjghost](https://discuss.ray.io/u/czjghost)\
**Replies:** 4\
**Last updated:** [December 9, 2025, 1:55am UTC](https://discuss.ray.io/t/about-ray-dag-api-for-serve-deployment-at-ray-2-44-1/23366 "2025-12-09T01:55:10Z")

</div>

1. Severity of the issue: (select one) Medium: Significantly affects my productivity but can find a workaround. 2. Environment: Ray version: 2.44.1 Python version: 3.9.9 OS: Linux 5.10.0-60.18.0.50.r509\_2.hce2.aarch…

---

## [Preprocessing in ray serve LLM](https://discuss.ray.io/t/preprocessing-in-ray-serve-llm/23346)

<div class="topic-metadata">

**Author:** [@manuel.ramblr](https://discuss.ray.io/u/manuel.ramblr)\
**Replies:** 3\
**Last updated:** [December 1, 2025, 11:50pm UTC](https://discuss.ray.io/t/preprocessing-in-ray-serve-llm/23346 "2025-12-01T23:50:06Z")

</div>

Hi everyone, I want to use ray serve LLM in the following way to host a model that is found on huggingface. In particular, I’m looking into this model: As you can see in their huggingface code, the preprocessing one…

---

## [TypeError: Failed to serialize the ASGI app.:](https://discuss.ray.io/t/typeerror-failed-to-serialize-the-asgi-app/23285)

<div class="topic-metadata">

**Author:** [@Ishika\_Sahu](https://discuss.ray.io/u/Ishika_Sahu)\
**Replies:** 2\
**Last updated:** [October 30, 2025, 9:27am UTC](https://discuss.ray.io/t/typeerror-failed-to-serialize-the-asgi-app/23285 "2025-10-30T09:27:28Z")

</div>

High: Completely blocks me. Ray version: 2.51.0 Python version: 3.12.12 OS: Windows Cloud/Infrastructure: Colab I keep getting this error when i try to run the below code. TypeError: self.handle cannot be converted t…

---

## [Serve deploy app support custom router with runtime\_env](https://discuss.ray.io/t/serve-deploy-app-support-custom-router-with-runtime-env/23284)

<div class="topic-metadata">

**Author:** [@wciq1208](https://discuss.ray.io/u/wciq1208)\
**Replies:** 1\
**Last updated:** [October 30, 2025, 2:30am UTC](https://discuss.ray.io/t/serve-deploy-app-support-custom-router-with-runtime-env/23284 "2025-10-30T02:30:49Z")

</div>

1. Severity of the issue: (select one) Medium: Significantly affects my productivity but can find a workaround. 2. Environment: Ray version:2.50.1 Python version:3.12 OS:ubuntu Cloud/Infrastructure:gcp Other libs/to…

---

## [\[Serve\] The \`ray start --head --node-ip-address ip\` is not working correctly in Docker. And it's not clear which ports to open](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214)

<div class="topic-metadata">

**Author:** [@psydok](https://discuss.ray.io/u/psydok)\
**Replies:** 8\
**Last updated:** [October 25, 2025, 11:41am UTC](https://discuss.ray.io/t/serve-the-ray-start-head-node-ip-address-ip-is-not-working-correctly-in-docker-and-its-not-clear-which-ports-to-open/13214 "2025-10-25T11:41:22Z")

</div>

I’m trying to connect nodes deployed via docker to the master node. I am having a number of problems, which locally I was able to solve by setting network\_mode: host to containers. But on the servers there is a firewall …

---

## [Nvidea-smi errors when deploying ray serve head on cpu only node](https://discuss.ray.io/t/nvidea-smi-errors-when-deploying-ray-serve-head-on-cpu-only-node/23138)

<div class="topic-metadata">

**Author:** [@Ruckley](https://discuss.ray.io/u/Ruckley)\
**Replies:** 2\
**Last updated:** [October 24, 2025, 5:51pm UTC](https://discuss.ray.io/t/nvidea-smi-errors-when-deploying-ray-serve-head-on-cpu-only-node/23138 "2025-10-24T17:51:48Z")

</div>

1. Severity of the issue: (select one) High: Completely blocks me. 2. Environment: Ray version: 2.49.2 Python version: 3.11.6 OS: linux Cloud/Infrastructure: Azure AKS Other libs/tools (if relevant): 3. What happen…

---

## [Running Multiple Ray Heads on Same Node - Safety & Best Practices?](https://discuss.ray.io/t/running-multiple-ray-heads-on-same-node-safety-best-practices/23234)

<div class="topic-metadata">

**Author:** [@michaelripa](https://discuss.ray.io/u/michaelripa)\
**Replies:** 0\
**Last updated:** [October 7, 2025, 7:48pm UTC](https://discuss.ray.io/t/running-multiple-ray-heads-on-same-node-safety-best-practices/23234 "2025-10-07T19:48:02Z")

</div>

Hello, I’m working on deploying multiple Ray Serve applications with head nodes residing on the same physical node on a system where I don’t have access to Docker. In the ideal scenario, I would be able to do this witho…

---

## [Ray Serve not distributing load to all replicas equally](https://discuss.ray.io/t/ray-serve-not-distributing-load-to-all-replicas-equally/22589)

<div class="topic-metadata">

**Author:** [@manickavela29](https://discuss.ray.io/u/manickavela29)\
**Replies:** 4\
**Last updated:** [September 19, 2025, 2:53pm UTC](https://discuss.ray.io/t/ray-serve-not-distributing-load-to-all-replicas-equally/22589 "2025-09-19T14:53:37Z")

</div>

1. Severity of the issue: (select one) High: Completely blocks me. 2. Environment: Ray version: rayproject/ray:2.41.0 Python version: OS: ubuntu Cloud/Infrastructure: aws Other libs/tools (if relevant): 3. What hap…

---

## [Non-linear throughput when scaling Ray Serve replicas](https://discuss.ray.io/t/non-linear-throughput-when-scaling-ray-serve-replicas/22959)

<div class="topic-metadata">

**Author:** [@raphael](https://discuss.ray.io/u/raphael)\
**Replies:** 3\
**Last updated:** [September 19, 2025, 2:50pm UTC](https://discuss.ray.io/t/non-linear-throughput-when-scaling-ray-serve-replicas/22959 "2025-09-19T14:50:15Z")

</div>

Severity: Medium – Significantly affects my productivity but I can find a workaround. Environment: Ray version: 2.48.0 Python version: 3.12.11 OS: Ubuntu 22, no docker Infra: Ray autoscaler with AWS EC2 …

---

## [FastAPI backend + Ray Core vs Ray Serve](https://discuss.ray.io/t/fastapi-backend-ray-core-vs-ray-serve/22972)

<div class="topic-metadata">

**Author:** [@vaporeon](https://discuss.ray.io/u/vaporeon)\
**Replies:** 1\
**Last updated:** [August 18, 2025, 8:21pm UTC](https://discuss.ray.io/t/fastapi-backend-ray-core-vs-ray-serve/22972 "2025-08-18T20:21:38Z")

</div>

I have an existing backend FastAPI server handling business logic. How can I extend this beckend to add a simple machine learning model that inference a POST request? I have read Have read Ray with FastAPI - Ray Core - …

---

## [Stop Ray Serve from overwriting LD\_LIBRARY\_PATH?](https://discuss.ray.io/t/stop-ray-serve-from-overwriting-ld-library-path/23014)

<div class="topic-metadata">

**Author:** [@jdwillard19](https://discuss.ray.io/u/jdwillard19)\
**Replies:** 1\
**Last updated:** [August 18, 2025, 4:24pm UTC](https://discuss.ray.io/t/stop-ray-serve-from-overwriting-ld-library-path/23014 "2025-08-18T16:24:25Z")

</div>

1. Severity of the issue: High 2. Environment: Ray version: 2.48.0 Python version: 3.12.10 OS: Red Hat Enterprise Linux 8.8 Cloud/Infrastructure: HPC Environment (no root) Other libs/tools (if relevant): 3. What happ…

---

## [Trouble deploying simple app with uv](https://discuss.ray.io/t/trouble-deploying-simple-app-with-uv/23003)

<div class="topic-metadata">

**Author:** [@Dominic\_Laflamme](https://discuss.ray.io/u/Dominic_Laflamme)\
**Replies:** 1\
**Last updated:** [August 17, 2025, 2:59pm UTC](https://discuss.ray.io/t/trouble-deploying-simple-app-with-uv/23003 "2025-08-17T14:59:57Z")

</div>

I’m trying to deploy a Ray Serve app to an already running cluster (connecting with ray.init(address="ray://\<HEAD\>:10001")). The project is a package: myapp/ (with \_\_init\_\_.py) contains app.py that builds the graph (entr…

[Next page](https://discuss.ray.io/c/ray-serve/6.md?page=1)
