# About the Ray Serve LLM APIs category

**URL:** <https://discuss.ray.io/t/about-the-ray-serve-llm-apis-category/22205>\
**Category:** Ray Serve LLM APIs\
**Created:** [April 2, 2025, 6:24pm UTC](https://discuss.ray.io/t/about-the-ray-serve-llm-apis-category/22205 "2025-04-02T18:24:58Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![christina](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/christina/32/7542_2.png) [@christina](https://discuss.ray.io/u/christina)\
**Post date:** [April 2, 2025, 6:24pm UTC](https://discuss.ray.io/t/about-the-ray-serve-llm-apis-category/22205/1 "2025-04-02T18:24:58Z")

</div>

Ray Serve has LLM APIs to provide an easy way to deploy and scale multiple LLM models with a unified API. It supports automatic scaling, multi-model deployment, OpenAI-compatible endpoints, and LoRA multiplexing. The engine-agnostic architecture works with frameworks like vLLM and SGLang, enabling efficient model serving across multiple nodes.

More documentation:

- [https://docs.ray.io/en/latest/serve/llm/serving-llms.html](https://docs.ray.io/en/latest/serve/llm/serving-llms.html)
