Ray Serve


Ray Serve LLM APIs Ray Serve has LLM APIs to provide an easy way to deploy and scale multiple LLM models with a unified API. It supports automatic scaling, multi-model deployment, OpenAI-compatible endpoints, and LoRA multiplexing. The engine-agnostic architecture works with frameworks like vLLM and SGLang, enabling efficient model serving across multiple nodes.
Topic Replies Views Activity
0 839 November 17, 2020
1 59 July 28, 2026
11 1628 July 27, 2026
2 98 July 27, 2026
4 101 July 27, 2026
50 1306 June 3, 2026
15 896 June 2, 2026
3 198 May 13, 2026
8 712 May 3, 2026
1 64 April 30, 2026
1 71 April 30, 2026
4 111 April 27, 2026
1 126 February 18, 2026
5 109 February 12, 2026
0 19 January 20, 2026
1 169 December 23, 2025
1 103 December 22, 2025
4 123 December 9, 2025
3 192 December 1, 2025
2 80 October 30, 2025
1 55 October 30, 2025
8 1069 October 25, 2025
2 115 October 24, 2025
0 128 October 7, 2025
4 184 September 19, 2025
3 157 September 19, 2025
1 138 August 18, 2025
1 58 August 18, 2025
1 151 August 17, 2025
1 950 August 14, 2025