# About the Ray Serve category

**URL:** <https://discuss.ray.io/t/about-the-ray-serve-category/48>\
**Category:** Ray Serve\
**Created:** [November 17, 2020, 12:02am UTC](https://discuss.ray.io/t/about-the-ray-serve-category/48 "2020-11-17T00:02:52Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![bill-anyscale](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/bill-anyscale/32/724_2.png) [@bill-anyscale](https://discuss.ray.io/u/bill-anyscale)\
**Post date:** [November 17, 2020, 12:02am UTC](https://discuss.ray.io/t/about-the-ray-serve-category/48/1 "2020-11-17T00:02:52Z")

</div>

Topics include model serving and inference. Use Serve to deploy and scale machine learning models with built-in support for APIs, batching, and multi-GPU inference.

Ray Serve is a scalable model serving library for building online inference APIs. Serve is framework-agnostic, so you can use a single toolkit to serve everything from deep learning models built with frameworks like PyTorch, TensorFlow, and Keras, to Scikit-Learn models, to arbitrary Python business logic. It has several features and performance optimizations for serving Large Language Models such as response streaming, dynamic request batching, multi-node/multi-GPU serving, etc.

Ray Serve is particularly well suited for [model composition](https://docs.ray.io/en/latest/serve/model_composition.html#serve-model-composition) and many model serving, enabling you to build a complex inference service consisting of multiple ML models and business logic all in Python code.

**Documentation** :

- [Getting Started — Ray 2.43.0](https://docs.ray.io/en/latest/serve/getting_started.html)
- [Key Concepts — Ray 2.43.0](https://docs.ray.io/en/latest/serve/key-concepts.html)
- [Develop and Deploy an ML Application — Ray 2.43.0](https://docs.ray.io/en/latest/serve/develop-and-deploy.html)
