# Does map\_batches avoid saturating the inference engine?

**URL:** <https://discuss.ray.io/t/does-map-batches-avoid-saturating-the-inference-engine/22538>\
**Category:** Ray Data LLM APIs\
**Created:** [May 21, 2025, 12:09am UTC](https://discuss.ray.io/t/does-map-batches-avoid-saturating-the-inference-engine/22538 "2025-05-21T00:09:57Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![shevateng0](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/shevateng0/32/7985_2.png) [@shevateng0](https://discuss.ray.io/u/shevateng0)\
**Post date:** [May 21, 2025, 12:09am UTC](https://discuss.ray.io/t/does-map-batches-avoid-saturating-the-inference-engine/22538/1 "2025-05-21T00:09:57Z")

</div>

I am experimenting with Ray Data for batch inference. Ray runs in standalone mode as I only use one GPU node.  
map\_batches() is used to perform batch inference. I want to know if batch inference via Ray will slow down the inference speed in standalone mode. The way I experiment is:

- Pick only one data batch of size 1000 (sorry not able to disclose the info of the data)
- Feed to (a) inference instantiated by `vllm.LLM` or `sglang.Engine`, and (b) vllm or sglang engine host in map\_batches()
- Compare the time used to finish inference.

I notice that map\_batches() is much slower than using the inference engine directly, like 50 mins vs 37 mins!

I want to understand what causes the slowness. Since I kept seeing the engine logs of SGLang popping “decode out of memory”, I wonder if Ray intentionally avoid saturating the engine too much to avoid crash. I also wonder if there is anyway we can speed up the inference.

Thanks.

---

<div class="post-metadata">

**Author:** ![lkchen](https://avatars.discourse-cdn.com/v4/letter/l/a88e57/32.png) [@lkchen](https://discuss.ray.io/u/lkchen)\
**Post date:** [May 25, 2025, 11:30pm UTC](https://discuss.ray.io/t/does-map-batches-avoid-saturating-the-inference-engine/22538/2 "2025-05-25T23:30:12Z")

</div>

Hi Sheva, thanks for the great quesiton!

Would you mind sharing the settings you have?

Previously I’ve done some investigations with vLLM 0.8.4 (with VLLM\_USE\_V1=1), you can see more details in [this PR](https://github.com/ray-project/ray/pull/52634):

- Prefer to use V1 over V0 for improved performance
- sync to the latest Ray version - the above mentioned PR landed after 2.46.0, and 2.47.0 is coming soon
- keep `batch_size` reasonably small (16-32) - the larger it is the more severe the long-tail problem, the smaller it is the more overhead
- increase `max_concurrent_batches` such that `max_concurrent_batches * batch_size` is large enough to saturate vLLM (monitor the warnings from `vllm_engine_stage.py`)

Utilizing ray.data does not magically boost LLM engines on single machine. However, the real value of `ray.data` is horizontal scaling, that is: if you plan to scale out the workload to multiple nodes/machines, `ray.data` abstracts the scheduling, orchestration, and fault tolerance, making distributed execution much easier without requiring manual coordination.

Would love to hear what results you get with the above tweaks—also keep in mind the long-tail behavior mentioned in the PR might vary based on your specific prompt.
