# Sequence/Tensor Parallelism with Ray Serve

**URL:** <https://discuss.ray.io/t/sequence-tensor-parallelism-with-ray-serve/14765>\
**Category:** Uncategorized\
**Created:** [May 22, 2024, 4:29pm UTC](https://discuss.ray.io/t/sequence-tensor-parallelism-with-ray-serve/14765 "2024-05-22T16:29:50Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![skippy](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/skippy/32/6151_2.png) [@skippy](https://discuss.ray.io/u/skippy)\
**Post date:** [May 22, 2024, 4:29pm UTC](https://discuss.ray.io/t/sequence-tensor-parallelism-with-ray-serve/14765/1 "2024-05-22T16:29:50Z")

</div>

Are there any examples/demos on how to do this for inference. Got a big model which needs sequence parallelism and looking to split the workload 8x on a node.

Thanks

---

<div class="post-metadata">

**Author:** ![Sam\_Chan](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sam_chan/32/5011_2.png) [@Sam\_Chan](https://discuss.ray.io/u/Sam_Chan)\
**Post date:** [May 22, 2024, 10:22pm UTC](https://discuss.ray.io/t/sequence-tensor-parallelism-with-ray-serve/14765/2 "2024-05-22T22:22:56Z")

</div>

vLLM on Ray Serve will give you Tensor Parallelism baked in and probably your best bet. Guide coming soon!

---

<div class="post-metadata">

**Author:** ![Akshay\_Malik](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/akshay_malik/32/3807_2.png) [@Akshay\_Malik](https://discuss.ray.io/u/Akshay_Malik)\
**Post date:** [May 23, 2024, 3:57pm UTC](https://discuss.ray.io/t/sequence-tensor-parallelism-with-ray-serve/14765/3 "2024-05-23T15:57:01Z")

</div>

here’s an example of setting up vllm with Ray Serve - [Serve a Large Language Model with vLLM — Ray 3.0.0.dev0](https://docs.ray.io/en/master/serve/tutorials/vllm-example.html)
