# vLLM v1 engine initialization workaround with vllm installation at runtime

**URL:** https://discuss.ray.io/t/vllm-v1-engine-initialization-workaround-with-vllm-installation-at-runtime/22827
**Category:** Ray Serve LLM APIs
**Created:** [July 11, 2025, 7:45pm UTC](https://discuss.ray.io/t/vllm-v1-engine-initialization-workaround-with-vllm-installation-at-runtime/22827 "2025-07-11T19:45:56Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![404notfound101](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/404notfound101/32/8091_2.png) [@404notfound101](https://discuss.ray.io/u/404notfound101)
#### Post date: [July 11, 2025, 7:45pm UTC](https://discuss.ray.io/t/vllm-v1-engine-initialization-workaround-with-vllm-installation-at-runtime/22827/1 "2025-07-11T19:45:56Z")

</div>

Hi team,

I am working on using `ray.serve.llm:build_openai_app` to serve LLM on an EKS cluster. Due to permission issues, I am not allowed to modify the worker image to pre-install `vllm`, but have to include it through `runtime_env` like:

```auto
...
    - import_path: ray.serve.llm:build_openai_app
      name: ...
      route_prefix: /
      runtime_env:
        pip:
        - vllm==0.8.5
...

```

It’s all good with v0 engine. However, by switching to v1 engine, I noticed the RayWorkerWrapper process failed to initialize because of the absence of `vllm`. I can see failed jobs with the entrypoint: `/tmp/ray/session_2025-07-11_18-56-02_402046_1/runtime_resources/pip/fffafa2881e929aa2b12b38ecc3f1e0f8255ad62/virtualenv/bin/python -c "from multiprocessing.spawn import spawn_main; spawn_main(tracker_fd=123, pipe_handle=125)" --multiprocessing-fork`  
I am not totally sure how this works, but it seems to be related to when and how the new process gets spawned.

Do you know what the cause is? Are there any workarounds

Thanks!

---

<div class="post-metadata">

### Author: ![kourosh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kourosh/32/2130_2.png) [@kourosh](https://discuss.ray.io/u/kourosh)
#### Post date: [July 11, 2025, 8:20pm UTC](https://discuss.ray.io/t/vllm-v1-engine-initialization-workaround-with-vllm-installation-at-runtime/22827/2 "2025-07-11T20:20:25Z")

</div>

hi @404notfound101 ,

not sure exactly why this is happening, but is there any way to use ray-llm pre-built images from here? They come with vLLM already.

[https://hub.docker.com/r/rayproject/ray-llm/tags](https://hub.docker.com/r/rayproject/ray-llm/tags)

---

<div class="post-metadata">

### Author: ![404notfound101](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/404notfound101/32/8091_2.png) [@404notfound101](https://discuss.ray.io/u/404notfound101)
#### Post date: [July 11, 2025, 11:12pm UTC](https://discuss.ray.io/t/vllm-v1-engine-initialization-workaround-with-vllm-installation-at-runtime/22827/3 "2025-07-11T23:12:06Z")

</div>

Hi @kourosh,

Thanks for the quick reply. Unfortunately, I cannot use other images nor pre-install `vllm` in our image.

I think it’s because the v1 engine spawns new processes instead of forking from the main actor. I really hope there’s a workaround for monkey-patching dependencies at the deployment level.

---

<div class="post-metadata">

### Author: ![kourosh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kourosh/32/2130_2.png) [@kourosh](https://discuss.ray.io/u/kourosh)
#### Post date: [July 11, 2025, 11:31pm UTC](https://discuss.ray.io/t/vllm-v1-engine-initialization-workaround-with-vllm-installation-at-runtime/22827/4 "2025-07-11T23:31:41Z")

</div>

can you maybe try later versions of vllm via runtime? I recall after some version we switched from fork to spawn in the vllm when ray is the context.

---

<div class="post-metadata">

### Author: ![ehiggi](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ehiggi/32/8109_2.png) [@ehiggi](https://discuss.ray.io/u/ehiggi)
#### Post date: [July 20, 2025, 6:22am UTC](https://discuss.ray.io/t/vllm-v1-engine-initialization-workaround-with-vllm-installation-at-runtime/22827/5 "2025-07-20T06:22:52Z")

</div>

We also ran into this issue - this PR should fix it [[Misc] allow pulling vllm in Ray runtime environment by eric-higgins-ai · Pull Request #21143 · vllm-project/vllm · GitHub](https://github.com/vllm-project/vllm/pull/21143)
