# Help designing fire and forget server for large batch inference

**URL:** <https://discuss.ray.io/t/help-designing-fire-and-forget-server-for-large-batch-inference/12953>\
**Category:** Ray Serve\
**Created:** [November 27, 2023, 2:25pm UTC](https://discuss.ray.io/t/help-designing-fire-and-forget-server-for-large-batch-inference/12953 "2023-11-27T14:25:10Z")\
**Posts on this page:** 1\
**Showing post:** 3

<div class="post-metadata">

**Author:** ![jonaz](https://avatars.discourse-cdn.com/v4/letter/j/e95f7d/32.png) [@jonaz](https://discuss.ray.io/u/jonaz)\
**Post date:** [November 29, 2023, 5:51pm UTC](https://discuss.ray.io/t/help-designing-fire-and-forget-server-for-large-batch-inference/12953/3 "2023-11-29T17:51:43Z")

</div>

I’ve trying to adapt my service to use Ray Workflow instead, but I’m running into some issues – getting OOMs due to memory not being released by Ray::IDLE processes (I posted the issue [here](https://discuss.ray.io/t/ray-idle-still-takes-a-lot-of-memory/12982/3)). Besides that, I still have the issue that each invocation of the workflow needs to load the model again from scratch, which is quite wasteful, since I always need the same model.

So, I’m still interested in my original questions from the first post;

1. Is there a risk of “losing” work by doing `deployment.remote()` but not `await`ing its result for example?

2. How large is the internal queue receiving requests when doing `deployment.remote()`? Is there a risk of it dropping requests?

---

_[View the full topic](https://discuss.ray.io/t/help-designing-fire-and-forget-server-for-large-batch-inference/12953)._
