# Parallel inference using CPUs

**URL:** <https://discuss.ray.io/t/parallel-inference-using-cpus/11177>\
**Category:** Ray Core\
**Created:** [June 27, 2023, 8:13pm UTC](https://discuss.ray.io/t/parallel-inference-using-cpus/11177 "2023-06-27T20:13:20Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![xzf0kgb0bqr.cev2RWU](https://avatars.discourse-cdn.com/v4/letter/x/65b543/32.png) [@xzf0kgb0bqr.cev2RWU](https://discuss.ray.io/u/xzf0kgb0bqr.cev2RWU)\
**Post date:** [June 27, 2023, 8:13pm UTC](https://discuss.ray.io/t/parallel-inference-using-cpus/11177/1 "2023-06-27T20:13:20Z")

</div>

- Medium: It contributes to significant difficulty to complete my task, but I can work around it.

**Main question**

I have a trained model. I have one GPU node which has 4 GPUs and 40 CPUs. I wish to apply the model in parallel over the 40 CPU nodes to test\_data (so each CPU gets 1/40 of the test\_data). How can I do this?

**More details**  
I would like to avoid using Ray AIR if possible for two reasons:

1. It is in beta testing.
2. I would need to convert my PyTorch DataLoader to a Ray AIR Dataset. The [github issues page](https://github.com/ray-project/ray/issues/31418) says that a tutorial on this is planned but not done yet, so I don’t know how to do this.

From googling this question, I see a lot of questions about parallelizing over the 4 GPUs, but since I have 40 CPUs, I think parallelizing over CPUs instead of GPUs would be faster. I am using PyTorch.

A skeleton code that loads the dataset, dataloader, and model are provided below.

```auto
from torch.utils.data import Dataset
from torch.utils.data import DataLoader

my_dataset = Dataset(...)
my_loader = DataLoader(my_dataset, ...)

state_dict = torch.load(model_save_location)
model.load_state_dict(state_dict)

device = torch.device('cpu')
model = model.to(device)

```

Notes:

1. I am aware that I can use Pool from torch.multiprocessing (as detailed [here](https://stackoverflow.com/questions/58150186/run-inference-on-cpu-using-pytorch-and-multiprocessing)). But I would prefer to use Ray because if I want to scale to multiple nodes in the future (I only have 1 GPU node now), it would be much easier with Ray than without, I think.

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [June 27, 2023, 11:49pm UTC](https://discuss.ray.io/t/parallel-inference-using-cpus/11177/2 "2023-06-27T23:49:11Z")

</div>

cc @amogkam can you address this question?

---

<div class="post-metadata">

**Author:** ![amogkam](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amogkam/32/17_2.png) [@amogkam](https://discuss.ray.io/u/amogkam)\
**Post date:** [July 7, 2023, 9:06pm UTC](https://discuss.ray.io/t/parallel-inference-using-cpus/11177/3 "2023-07-07T21:06:06Z")

</div>

@xzf0kgb0bqr.cev2RWU we recently added the guide to move from Torch Datasets/DataLoader to Ray Datasets! [Working with PyTorch — Ray 3.0.0.dev0](https://docs.ray.io/en/master/data/working-with-pytorch.html#migrating-from-pytorch-datasets-and-dataloaders)
