# Does Ray accelerate matrix multiplication?

**URL:** <https://discuss.ray.io/t/does-ray-accelerate-matrix-multiplication/8534>\
**Category:** Ray Core\
**Created:** [December 2, 2022, 11:55pm UTC](https://discuss.ray.io/t/does-ray-accelerate-matrix-multiplication/8534 "2022-12-02T23:55:29Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![xzf0kgb0bqr.cev2RWU](https://avatars.discourse-cdn.com/v4/letter/x/65b543/32.png) [@xzf0kgb0bqr.cev2RWU](https://discuss.ray.io/u/xzf0kgb0bqr.cev2RWU)\
**Post date:** [December 2, 2022, 11:55pm UTC](https://discuss.ray.io/t/does-ray-accelerate-matrix-multiplication/8534/1 "2022-12-02T23:55:29Z")

</div>

Low: It annoys or frustrates me for a moment.

**Question**

If I multiply two numpy arrays, will Ray accelerate the computation?

**Some Details**

I have one node with 20 CPUs. I am considering using a python script with numpy matrix multiplication (which presumably just uses one CPU) vs using Ray to do matrix multiplication.

**Context**  
I know Ray is great at reducing memory using zero-copy reads, but I would like to know if it speeds up the calculation too.

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [December 8, 2022, 4:34am UTC](https://discuss.ray.io/t/does-ray-accelerate-matrix-multiplication/8534/2 "2022-12-08T04:34:10Z")

</div>

cc @Clark_Zinzow do you know what’s the best solution for this now?

---

<div class="post-metadata">

**Author:** ![Clark\_Zinzow](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/clark_zinzow/32/445_2.png) [@Clark\_Zinzow](https://discuss.ray.io/u/Clark_Zinzow)\
**Post date:** [December 8, 2022, 6:25pm UTC](https://discuss.ray.io/t/does-ray-accelerate-matrix-multiplication/8534/3 "2022-12-08T18:25:32Z")

</div>

@xzf0kgb0bqr.cev2RWU NumPy should already use BLAS under-the-hood which will already be very fast, and I believe that you can have NumPy parallelize matrix multiplication over multiple cores by setting the `OMP_NUM_THREADS` environment variable, e.g. `OMP_NUM_THREADS=20`. I think that this will probably give you better performance than manually partitioning your matrix and doing your own parallel matrix multiplication on Ray, and if you’re having difficulties saturating your CPU resources on your machine with NumPy, then your next best bet is probably using a BLAS library directly, e.g. `scipy.linalg.blas.sgemm(...)`.

Some links:

- [Global State — NumPy v1.25.dev0 Manual](https://numpy.org/devdocs/reference/global_state.html#number-of-threads-used-for-linear-algebra)
- [python - Multiprocessing.Pool makes Numpy matrix multiplication slower - Stack Overflow](https://stackoverflow.com/a/15415690)
- [python - How to get faster code than numpy.dot for matrix multiplication? - Stack Overflow](https://stackoverflow.com/a/19839985)
- [Benjamin Johnston - Faster Matrix Multiplications in Numpy](https://www.benjaminjohnston.com.au/matmul)
