# \[Core\] Ray Cluster with Shared Global Pandas Dataframe

**URL:** <https://discuss.ray.io/t/core-ray-cluster-with-shared-global-pandas-dataframe/2186>\
**Category:** Ray Core\
**Created:** [May 18, 2021, 4:02am UTC](https://discuss.ray.io/t/core-ray-cluster-with-shared-global-pandas-dataframe/2186 "2021-05-18T04:02:54Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![psytron](https://avatars.discourse-cdn.com/v4/letter/p/7feea3/32.png) [@psytron](https://discuss.ray.io/u/psytron)\
**Post date:** [May 18, 2021, 4:02am UTC](https://discuss.ray.io/t/core-ray-cluster-with-shared-global-pandas-dataframe/2186/1 "2021-05-18T04:02:55Z")

</div>

Is it possible to create a Ray configuration where 100 Actors can read+write to a global shared pandas data frame ? What is the fastest most real-time way of accomplishing this? Trying to create a system for real-time analytics / OLAP style like Druid.

Previous responses suggested:  
**Antipattern: Accessing Global Variable**

> **[Ray Design Patterns](https://docs.google.com/document/d/167rnnDFIVRhHhK4mznEIemOtj63IOhtIPvSYaPgI4Fg/edit#heading=h.eg7m6lz2y48u)**
>
> Ray Design Patterns Created: November 2020 ⭐ This is a community maintained document; suggested edits and comments are welcome! This document is a collection of common design patterns (and anti-patterns) for Ray programs. It is meant as a handbook...

this works well, but forces the entire DataFrame to be copied to each Actor which is not as fast. Is there a real-time OLAP style way of sharing data between Ray Actors ?

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [May 19, 2021, 5:32pm UTC](https://discuss.ray.io/t/core-ray-cluster-with-shared-global-pandas-dataframe/2186/2 "2021-05-19T17:32:10Z")

</div>

One way is to implement a separate actor that will hold the dataframe for you?

It also depends on what types of operations you want to support on the dataframe.

---

<div class="post-metadata">

**Author:** ![psytron](https://avatars.discourse-cdn.com/v4/letter/p/7feea3/32.png) [@psytron](https://discuss.ray.io/u/psytron)\
**Post date:** [May 19, 2021, 9:52pm UTC](https://discuss.ray.io/t/core-ray-cluster-with-shared-global-pandas-dataframe/2186/3 "2021-05-19T21:52:54Z")

</div>

Thanks @rliaw this is the solution I’m currently using ( from Ray Design Patterns doc ) but I wanted to know if there is a high performance faster method.

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [May 21, 2021, 1:40am UTC](https://discuss.ray.io/t/core-ray-cluster-with-shared-global-pandas-dataframe/2186/4 "2021-05-21T01:40:48Z")

</div>

You could:

1. shard the dataframe (and have multiple actors hold different shards)
2. Use `max_concurrency>1` for your actor, allowing multiple processes to interact with it. Note that you’ll still be GIL bottlenecked.
