# Ray/Plasma backed array

**URL:** <https://discuss.ray.io/t/ray-plasma-backed-array/1086>\
**Category:** Uncategorized\
**Created:** [March 2, 2021, 7:59pm UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086 "2021-03-02T19:59:08Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sam\_Shleifer](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sam_shleifer/32/567_2.png) [@Sam\_Shleifer](https://discuss.ray.io/u/Sam_Shleifer)\
**Post date:** [March 2, 2021, 7:59pm UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/1 "2021-03-02T19:59:08Z")

</div>

I’m trying to reduce CPU ram usage of a pytorch dataset, namely fairseq [`TokenBlockDataset`](https://github.com/pytorch/fairseq/blob/master/fairseq/data/token_block_dataset.py) by writing arrays to shared memory and then reading array **entries** from shared memory.  
Without shared memory, this would be something like:

```python
class InMemoryArray:
    
    def __init__ (self, array):
        self.array = array
        
    def read_one(self, i):
        return self.array[i]
        

```

but `self.array` would be stored fully on each GPU.

Is there a way to use Ray such that only 1 rank actually needs to write the array and all workers can read? Or some other way to use ray to better manage memory?

```auto
class RayBackedArray:

    def __init__ (self, array, rank):
        if rank == 0:
            self.refs = [ray.put(x) for x in array]
        else: self.refs = magically_gather_refs()
    def read_one(self, i):
        return ray.get(self.refs[i])

```

For reference, In raw `pyarrow.plasma`, you can invent a hashing scheme where clients can know the object id without calling `put`, but the writes (I think) are slow and block each other. I have deleted some stuff for brevity, but happy to share working slow plasma example if that’s helpful!

```auto
class PlasmaView:
    """
    New callers of this method should read https://tinyurl.com/25xx7j7y to avoid hash collisions
    """

    def __init__ (self, array, data_path: str, object_id: int):
        self.object_id = self.get_object_id(data_path, object_id) # hash the data path and a number declared by the caller
        self._client = None # Initialize lazily for pickle, (TODO: needed?)
        self.use_lock = True
        self.full_shape = array.shape
        self._partial_hash = None
        if array is None: return # another rank will write to the plasma store
        if not self.client.contains(self.object_id):
            try:
                for i, x in tqdm(list(enumerate(array))):
                    oid = self._index_to_object_id(i)
                    self.client.put(x, oid)

            except plasma.PlasmaObjectExists:
                self.msg(f"PlasmaObjectExists {oid}")
    @property
    def partial_hash(self):
        if self._partial_hash is None:
            hash = hashlib.blake2b(b'0', digest_size=20)
            for dim in self.full_shape:
                hash.update(dim.to_bytes(4, byteorder='big'))
            self._partial_hash = hash
        return self._partial_hash

    @property
    def client(self):
        if self._client is None:
            self._client = plasma.connect('/tmp/plasma', num_retries=200)
        return self._client
  
   def read_one(self, index):
        oid = self._index_to_object_id(index)
        return self._get_with_lock(oid)

    def _index_to_object_id(self, index):
        hash = self.partial_hash.copy()
        hash.update(self.int_to_bytes(int(index)))
        return plasma.ObjectID(hash.digest())

```

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 2, 2021, 8:17pm UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/2 "2021-03-02T20:17:07Z")

</div>

> [@Sam\_Shleifer](#):
>
> Is there a way to use Ray such that only 1 rank actually needs to write the array and all workers can read? Or some other way to use ray to better manage memory?
> 
> ```auto
> 
> ```

Hey @Sam_Shleifer great to see you again!

RE: the example you shared - it looks like the writes only happen upon init? Is your main problem with that example the fact that the tqdm + client.put call is too slow?

My understanding is that it seems like you have some numerical array. If you simply call `ray/plasma.put(array)`, and on other workers you call `ray/plasma.get(object_id)`, you will not duplicate the memory usage.

I have a question though; doesn’t TokenBlockDataset already leverage plasma to reduce memory usage?

---

<div class="post-metadata">

**Author:** ![Sam\_Shleifer](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sam_shleifer/32/567_2.png) [@Sam\_Shleifer](https://discuss.ray.io/u/Sam_Shleifer)\
**Post date:** [March 2, 2021, 10:06pm UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/3 "2021-03-02T22:06:52Z")

</div>

> If you simply call `ray/plasma.put(array)` , and on other workers you call `ray/plasma.get(object_id)` , you will not duplicate the memory usage.

If I call `plasma.put(array, object_id)` I only get one `object_id`, so I can’t randomly index into the array.

Great Q. The existing `PlasmaArray` called by `TokenBlockDataset` only writes to plasma when a Dataloader pickles a dataset. Then each worker’s instance reads back the full array whenever it is needed. So if `DataLoader(num_workers=0)` there are no savings. The pickling of `np.array` was very expensive (or maybe there was more pickling than actual usage). Also the plasma is only shared by multiple workers on the same cuda rank, not across cuda ranks.

Sam

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 2, 2021, 10:21pm UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/4 "2021-03-02T22:21:35Z")

</div>

> [@Sam\_Shleifer](#):
>
> If I call `plasma.put(array, object_id)` I only get one `object_id` , so I can’t randomly index into the array.

If you do `array_view = plasma.get(object_id)` after the put, then you get a read-only view of the array (assuming it is a numpy array) where you could randomly index and would not use extra worker ram, right? (Maybe I’m missing something here).

---

<div class="post-metadata">

**Author:** ![Sam\_Shleifer](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sam_shleifer/32/567_2.png) [@Sam\_Shleifer](https://discuss.ray.io/u/Sam_Shleifer)\
**Post date:** [March 3, 2021, 12:29am UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/5 "2021-03-03T00:29:32Z")

</div>

`plasma.get(object_id)` in this example returns a full deserialized `np.array`, which makes me thing that  
`plasma.get(object_id)[0]` uses lots of worker memory. Is that a bad assumption?

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 3, 2021, 1:11am UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/6 "2021-03-03T01:11:24Z")

</div>

Hypothetically it shouldn’t use a lot of worker memory (since it’s a just a readonly view backed by shared memory). Can you check top/htop to see if that’s the case?

---

<div class="post-metadata">

**Author:** ![Sam\_Shleifer](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sam_shleifer/32/567_2.png) [@Sam\_Shleifer](https://discuss.ray.io/u/Sam_Shleifer)\
**Post date:** [March 3, 2021, 1:11am UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/7 "2021-03-03T01:11:43Z")

</div>

Correction: The `np.array` returned is read only. Working on benchmarking.

---

<div class="post-metadata">

**Author:** ![Sam\_Shleifer](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sam_shleifer/32/567_2.png) [@Sam\_Shleifer](https://discuss.ray.io/u/Sam_Shleifer)\
**Post date:** [March 3, 2021, 1:57am UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/8 "2021-03-03T01:57:06Z")

</div>

My benchmarking [durbango/benchmark\_plasma\_reads.py at master · sshleifer/durbango · GitHub](https://github.com/sshleifer/durbango/blob/master/scripts/benchmark_plasma_reads.py) suggests that

- `client.put` uses a lot of worker memory that I can’t manually garbage collect
- the reading of the full array consumes 0 memory.

So I guess “read only view into shared memory” is some sick thing that I should use more of and I don’t need to do my `tqdm` write each entry.

Do you know whether plasma can handle simultaneous reads? I get segfault+ terrible traceback when too many workers read the same id from plasma. I can fix it with a lock, but that seems suboptimal.

Anyways, thanks for your help Richard, phenomenally useful. Are you on github sponsors? Alternatively, I can just owe you some retweets.

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 4, 2021, 1:57am UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/9 "2021-03-04T01:57:18Z")

</div>

> I get segfault+ terrible traceback when too many workers read the same id from plasma.

Oh uhh that seems bad. I’ve never seen this before - could you provide a traceback here?

> [@Sam\_Shleifer](#):
>
> Anyways, thanks for your help Richard, phenomenally useful. Are you on github sponsors? Alternatively, I can just owe you some retweets.

No problem! Retweets are an acceptable currency (and probably the most important one for me!)

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [March 4, 2021, 4:03am UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/10 "2021-03-04T04:03:21Z")

</div>

@Sam_Shleifer I’d love to hear what was the issue you were facing.

---

<div class="post-metadata">

**Author:** ![Sam\_Shleifer](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sam_shleifer/32/567_2.png) [@Sam\_Shleifer](https://discuss.ray.io/u/Sam_Shleifer)\
**Post date:** [March 4, 2021, 5:05pm UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/11 "2021-03-04T17:05:28Z")

</div>

Here are two tracebacks (1 for 1 GPU, another for 2 GPU)

> <https://gist.github.com/sshleifer/bd6982b3f632f1d4bcefc9feceb30b1a>

that demonstrate various errors.

I might also be leaving too many plasma connections open.

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 4, 2021, 6:06pm UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/12 "2021-03-04T18:06:56Z")

</div>

Hmm, ok. BTW, Ray has moved off of the Arrow-version of Plasma and now vendors its own, and many of these issues may be resolved there (@sangcho and @pcmoritz may have more context here)

---

<div class="post-metadata">

**Author:** ![Sam\_Shleifer](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sam_shleifer/32/567_2.png) [@Sam\_Shleifer](https://discuss.ray.io/u/Sam_Shleifer)\
**Post date:** [March 4, 2021, 9:37pm UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/13 "2021-03-04T21:37:22Z")

</div>

You can reproduce the plasma issue without fairseq or CUDA:

### Setup

- Put [this script](https://github.com/sshleifer/durbango/blob/master/scripts/plasma_demo.py) at `plasma_demo.py`
- Install Dependencies: `pip install pyarrow torch numpy`
- run `python plasma_demo.py --num-workers 2`.

Traceback [here](https://gist.github.com/sshleifer/aa956446f2d52175edcf2f85e566006d)

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 5, 2021, 12:56am UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/14 "2021-03-05T00:56:16Z")

</div>

@Sam_Shleifer could you try using the Ray version of Plasma instead?

> **[ray-project/ray](https://github.com/ray-project/ray/search?l=Python&q=plasma)**
>
> h

---

<div class="post-metadata">

**Author:** ![Sam\_Shleifer](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sam_shleifer/32/567_2.png) [@Sam\_Shleifer](https://discuss.ray.io/u/Sam_Shleifer)\
**Post date:** [March 5, 2021, 7:39pm UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/15 "2021-03-05T19:39:26Z")

</div>

I dont understand the link you sent. For ray plasma, do you just mean `ray.put` and `ray.get`, or something more direct?

- When I run this (`ray.put`, `ray.get`) [durbango/ray\_demo.py at master · sshleifer/durbango · GitHub](https://github.com/sshleifer/durbango/blob/master/scripts/ray_demo.py)  
with num-workers \>0 the non-master workers think `ray` hasn’t been initialized.

---

<div class="post-metadata">

**Author:** ![rliaw](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rliaw/32/24_2.png) [@rliaw](https://discuss.ray.io/u/rliaw)\
**Post date:** [March 8, 2021, 12:10am UTC](https://discuss.ray.io/t/ray-plasma-backed-array/1086/16 "2021-03-08T00:10:16Z")

</div>

Hmm, I meant try using the Ray vendored Plasma ([ray/services.py at 1d2136959ff8f8273dfbe5288bb58baa3ff21139 · ray-project/ray · GitHub](https://github.com/ray-project/ray/blob/1d2136959ff8f8273dfbe5288bb58baa3ff21139/python/ray/_private/services.py#L44)) instead of the standard Plasma executable
