# Large Read-Only Items for NLP task

**URL:** <https://discuss.ray.io/t/large-read-only-items-for-nlp-task/675>\
**Category:** Ray Core\
**Created:** [January 31, 2021, 7:55pm UTC](https://discuss.ray.io/t/large-read-only-items-for-nlp-task/675 "2021-01-31T19:55:51Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![rmoughan](https://avatars.discourse-cdn.com/v4/letter/r/eb9ed0/32.png) [@rmoughan](https://discuss.ray.io/u/rmoughan)\
**Post date:** [January 31, 2021, 7:55pm UTC](https://discuss.ray.io/t/large-read-only-items-for-nlp-task/675/1 "2021-01-31T19:55:51Z")

</div>

Hi! I’m new to Ray and have been using it for some basic NLP tasks. I had a question about the general setup of memory in the system. Ideally, I’d like to have something like:

 ![Screen Shot 2021-01-31 at 10.45.32 AM](https://us1.discourse-cdn.com/flex020/uploads/ray/original/1X/6a95ad5e2a22090e181ac92f5d8776af80acd61a.jpeg)

There’s a text array and some other read-only items that are fairly large but only need to be read by each worker. Then each worker does some computation and local writes that I can combine as an output later on. It’s fairly similar to map-reduce, and I’ve done the GitHub example and seen the pattern document. My confusion comes with whether or not text\_array and the other read-only items are being copied over to the workers. From the readouts in !ray memory it seems like they’re not, but when I set up the futures it gives me the warning that the function has a large size when pickled, implying that they are. It would be great if they didn’t have to be copied over since they’re read-only and should be the same across all workers, so I was hoping someone could shed some light on this.

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [February 1, 2021, 7:31pm UTC](https://discuss.ray.io/t/large-read-only-items-for-nlp-task/675/2 "2021-02-01T19:31:18Z")

</div>

Hmmm the large pickle warnings are referring to issues when serializing the remote function. They tend to come accidentally including some large object in a function’s scope.

For example,

```auto
big_obj = np.zeros(1000**3)

@ray.remote
def foo():
    return np.sum(big_obj)

```

has to capture and serialize the entire object with the function, whereas we prefer for the object to not be part of the function definition.

Maybe it could help if you could share some sample code?

---

<div class="post-metadata">

**Author:** ![rmoughan](https://avatars.discourse-cdn.com/v4/letter/r/eb9ed0/32.png) [@rmoughan](https://discuss.ray.io/u/rmoughan)\
**Post date:** [February 2, 2021, 1:59am UTC](https://discuss.ray.io/t/large-read-only-items-for-nlp-task/675/3 "2021-02-02T01:59:57Z")

</div>

Hi Alex! I think you captured the heart of my issue pretty well. I created a simple example of the idea of what my code is doing below. The actual code is doing something similar to latent semantic analysis but is probably a lot harder to read. So it’s taking in a large text array and doing computation on it in parallel. You can imagine arr is the text array that ideally I don’t want to be copied over to each local worker since it’s read-only.

 ![Screen Shot 2021-02-01 at 5.52.30 PM](https://us1.discourse-cdn.com/flex020/uploads/ray/original/1X/4080003320c0bf74265973177f3798b64272246f.png)

I’m also confused why the warning says the pickled function has size ~320MB (so almost entirely because of the array arr) while !ray memory says the much more modest ~80KB. I don’t think I’m understanding the difference and the implications of the two.

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [February 2, 2021, 2:14am UTC](https://discuss.ray.io/t/large-read-only-items-for-nlp-task/675/4 "2021-02-02T02:14:19Z")

</div>

Yep, so the issue is that function definitions are stored in GCS (Redis) which is only suitable for small function definitions. In general, you should try to put the large object in the object store, then pass that object as an argument. In your case, that might look like:

```auto
arr = np.zeros(100000, 4000)

@ray.remote
def foo(arr):
    x = np.zeros(shape=(1000,))
    x[0] = arr[0][0]

arr_ref = ray.put(arr)
foo.remote(arr_ref)

```

You can find some more details here: [Ray Design Patterns - Google Docs](https://docs.google.com/document/d/167rnnDFIVRhHhK4mznEIemOtj63IOhtIPvSYaPgI4Fg/edit#heading=h.1afmymq455wu)

---

<div class="post-metadata">

**Author:** ![rmoughan](https://avatars.discourse-cdn.com/v4/letter/r/eb9ed0/32.png) [@rmoughan](https://discuss.ray.io/u/rmoughan)\
**Post date:** [February 2, 2021, 3:28am UTC](https://discuss.ray.io/t/large-read-only-items-for-nlp-task/675/5 "2021-02-02T03:28:08Z")

</div>

Ah I had tried something similar but was lacking some conceptual understanding. I see the issue now, thanks Alex! And to confirm my understanding, ray.put is putting the object in a shared memory object store and then all the workers can reference it using the returned reference id?

---

<div class="post-metadata">

**Author:** ![Alex](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/alex/32/341_2.png) [@Alex](https://discuss.ray.io/u/Alex)\
**Post date:** [February 2, 2021, 3:40am UTC](https://discuss.ray.io/t/large-read-only-items-for-nlp-task/675/6 "2021-02-02T03:40:58Z")

</div>

Yup! Sounds like you understand now.
