# Memory sharing with nested list

**URL:** <https://discuss.ray.io/t/memory-sharing-with-nested-list/2626>\
**Category:** Ray Core\
**Created:** [June 24, 2021, 11:17am UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626 "2021-06-24T11:17:51Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![hbodory](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hbodory/32/1176_2.png) [@hbodory](https://discuss.ray.io/u/hbodory)\
**Post date:** [June 24, 2021, 11:17am UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626/1 "2021-06-24T11:17:51Z")

</div>

In our ML project, we assign a large non-standard Python object (nested list) to processes in parallel. The problem is that this nested list (e.g. 3 GB) is first loaded into the Ray object store and then it will be copied for each child process, i.e. memory sharing not does not work between object store and child processes. Below you find a toy example. Our users work on Windows machines. Can you please help me to find a solution to this problem, which often leads to program abortion because we run out of memory?

```auto
# -*- coding: utf-8 -*-
"""
Created on Wed Mar 10 15:25:36 2021.

@author: HBodory
"""

import numpy as np
import ray
import time
from scipy.sparse import lil_matrix

@ray.remote
def toy_function(x, nested_list, sparse_matrix):
    """Test."""
    idx = np.array(list(range(30)))
    data = np.array([i / 1 for i in range(30)], dtype="float32")
    sparse_matrix[0, idx] = data
    sparse_matrix[1, 30:60] = sparse_matrix[0, :30]
    for i in range(1 * 10**2):
        m1 = np.random.random((100, 100))
        matrix = m1 @ np.transpose(m1)
        sol = np.linalg.inv(matrix)
        c_sparse = sparse_matrix[0, 2]
        result = x + nested_list[0][0]
    return result, sol, c_sparse

if __name__ == " __main__":

    N = 10**7

    ray.init(num_cpus=4, include_dashboard=False)

    # Create a nested list
    nested_list = [list(range(N)), list(range(2*N))]
    nested_list_id = ray.put(nested_list)

    sparse_matrix = lil_matrix((N, N), dtype=np.float32)
    sparse_matrix[0, :30] = range(30)
    sparse_matrix[1, 30:60] = sparse_matrix[0, :30]
    sparse_matrix = ray.put(sparse_matrix)

    t0 = time.time()
    result = ray.get([toy_function.remote(i, nested_list_id, sparse_matrix)
                      for i in range(10)])
    t1 = time.time() - t0
    print(f"Execution time: {t1} sec.")
    ray.shutdown()

```

---

<div class="post-metadata">

**Author:** ![architkulkarni](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/architkulkarni/32/10_2.png) [@architkulkarni](https://discuss.ray.io/u/architkulkarni)\
**Post date:** [June 25, 2021, 7:03pm UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626/2 "2021-06-25T19:03:39Z")

</div>

Hi @hbodory, thanks for the question! I’m not sure what’s causing this issue, as your code sample seems to be following the recommendation at [Tips for first-time users — Ray v2.0.0.dev0](https://docs.ray.io/en/master/auto_examples/tips-for-first-time.html#tip-3-avoid-passing-same-object-repeatedly-to-remote-tasks) to the letter. Is it possible for you to use a numpy array or collection of numpy arrays instead of a nested list? That may save on deserialization costs in the workers: [Ray Core Walkthrough — Ray v2.0.0.dev0](https://docs.ray.io/en/master/walkthrough.html#fetching-results)

@Stephanie_Wang any ideas about this?

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [June 25, 2021, 9:58pm UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626/3 "2021-06-25T21:58:49Z")

</div>

If you are talking about zero-copy read, it only works on data structure that supports pickle 5 protocol (e.g., numpy array). I believe the scipy matrix probalby uses the numpy under the hood, but for the python list, it doesn’t work (it is copied to the process memory when it is used).

---

<div class="post-metadata">

**Author:** ![hbodory](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hbodory/32/1176_2.png) [@hbodory](https://discuss.ray.io/u/hbodory)\
**Post date:** [June 26, 2021, 3:00pm UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626/4 "2021-06-26T15:00:14Z")

</div>

We thought about using objects that support the pickle 5 protocol for zero-copy read. But if we used, for example, numpy arrays instead of the single nested list object, we would have to generate tens of thousends or even hundreds of thousends numpy arrays, depending on the data size, number of trees, leaf sizes, and so on. Is there a way to handle such a large collection of numpy arrays in a similar (simple) fashion like a single nested list?

---

<div class="post-metadata">

**Author:** ![hbodory](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hbodory/32/1176_2.png) [@hbodory](https://discuss.ray.io/u/hbodory)\
**Post date:** [June 26, 2021, 3:06pm UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626/5 "2021-06-26T15:06:19Z")

</div>

Dear @architkulkarni, thanks for you recommendation regarding the use of numpy arrays. We thought about using objects that support the pickle 5 protocol for zero-copy read. But if we used, for example, numpy arrays instead of the single nested list object, we would have to generate tens of thousends or even hundreds of thousends numpy arrays, depending on the data size, number of trees, leaf sizes, and so on. Is there a way to handle such a large collection of numpy arrays in a similar (simple) fashion like a single nested list?

---

<div class="post-metadata">

**Author:** ![hbodory](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hbodory/32/1176_2.png) [@hbodory](https://discuss.ray.io/u/hbodory)\
**Post date:** [June 26, 2021, 3:08pm UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626/6 "2021-06-26T15:08:17Z")

</div>

Dear @sangcho, thank you very much for your reply. Please see my answer sent to @architkulkarni.

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [June 26, 2021, 8:58pm UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626/7 "2021-06-26T20:58:24Z")

</div>

I think if you are storing the numpy array in a nested list, only the list part is copied (and the numpy buffer itself would be zero-copy read). @suquark can you confirm this?

---

<div class="post-metadata">

**Author:** ![suquark](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/suquark/32/227_2.png) [@suquark](https://discuss.ray.io/u/suquark)\
**Post date:** [June 27, 2021, 2:17am UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626/8 "2021-06-27T02:17:36Z")

</div>

Yes, only the list part is copied

---

<div class="post-metadata">

**Author:** ![hbodory](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hbodory/32/1176_2.png) [@hbodory](https://discuss.ray.io/u/hbodory)\
**Post date:** [June 28, 2021, 9:11am UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626/9 "2021-06-28T09:11:35Z")

</div>

Dear @sangcho, I ran some small tests, you are right, storing numpy arrays in a nested list can lead to improvements in several dimensions, e.g. (i) requiring less memory for the nested list itself, (ii) saving RAM for the child processes because of zero-copy reading, and (iii) reducing the run time. Thank you very much for your valuable and helpful support.

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [June 30, 2021, 5:19am UTC](https://discuss.ray.io/t/memory-sharing-with-nested-list/2626/10 "2021-06-30T05:19:16Z")

</div>

Glad it helped!! And thanks for the verification :)!!
