# Memory management with non-exclusive node access

**URL:** <https://discuss.ray.io/t/memory-management-with-non-exclusive-node-access/3620>\
**Category:** RLlib\
**Created:** [September 24, 2021, 12:08pm UTC](https://discuss.ray.io/t/memory-management-with-non-exclusive-node-access/3620 "2021-09-24T12:08:47Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![hartzj](https://avatars.discourse-cdn.com/v4/letter/h/a3d4f5/32.png) [@hartzj](https://discuss.ray.io/u/hartzj)\
**Post date:** [September 24, 2021, 12:08pm UTC](https://discuss.ray.io/t/memory-management-with-non-exclusive-node-access/3620/1 "2021-09-24T12:08:47Z")

</div>

Hi everyone,

I am running simple RLLIB runs using Tune on a cluster that uses the MOAB workload manager (similar to SLURM).

The jobs themselves are really simple; eg. evaluating DQN on a custom variation of Breakout, so they should not need huge amounts of RAM.

Thus, when I schedule the job on the cluster, I might ask for something like 1 GPU, 2 CPUs and 16 GB of RAM per core.

However. Ray seems to think it has exclusive access to the node:

> NodeManager:  
> InitialConfigResources: {GPU: 1.000000}, {CPU: 128.000000}, {memory: 339.085533 GiB}, {object\_store\_memory: 149.313754 GiB}, {node:10.16.46.69: 1.000000}, {accelerator\_type:T4: 1.000000}

With DQN specifically that leads to the good old `the actor died unexpectedly` error - I assume that the DQN starts to use too much memory and is thus killed by MOAB. (Interestingly, PPO, DDPG and A2C work fine.)

I tried to control the problem by specifying the resources Ray is supposed to use.  
For testing purposes, I tried to run `num_cpus=2, memory=2e8, object_store_memory=4e8` locally. I just picked those small values to see whether this works at all.

In the Dashboard, I find two workers. One of them is idle, the other one uses this much memory:

> rss:1.47GB  
> vms:6.55GB  
> shared:284.73MB  
> text:1.84MB  
> lib:0KB  
> data:1.54GB  
> dirty:0KB  
> The RSS and data grew even more over training.

How does this relate to memory settings I passed to Ray? The shared memory seems to be in the order of the specified object\_store\_memory, but overall the worker uses much more memory than it is supposed to (and I don’t assume that the redis and raylet both need that much memory).

So much questions are:

1. What exactly does the \_memory parameter for Ray.init() control?
2. How to ensure that Ray only uses the resources that it is allowed to on a shared node?

Thanks in advance!

Have a great day,  
Jan

---

<div class="post-metadata">

**Author:** ![sangcho](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sangcho/32/425_2.png) [@sangcho](https://discuss.ray.io/u/sangcho)\
**Post date:** [September 24, 2021, 9:53pm UTC](https://discuss.ray.io/t/memory-management-with-non-exclusive-node-access/3620/2 "2021-09-24T21:53:50Z")

</div>

> [@hartzj](#):
>
> \_memory

Currently, memory resources specification is just used for bookeeping, but it doesn’t enforce the memory usage of each worker. That says, imagine you have 4GB of memory for Ray, and create 2 actors with 2 GB memory each. This will schedule 2 actors on that node, but if one of actor uses 4GB of memory, that still can crash the node.

That says, to resolve the issue, you should reduce the RSS consumption of the worker (which could be application specific in this case)

---

<div class="post-metadata">

**Author:** ![gjoliver](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/gjoliver/32/1490_2.png) [@gjoliver](https://discuss.ray.io/u/gjoliver)\
**Post date:** [September 24, 2021, 10:01pm UTC](https://discuss.ray.io/t/memory-management-with-non-exclusive-node-access/3620/3 "2021-09-24T22:01:10Z")

</div>

Another tip from the RLlib team is to reduce the replay buffer size of your DQN agent, so it consumes less memory.  
In case a hard limit is not feasible.

---

<div class="post-metadata">

**Author:** ![hartzj](https://avatars.discourse-cdn.com/v4/letter/h/a3d4f5/32.png) [@hartzj](https://discuss.ray.io/u/hartzj)\
**Post date:** [October 5, 2021, 9:44am UTC](https://discuss.ray.io/t/memory-management-with-non-exclusive-node-access/3620/4 "2021-10-05T09:44:38Z")

</div>

Thank you two for the clarifications.

I haven’t specifically set DQN’s buffer size, hence it should be at the default 50k. That’s also why I was wondering that this problem occurs at all.
