# Memory leak in ray head

**URL:** <https://discuss.ray.io/t/memory-leak-in-ray-head/4381>\
**Category:** Ray Clusters\
**Created:** [December 7, 2021, 7:57am UTC](https://discuss.ray.io/t/memory-leak-in-ray-head/4381 "2021-12-07T07:57:33Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![kubav](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kubav/32/1576_2.png) [@kubav](https://discuss.ray.io/u/kubav)\
**Post date:** [December 7, 2021, 7:57am UTC](https://discuss.ray.io/t/memory-leak-in-ray-head/4381/1 "2021-12-07T07:57:33Z")

</div>

CPU and memory usage on ray-head pod is still increasing and has to be restarted every 3 days.  
 ![ray-head](https://us1.discourse-cdn.com/flex020/uploads/ray/original/2X/5/5f98e1a3b7d804a3f56f0be177ff05cd3ff3398d.png)

I have checked that it is not caused by storing objects in cluster but it is probably caused by redis database used by GCS. Records in database are being created but they are never deleted.

I have tried to clean database manually and some of the record can be safely deleted. i.e. “DASHBOARD\*” keys are not needed and deleting them delays the time when head node needs to be restarted.

Do you know if this is ray bug or some configuration issue on our side?

---

<div class="post-metadata">

**Author:** ![zhz](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/zhz/32/12_2.png) [@zhz](https://discuss.ray.io/u/zhz)\
**Post date:** [December 10, 2021, 1:47pm UTC](https://discuss.ray.io/t/memory-leak-in-ray-head/4381/2 "2021-12-10T13:47:26Z")

</div>

Thanks @kubav ! FYI @sangcho is investigating a memory issue that could be related

---

<div class="post-metadata">

**Author:** ![Chen\_Shen](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/chen_shen/32/1486_2.png) [@Chen\_Shen](https://discuss.ray.io/u/Chen_Shen)\
**Post date:** [December 10, 2021, 7:29pm UTC](https://discuss.ray.io/t/memory-leak-in-ray-head/4381/3 "2021-12-10T19:29:43Z")

</div>

@kubav were you able to identify which process(es) is the offender? also if possible let’s move the conversation [[P0][Bug] Memory leak in ray head · Issue #21016 · ray-project/ray · GitHub](https://github.com/ray-project/ray/issues/21016)

---

<div class="post-metadata">

**Author:** ![ericl](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ericl/32/133_2.png) [@ericl](https://discuss.ray.io/u/ericl)\
**Post date:** [December 15, 2021, 11:20pm UTC](https://discuss.ray.io/t/memory-leak-in-ray-head/4381/4 "2021-12-15T23:20:39Z")

</div>

To clarify the triggering condition of this leak, is this when running multiple jobs over time? If so, that’s likely [Remote function and actor definitions are not garbage collected when drivers exit, so memory increases in cluster setting · Issue #8822 · ray-project/ray · GitHub](https://github.com/ray-project/ray/issues/8822)

Or is it something else (e.g., memory increases without new jobs being run at all?)

---

<div class="post-metadata">

**Author:** ![kubav](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kubav/32/1576_2.png) [@kubav](https://discuss.ray.io/u/kubav)\
**Post date:** [December 16, 2021, 7:56am UTC](https://discuss.ray.io/t/memory-leak-in-ray-head/4381/5 "2021-12-16T07:56:57Z")

</div>

Yes, it is the issue you linked.
