# Ray data read hdfs slowly and process slowly

**URL:** <https://discuss.ray.io/t/ray-data-read-hdfs-slowly-and-process-slowly/11958>\
**Category:** Ray Train\
**Created:** [August 28, 2023, 7:03am UTC](https://discuss.ray.io/t/ray-data-read-hdfs-slowly-and-process-slowly/11958 "2023-08-28T07:03:04Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![yiwei00000](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/yiwei00000/32/4985_2.png) [@yiwei00000](https://discuss.ray.io/u/yiwei00000)\
**Post date:** [August 28, 2023, 7:03am UTC](https://discuss.ray.io/t/ray-data-read-hdfs-slowly-and-process-slowly/11958/1 "2023-08-28T07:03:04Z")

</div>

Our data volume is 2 million pieces, each with 4000 columns, totaling about 80G. The data is stored on hdfs.  
Ray cluster resources 6 cores and 24GB of memory per node. There are a total of three nodes.

We use ray. data. read\_ csv() reads hdfs data, and all data is read into memory and overflowed to disk. When we perform ds. show(), ds. max (“x0”), or other operators, they all run very slowly, much slower than Spark under the same configuration. And when we run ds. max (“x0”) multiple times, the time consumed is relatively long. I can’t understand why the data is already in memory, and the time is still so long after the first time?

code eg:  
import ray  
from pyarrow import fs  
ray.init(‘auto’)

hdfs\_fs = fs.HadoopFileSystem.from\_uri(“hdfs://xx-hadoop/?user=root”)  
ds = ray.data.read\_csv(‘/data/eps-files/’, filesystem=hdfs\_fs,parallelism=2000)  
ds.max(“x0”) #I saw that for the first time, all data will be placed in the object store and overflowed to disk,  
#which takes a long time  
ds.max(“x0”) #The time consumed for the second time is similar to the first time

### Versions / Dependencies

ray 2.6.0  
hdfs 3.2.2

---

<div class="post-metadata">

**Author:** ![Jules\_Damji](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jules_damji/32/4058_2.png) [@Jules\_Damji](https://discuss.ray.io/u/Jules_Damji)\
**Post date:** [August 29, 2023, 8:38pm UTC](https://discuss.ray.io/t/ray-data-read-hdfs-slowly-and-process-slowly/11958/2 "2023-08-29T20:38:19Z")

</div>

cc: @chengsu Any idea or insights into this? Anything unique about HDFS URL here?

---

<div class="post-metadata">

**Author:** ![yiwei00000](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/yiwei00000/32/4985_2.png) [@yiwei00000](https://discuss.ray.io/u/yiwei00000)\
**Post date:** [August 31, 2023, 8:55am UTC](https://discuss.ray.io/t/ray-data-read-hdfs-slowly-and-process-slowly/11958/3 "2023-08-31T08:55:44Z")

</div>

I have solved this problem by adding ds.materialize() before the max operation

---

<div class="post-metadata">

**Author:** ![Jules\_Damji](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/jules_damji/32/4058_2.png) [@Jules\_Damji](https://discuss.ray.io/u/Jules_Damji)\
**Post date:** [August 31, 2023, 3:54pm UTC](https://discuss.ray.io/t/ray-data-read-hdfs-slowly-and-process-slowly/11958/4 "2023-08-31T15:54:16Z")

</div>

Excellent @yiwei00000. I’ll close this issue.
