# Ray Column With Custom Python Dataclass Type

**URL:** <https://discuss.ray.io/t/ray-column-with-custom-python-dataclass-type/14092>\
**Category:** Ray Data\
**Created:** [March 21, 2024, 12:32am UTC](https://discuss.ray.io/t/ray-column-with-custom-python-dataclass-type/14092 "2024-03-21T00:32:10Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![kaylahardie](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kaylahardie/32/5839_2.png) [@kaylahardie](https://discuss.ray.io/u/kaylahardie)\
**Post date:** [March 21, 2024, 12:32am UTC](https://discuss.ray.io/t/ray-column-with-custom-python-dataclass-type/14092/1 "2024-03-21T00:32:10Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- Low: It annoys or frustrates me for a moment.

Can you store custom python dataclasses in ray datasets? If not, what is the recommended way to create a column with structured data?

For example:  
@dataclass  
class ImageMetadata:  
x\_resolution: float  
y\_resolution: float  
file\_name: str

@dataclass  
class ImageTile:  
tile\_values: np.array  
tile\_metadata: ImageTileMetadata

I’d like to have ImageTile be a column type in my ray dataset.

---

<div class="post-metadata">

**Author:** ![raulchen](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/raulchen/32/330_2.png) [@raulchen](https://discuss.ray.io/u/raulchen)\
**Post date:** [April 9, 2024, 7:55pm UTC](https://discuss.ray.io/t/ray-column-with-custom-python-dataclass-type/14092/2 "2024-04-09T19:55:10Z")

</div>

you can do that, but more recommended to use numpy ndarrays.

---

<div class="post-metadata">

**Author:** ![kaylahardie](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kaylahardie/32/5839_2.png) [@kaylahardie](https://discuss.ray.io/u/kaylahardie)\
**Post date:** [April 15, 2024, 7:18pm UTC](https://discuss.ray.io/t/ray-column-with-custom-python-dataclass-type/14092/3 "2024-04-15T19:18:12Z")

</div>

Thanks! Is there any advice on how to make the numpy arrays more readable? For example, if I had an array of [x\_resolution, y\_resolution] it would be error prone to remember that the x\_resolution is in position 0.

Using numpy arrays might also make it hard for me to show relationships between the data in the columns. For example, I have an image and then I apply a flat\_map and create tiles it might be nice to represent that relationship (an image has tiles) in the dataset. Do you have any thoughts on how to do that in a readable way?

---

<div class="post-metadata">

**Author:** ![Ryan\_Avery](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ryan_avery/32/7281_2.png) [@Ryan\_Avery](https://discuss.ray.io/u/Ryan_Avery)\
**Post date:** [May 22, 2025, 11:04pm UTC](https://discuss.ray.io/t/ray-column-with-custom-python-dataclass-type/14092/4 "2025-05-22T23:04:38Z")

</div>

This is a great question and would be great if Ray supported serializing dataclasses zero copy if the dataclass fields conform to pyarrow types. I would like to have labels for ndarrays and slices for writing these arrays out to a single array store.
