# None value in ds

**URL:** <https://discuss.ray.io/t/none-value-in-ds/15414>\
**Category:** Uncategorized\
**Created:** [August 1, 2024, 2:55am UTC](https://discuss.ray.io/t/none-value-in-ds/15414 "2024-08-01T02:55:55Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![duong\_phuc](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/duong_phuc/32/6371_2.png) [@duong\_phuc](https://discuss.ray.io/u/duong_phuc)\
**Post date:** [August 1, 2024, 2:55am UTC](https://discuss.ray.io/t/none-value-in-ds/15414/1 "2024-08-01T02:55:55Z")

</div>

I am using ray data map\_batches() to transform my dataset.  
My dataset have all value with the type String, and I need to specify the type before write it to parquet file.

but I have a problem, I want to do this line in the fn in map\_batches(): batch\_df[col][batch\_df[col] == ‘’] = None

after that, the ds turn to PandasBlockSchema instead of normal Dataset.  
so I can not using dataset.to\_arrow\_refs() to make the specific types of fields before write it to parquet file.

please help me

---

<div class="post-metadata">

**Author:** ![ussesjenny](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/ussesjenny/32/6369_2.png) [@ussesjenny](https://discuss.ray.io/u/ussesjenny)\
**Post date:** [August 1, 2024, 4:42am UTC](https://discuss.ray.io/t/none-value-in-ds/15414/2 "2024-08-01T04:42:54Z")

</div>

Hi everyone,

I think you should using `ray.data.map_batches()` to transform your dataset with string values and need to replace empty strings with `None` before writing to a Parquet file. When you do `batch_df[col][batch_df[col] == ''] = None`, it converts your dataset to a `PandasBlockSchema`, preventing you from using `dataset.to_arrow_refs()` to specify field types before saving. try to this may be it is useful.

Thanks

---

<div class="post-metadata">

**Author:** ![duong\_phuc](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/duong_phuc/32/6371_2.png) [@duong\_phuc](https://discuss.ray.io/u/duong_phuc)\
**Post date:** [August 1, 2024, 4:55am UTC](https://discuss.ray.io/t/none-value-in-ds/15414/3 "2024-08-01T04:55:40Z")

</div>

are you a bot :)). your answer do not make sense :))

---

<div class="post-metadata">

**Author:** ![Sam\_Chan](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sam_chan/32/5011_2.png) [@Sam\_Chan](https://discuss.ray.io/u/Sam_Chan)\
**Post date:** [August 6, 2024, 6:54am UTC](https://discuss.ray.io/t/none-value-in-ds/15414/4 "2024-08-06T06:54:31Z")

</div>

I think @ussesjenny response is on the right track; to put another way you basically need to fill in the Non empty strings before your write to Parquet while it’s still a Dataset object.

Your explicit call to batch\_df implicitly converts it to PandasBlockSchema, which, if I’m understanding correctly, is what you are trying to avoid.
