# ValueError: buffer source array is read-only with ds.map\_batches and pandas as the batch format

**URL:** <https://discuss.ray.io/t/valueerror-buffer-source-array-is-read-only-with-ds-map-batches-and-pandas-as-the-batch-format/8418>\
**Category:** Ray Data\
**Created:** [November 23, 2022, 10:19pm UTC](https://discuss.ray.io/t/valueerror-buffer-source-array-is-read-only-with-ds-map-batches-and-pandas-as-the-batch-format/8418 "2022-11-23T22:19:19Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![harshit206](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/harshit206/32/3483_2.png) [@harshit206](https://discuss.ray.io/u/harshit206)\
**Post date:** [November 23, 2022, 10:19pm UTC](https://discuss.ray.io/t/valueerror-buffer-source-array-is-read-only-with-ds-map-batches-and-pandas-as-the-batch-format/8418/1 "2022-11-23T22:19:19Z")

</div>

Hi  
I am facing problems processing the text data using ds.map\_batches with pandas as the batch format. Getting `ValueError: buffer source array is read-only` . I have described my code below.  
I am using ray dataset api to read parquet files stored in S3 using:

> ds = ray.data.read\_parquet(“S3//PATH”)

The schema looks like this:

> schema={‘col A’: string, ‘col B’: string, ‘col C’: list\<element: string\>}

Load spacy model:

> nlp = spacy.load(“en\_core\_web\_lg”)

I am doing basic stuff like lowercasing the text and converting the text to spacy doc. My transformation function:

```auto
def transform_batch(batch: pd.DataFrame) -> pd.DataFrame:
        batch = batch.copy(deep=True)
        batch['lower_text'] = batch['text'].map(str.lower)
        batch['spacy_docs'] = batch['lower_text'].map(nlp)
        return batch

```

Finally, I do:

> transformed\_ds = ds.map\_batches(transform\_batch, batch\_format=‘pandas’)

The transform\_batch function above works fine as a standalone pandas function but using it with ray throws the error

> ValueError: buffer source array is read-only

I understand ray uses plasma store to store objects that are immutable which doesn’t allow mutating the object in place. Ray doc and ray team member from the slack community suggested creating a copy of the object as shown in the transform\_batch function. However, am facing the same error. Can someone suggest a workaround for this?

---

<div class="post-metadata">

**Author:** ![jianxiao](https://avatars.discourse-cdn.com/v4/letter/j/e480ec/32.png) [@jianxiao](https://discuss.ray.io/u/jianxiao)\
**Post date:** [November 30, 2022, 6:44pm UTC](https://discuss.ray.io/t/valueerror-buffer-source-array-is-read-only-with-ds-map-batches-and-pandas-as-the-batch-format/8418/2 "2022-11-30T18:44:15Z")

</div>

Hi @harshit206, welcome to the Ray community!

I tried to make a minimal repro based on what you are doing:

```auto
import ray
import pandas as pd

ds = ray.data.from_items([
    {
        "A": "hello",
        "B": "world",
    }
])
ds.show()

def transform_batch(batch: pd.DataFrame) -> pd.DataFrame:
    batch["C"] = "welcome"
    return batch
ds2 = ds.map_batches(transform_batch)
ds2.show()

```

As you can see, the `transform_batch` is mutating the batch. And this runs without issue:

```auto
2022-11-30 18:40:27,008	INFO worker.py:1529 -- Started a local Ray instance. View the dashboard at 127.0.0.1:8265
{'A': 'hello', 'B': 'world'}
Map_Batches: 100%|████████████████████████████████████████████████████| 1/1 [00:00<00:00, 1.43it/s]
{'A': 'hello', 'B': 'world', 'C': 'welcome'}

```

Do you mind creating a repro script? And also which Ray version are you using?

---

<div class="post-metadata">

**Author:** ![harshit206](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/harshit206/32/3483_2.png) [@harshit206](https://discuss.ray.io/u/harshit206)\
**Post date:** [November 30, 2022, 9:50pm UTC](https://discuss.ray.io/t/valueerror-buffer-source-array-is-read-only-with-ds-map-batches-and-pandas-as-the-batch-format/8418/3 "2022-11-30T21:50:08Z")

</div>

Hi @jianxiao  
Thanks for your response  
I found that problem has something to do with the spacy model (nlp object) as I executed the transform batch function without using the spacy model to see if i still get the same error but it ran successfully as you did. I wonder if spacy tries to mutate(convert to spacy doc) the `'text'` in place and hence the `ValueError: buffer source array is read-only`. This is how I solved my problem by using Ray Actors:

```auto
@ray.remote
class Textprocessor:
    def __init__ (self):

       #setup instructions for spacy model
        import pytextrank
        self.nlp = spacy.load("en_core_web_lg")
        self.nlp.max_length = 1080000  
        self.nlp.add_pipe("textrank")

  def process(self, args):
         ## some processing ###
        return

```

```auto
actors = []
for actor in range(int(ray.cluster_resources()['CPU'])):
    actors.append(Textprocessor.remote())
pool = ActorPool(actors)

```

```auto
for output in pool.map_unordered(lambda a, v: a.process.remote(v), args):
 ## processing

```

I am using Ray version 2.0.0

---

<div class="post-metadata">

**Author:** ![jianxiao](https://avatars.discourse-cdn.com/v4/letter/j/e480ec/32.png) [@jianxiao](https://discuss.ray.io/u/jianxiao)\
**Post date:** [November 30, 2022, 11:32pm UTC](https://discuss.ray.io/t/valueerror-buffer-source-array-is-read-only-with-ds-map-batches-and-pandas-as-the-batch-format/8418/4 "2022-11-30T23:32:59Z")

</div>

Note that you can use actor in `.map_batches(UDF, compute=ActorPoolStrategy(min, max), ....)` if actor is what you need. And this is a recommended way because the Dataset actorpool can autoscale dynamically between min/max.
