# Can Ray Dataset be used between S3 and PyTorch?

**URL:** <https://discuss.ray.io/t/can-ray-dataset-be-used-between-s3-and-pytorch/4998>\
**Category:** Ray Data\
**Created:** [February 9, 2022, 7:26pm UTC](https://discuss.ray.io/t/can-ray-dataset-be-used-between-s3-and-pytorch/4998 "2022-02-09T19:26:36Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lacruche](https://avatars.discourse-cdn.com/v4/letter/l/ecccb3/32.png) [@Lacruche](https://discuss.ray.io/u/Lacruche)\
**Post date:** [February 9, 2022, 7:26pm UTC](https://discuss.ray.io/t/can-ray-dataset-be-used-between-s3-and-pytorch/4998/1 "2022-02-09T19:26:36Z")

</div>

Hi,

I have a PyTorch training on EC2 (plain pytorch, not Ray Train) that trains on thousands of large images (100Mb each). I’m hesitating between:

1. creating a map-style PyTorch Dataset class fetching images from S3 in the getitem (possibly with a local cache to make epoch 2 read locally),
2. reading multi-image files sequentially via an iterable dataset, powered by [Webdataset](https://github.com/webdataset/webdataset). This seems to be SOTA yet doc is a bit sparse

I’m satisfied with neither: option 1 is transparent, easy, but verbose. Option 2 is performant but more opaque (+ requires an additional tar compression step)

**I’m wondering: can Ray Datasets save my day here? And power efficient dataloading of large pictures in S3 to a PyTorch script? How?**

---

<div class="post-metadata">

**Author:** ![Clark\_Zinzow](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/clark_zinzow/32/445_2.png) [@Clark\_Zinzow](https://discuss.ray.io/u/Clark_Zinzow)\
**Post date:** [February 9, 2022, 8:35pm UTC](https://discuss.ray.io/t/can-ray-dataset-be-used-between-s3-and-pytorch/4998/2 "2022-02-09T20:35:34Z")

</div>

Hi @Lacruche, Datasets should work great here! Datasets provides an [API](https://docs.ray.io/en/master/data/package-ref.html#ray.data.read_binary_files) for reading binary files such as imagery, and has a `.to_torch()` [API](https://docs.ray.io/en/master/data/package-ref.html#ray.data.Dataset.to_torch) that yields a familiar PyTorch `IterableDataset` that you can directly consume within your PyTorch trainer:

```auto
ds_pipe = ray.data.read_binary_files("s3://some/bucket") \
    .map(lambda bytes: {"data": to_image(bytes), "label": to_label(bytes)}) \
    .repeat(num_epochs)

for epoch, ds in enumerate(ds_pipe.iter_epochs()):
    torch_ds: torch.utils.data.IterableDataset = ds.to_torch(...)
    for batch_idx, data in enumerate(torch_ds):
        # get the inputs; data is a list of [inputs, labels]
        inputs, labels = data

        # zero the parameter gradients
        optimizer.zero_grad()

        # forward + backward + optimize
        outputs = net(inputs)
        loss = criterion(outputs, labels)
        loss.backward()
        optimizer.step()

```

The imagery will be read into and held in distributed memory, and will be fed to your trainer(s) with options for prefetching, batching, etc.

---

<div class="post-metadata">

**Author:** ![amogkam](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/amogkam/32/17_2.png) [@amogkam](https://discuss.ray.io/u/amogkam)\
**Post date:** [February 10, 2022, 7:25pm UTC](https://discuss.ray.io/t/can-ray-dataset-be-used-between-s3-and-pytorch/4998/3 "2022-02-10T19:25:21Z")

</div>

Btw @Lacruche for high performant image loading, I was wondering if you have checked out [https://github.com/libffcv/ffcv](https://github.com/libffcv/ffcv)

They tout better performance than Webdataset

---

<div class="post-metadata">

**Author:** ![Lacruche](https://avatars.discourse-cdn.com/v4/letter/l/ecccb3/32.png) [@Lacruche](https://discuss.ray.io/u/Lacruche)\
**Post date:** [February 15, 2022, 10:54am UTC](https://discuss.ray.io/t/can-ray-dataset-be-used-between-s3-and-pytorch/4998/4 "2022-02-15T10:54:44Z")

</div>

interesting, I wasn’t aware of it! thanks

---

<div class="post-metadata">

**Author:** ![istranic](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/istranic/32/2171_2.png) [@istranic](https://discuss.ray.io/u/istranic)\
**Post date:** [February 17, 2022, 11:40pm UTC](https://discuss.ray.io/t/can-ray-dataset-be-used-between-s3-and-pytorch/4998/5 "2022-02-17T23:40:27Z")

</div>

Hey @Lacruche . Another alternate you could consider is [Activeloop Hub](https://github.com/activeloopai/Hub). It’s a OSS format optimized for computer vision data, and it’s primary benefit is the ability to stream data while training models. The Python API is also very simple, and you can continue to store your data in your S3 Bucket.

Full disclosure, I work for Activeloop, but we’d love if you give it a shot.
