# Ray Data

**URL:** https://discuss.ray.io/c/ray-data/15.md

[Latest](https://discuss.ray.io/latest.md) · [Categories](https://discuss.ray.io/categories.md) · [Tags](https://discuss.ray.io/tags.md)

---

## [About the Ray Data category](https://discuss.ray.io/t/about-the-ray-data-category/3270)

<div class="topic-metadata">

**Author:** [@rliaw](https://discuss.ray.io/u/rliaw)\
**Replies:** 1\
**Last updated:** [April 14, 2025, 11:25am UTC](https://discuss.ray.io/t/about-the-ray-data-category/3270 "2025-04-14T11:25:50Z")

</div>

For questions on large-scale data loading, preprocessing, and batch transformations in distributed pipelines. Ray Datasets are the standard way to load and exchange data in Ray applications and provide transformations su…

---

## [Ray Data AutoScaler, scale down very slow](https://discuss.ray.io/t/ray-data-autoscaler-scale-down-very-slow/23548)

<div class="topic-metadata">

**Author:** [@allendang](https://discuss.ray.io/u/allendang)\
**Replies:** 1\
**Last updated:** [May 13, 2026, 3:30am UTC](https://discuss.ray.io/t/ray-data-autoscaler-scale-down-very-slow/23548 "2026-05-13T03:30:28Z")

</div>

Hi team :waving\_hand: Question on Ray Data streaming executor’s autoscaler behavior. When an upstream op fully finishes (no more inputs, all tasks drained), its resources seem to be released gradually step-by-step rath…

---

## [Confusing behavior in setting object store size](https://discuss.ray.io/t/confusing-behavior-in-setting-object-store-size/23486)

<div class="topic-metadata">

**Author:** [@mk6](https://discuss.ray.io/u/mk6)\
**Replies:** 7\
**Last updated:** [February 24, 2026, 6:00am UTC](https://discuss.ray.io/t/confusing-behavior-in-setting-object-store-size/23486 "2026-02-24T06:00:55Z")

</div>

1. Severity of the issue: (select one) Low: Annoying but doesn’t hinder my work. 2. Environment: Ray version: 2.54.0 Python version: 3.11.10 OS: Ubuntu 22.04.3 LTS 3. What happened vs. what you expected: ray.init(…

---

## [How to interpret ray progress bar](https://discuss.ray.io/t/how-to-interpret-ray-progress-bar/23472)

<div class="topic-metadata">

**Author:** [@t-monsel](https://discuss.ray.io/u/t-monsel)\
**Replies:** 1\
**Last updated:** [February 11, 2026, 4:09pm UTC](https://discuss.ray.io/t/how-to-interpret-ray-progress-bar/23472 "2026-02-11T16:09:22Z")

</div>

I’m streaming a dataset with read\_parquet and enabled the rich\_progress bar with DatasetContext.get\_current().enable\_rich\_progress\_bars = True. I’m having logs that look like : (pid=94348) Running Dataset: train\_30\_1. …

---

## [OOM of read\_parquet for a pyspark created dataframe](https://discuss.ray.io/t/oom-of-read-parquet-for-a-pyspark-created-dataframe/23471)

<div class="topic-metadata">

**Author:** [@Sdc](https://discuss.ray.io/u/Sdc)\
**Replies:** 1\
**Last updated:** [February 11, 2026, 3:45pm UTC](https://discuss.ray.io/t/oom-of-read-parquet-for-a-pyspark-created-dataframe/23471 "2026-02-11T15:45:50Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Why is Ray Data running out of memory?](https://discuss.ray.io/t/why-is-ray-data-running-out-of-memory/23373)

<div class="topic-metadata">

**Author:** [@btian](https://discuss.ray.io/u/btian)\
**Replies:** 2\
**Last updated:** [December 23, 2025, 11:21pm UTC](https://discuss.ray.io/t/why-is-ray-data-running-out-of-memory/23373 "2025-12-23T23:21:56Z")

</div>

1. Severity of the issue: (select one) Medium: Significantly affects my productivity but can find a workaround. 2. Environment: Ray version: 2.52.1 Python version: 3.12.3 OS: Ubuntu Cloud/Infrastructure: Other libs/…

---

## [Dataset sort on lists/tuples](https://discuss.ray.io/t/dataset-sort-on-lists-tuples/23394)

<div class="topic-metadata">

**Author:** [@achimgaedke](https://discuss.ray.io/u/achimgaedke)\
**Replies:** 5\
**Last updated:** [December 21, 2025, 7:28am UTC](https://discuss.ray.io/t/dataset-sort-on-lists-tuples/23394 "2025-12-21T07:28:39Z")

</div>

1. Severity of the issue: (select one) Low: Annoying but doesn’t hinder my work. 2. Environment: Ray version: 2.52.1 Python version: 3.12/3.13 OS: MacOS/Linux 3. What happened vs. what you expected: Here’s a minim…

---

## [How can I limit the number of blocks ray data precomputes?](https://discuss.ray.io/t/how-can-i-limit-the-number-of-blocks-ray-data-precomputes/23365)

<div class="topic-metadata">

**Author:** [@pesto](https://discuss.ray.io/u/pesto)\
**Replies:** 1\
**Last updated:** [December 8, 2025, 8:45am UTC](https://discuss.ray.io/t/how-can-i-limit-the-number-of-blocks-ray-data-precomputes/23365 "2025-12-08T08:45:32Z")

</div>

1. Severity of the issue: (select one) Medium: Significantly affects my productivity but can find a workaround. 2. Environment: Ray version: 2.52.1 Python version: 3.10.19 OS: linux/ubuntu24.04 Cloud/Infrastructure:…

---

## [How to download data?](https://discuss.ray.io/t/how-to-download-data/23355)

<div class="topic-metadata">

**Author:** [@zengxinliang](https://discuss.ray.io/u/zengxinliang)\
**Replies:** 1\
**Last updated:** [December 4, 2025, 4:06pm UTC](https://discuss.ray.io/t/how-to-download-data/23355 "2025-12-04T16:06:25Z")

</div>

---

## [“Debugging ray.data”](https://discuss.ray.io/t/debugging-ray-data/23343)

<div class="topic-metadata">

**Author:** [@itswzz8](https://discuss.ray.io/u/itswzz8)\
**Replies:** 1\
**Last updated:** [November 28, 2025, 2:51pm UTC](https://discuss.ray.io/t/debugging-ray-data/23343 "2025-11-28T14:51:17Z")

</div>

1. Severity of the issue: (select one) Medium: Significantly affects my productivity but can find a workaround. 2. Environment: Ray version: Python version: OS: Cloud/Infrastructure: Other libs/tools (if relevant): …

---

## [PrepareImageUDF Error "The actor is temporarily unavailable" for ray.data.llm multimodal batch inference](https://discuss.ray.io/t/prepareimageudf-error-the-actor-is-temporarily-unavailable-for-ray-data-llm-multimodal-batch-inference/23249)

<div class="topic-metadata">

**Author:** [@Erickek](https://discuss.ray.io/u/Erickek)\
**Replies:** 4\
**Last updated:** [October 16, 2025, 5:57pm UTC](https://discuss.ray.io/t/prepareimageudf-error-the-actor-is-temporarily-unavailable-for-ray-data-llm-multimodal-batch-inference/23249 "2025-10-16T17:57:35Z")

</div>

tested: ds = ray.data.from\_items( \[ \[{ "role": "system", "content": \[ { "type": "text", "text": "You are a helpful assistant."}…

---

## [KubeRay Won't Scale to Zero: datasets\_stats\_actor Persists](https://discuss.ray.io/t/kuberay-wont-scale-to-zero-datasets-stats-actor-persists/23195)

<div class="topic-metadata">

**Author:** [@avborg](https://discuss.ray.io/u/avborg)\
**Replies:** 0\
**Last updated:** [September 29, 2025, 11:31am UTC](https://discuss.ray.io/t/kuberay-wont-scale-to-zero-datasets-stats-actor-persists/23195 "2025-09-29T11:31:18Z")

</div>

1. Severity of the issue: Medium: Significantly affects my productivity but can find a workaround. 2. Environment: Ray version: 2.48.0 Python version: 3.12.9 OS: ubuntu (from standard kuberay worker image rayproject/…

---

## [Ray parquet data streaming causing massive spill and storage issues](https://discuss.ray.io/t/ray-parquet-data-streaming-causing-massive-spill-and-storage-issues/23107)

<div class="topic-metadata">

**Author:** [@kohlisimranjitAD](https://discuss.ray.io/u/kohlisimranjitAD)\
**Replies:** 0\
**Last updated:** [September 15, 2025, 11:15pm UTC](https://discuss.ray.io/t/ray-parquet-data-streaming-causing-massive-spill-and-storage-issues/23107 "2025-09-15T23:15:48Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Ray data concurrency \> 1 fails](https://discuss.ray.io/t/ray-data-concurrency-1-fails/23082)

<div class="topic-metadata">

**Author:** [@anaykulkarni](https://discuss.ray.io/u/anaykulkarni)\
**Replies:** 1\
**Last updated:** [September 8, 2025, 6:43pm UTC](https://discuss.ray.io/t/ray-data-concurrency-1-fails/23082 "2025-09-08T18:43:58Z")

</div>

Hi guys, I’m new to Ray and I’m facing issues with concurrency while using ray data’s map\_batches. I’m trying to launch 8 vLLM instances in parallel. At first it tries to launch 8 actors, then the model loading starts of…

---

## [Where are the results produced during execution pipeline](https://discuss.ray.io/t/where-are-the-results-produced-during-execution-pipeline/23045)

<div class="topic-metadata">

**Author:** [@Leslie-Chung](https://discuss.ray.io/u/Leslie-Chung)\
**Replies:** 0\
**Last updated:** [August 26, 2025, 12:28pm UTC](https://discuss.ray.io/t/where-are-the-results-produced-during-execution-pipeline/23045 "2025-08-26T12:28:48Z")

</div>

In a streaming execution pipeline like read\_text → map → iter\_rows, will the output blocks produced by read\_text be stored in the Ray Object Store, or are they only materialized temporarily during the iteration?

---

## [Multi-model batch inference - problem with scaling of actors](https://discuss.ray.io/t/multi-model-batch-inference-problem-with-scaling-of-actors/22998)

<div class="topic-metadata">

**Author:** [@krzwaraksa](https://discuss.ray.io/u/krzwaraksa)\
**Replies:** 2\
**Last updated:** [August 19, 2025, 7:36am UTC](https://discuss.ray.io/t/multi-model-batch-inference-problem-with-scaling-of-actors/22998 "2025-08-19T07:36:51Z")

</div>

Hello everyone, we’re building a multi-model batch inference pipeline and so far, we’re enjoying Ray a lot! However, I’ve noticed one problem with auto scaling when using stateful maps (with Actors). Our pipeline looks a…

---

## [LLM Batch Model loading using runai\_streamer is very slow](https://discuss.ray.io/t/llm-batch-model-loading-using-runai-streamer-is-very-slow/22922)

<div class="topic-metadata">

**Author:** [@sologymming](https://discuss.ray.io/u/sologymming)\
**Replies:** 1\
**Last updated:** [August 13, 2025, 6:14pm UTC](https://discuss.ray.io/t/llm-batch-model-loading-using-runai-streamer-is-very-slow/22922 "2025-08-13T18:14:38Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. \[V\] High: Com…

---

## [Problem regarding doing a vector search on a Lance Dataset generated using ray.data.Dataset.write\_lance](https://discuss.ray.io/t/problem-regarding-doing-a-vector-search-on-a-lance-dataset-generated-using-ray-data-dataset-write-lance/22978)

<div class="topic-metadata">

**Author:** [@king4face](https://discuss.ray.io/u/king4face)\
**Replies:** 0\
**Last updated:** [August 11, 2025, 8:32am UTC](https://discuss.ray.io/t/problem-regarding-doing-a-vector-search-on-a-lance-dataset-generated-using-ray-data-dataset-write-lance/22978 "2025-08-11T08:32:01Z")

</div>

1. Severity of the issue: High: Completely blocks me. 2. Environment: Ray version: 2.48.0 Python version: 3.10.12 OS: Ubuntu 22.04.5 LTS Other libs/tools (if relevant): pylance 0.32.0 3. What happened vs. what you e…

---

## [Can multiple Ray Data pipeline steps share the same large model instance for inference?](https://discuss.ray.io/t/can-multiple-ray-data-pipeline-steps-share-the-same-large-model-instance-for-inference/22876)

<div class="topic-metadata">

**Author:** [@donglin\_hao](https://discuss.ray.io/u/donglin_hao)\
**Replies:** 3\
**Last updated:** [August 6, 2025, 10:14pm UTC](https://discuss.ray.io/t/can-multiple-ray-data-pipeline-steps-share-the-same-large-model-instance-for-inference/22876 "2025-08-06T22:14:06Z")

</div>

Hi, In a Ray Data processing pipeline, I have multiple steps that all need to call the same large model for inference. Although their preprocessing and postprocessing may differ, the inference stage itself is exactly th…

---

## [XGBoost + large data with Ray Tune: How?](https://discuss.ray.io/t/xgboost-large-data-with-ray-tune-how/22907)

<div class="topic-metadata">

**Author:** [@Ivan](https://discuss.ray.io/u/Ivan)\
**Replies:** 0\
**Last updated:** [July 29, 2025, 11:50am UTC](https://discuss.ray.io/t/xgboost-large-data-with-ray-tune-how/22907 "2025-07-29T11:50:34Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Adjusting number of pending tasks per actor?](https://discuss.ray.io/t/adjusting-number-of-pending-tasks-per-actor/22838)

<div class="topic-metadata">

**Author:** [@kbumsik](https://discuss.ray.io/u/kbumsik)\
**Replies:** 0\
**Last updated:** [July 14, 2025, 7:06am UTC](https://discuss.ray.io/t/adjusting-number-of-pending-tasks-per-actor/22838 "2025-07-14T07:06:36Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Access aws s3 for vllm v0.9+](https://discuss.ray.io/t/access-aws-s3-for-vllm-v0-9/22755)

<div class="topic-metadata">

**Author:** [@cnmdestroyer](https://discuss.ray.io/u/cnmdestroyer)\
**Replies:** 2\
**Last updated:** [July 10, 2025, 11:29pm UTC](https://discuss.ray.io/t/access-aws-s3-for-vllm-v0-9/22755 "2025-07-10T23:29:28Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [Ffmpeg hanging in a remote task when trying to convert a .MP4 file to .H265 file](https://discuss.ray.io/t/ffmpeg-hanging-in-a-remote-task-when-trying-to-convert-a-mp4-file-to-h265-file/22805)

<div class="topic-metadata">

**Author:** [@rk14](https://discuss.ray.io/u/rk14)\
**Replies:** 0\
**Last updated:** [July 5, 2025, 5:33pm UTC](https://discuss.ray.io/t/ffmpeg-hanging-in-a-remote-task-when-trying-to-convert-a-mp4-file-to-h265-file/22805 "2025-07-05T17:33:34Z")

</div>

Hi all! I’m new to Ray and I’ve stumbled upon an issue where my process is hanging when running an ffmpeg command to convert a .MP4 file into a .H265 file in my remote task. However, I’m able to convert a .H264 file into…

---

## [Ray data read\_text calls read all of input hogging memory and spilling](https://discuss.ray.io/t/ray-data-read-text-calls-read-all-of-input-hogging-memory-and-spilling/22667)

<div class="topic-metadata">

**Author:** [@vgill](https://discuss.ray.io/u/vgill)\
**Replies:** 1\
**Last updated:** [June 17, 2025, 11:05pm UTC](https://discuss.ray.io/t/ray-data-read-text-calls-read-all-of-input-hogging-memory-and-spilling/22667 "2025-06-17T23:05:11Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. \[X \] High: Co…

---

## [Adding model compilation like TensorRT in offline inference](https://discuss.ray.io/t/adding-model-compilation-like-tensorrt-in-offline-inference/22665)

<div class="topic-metadata">

**Author:** [@Monil](https://discuss.ray.io/u/Monil)\
**Replies:** 1\
**Last updated:** [June 17, 2025, 11:03pm UTC](https://discuss.ray.io/t/adding-model-compilation-like-tensorrt-in-offline-inference/22665 "2025-06-17T23:03:02Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. Low: Annoying but doesn’t hinder my work. Medium: Significantly affects my productivity but can find a workaround. High: Comple…

---

## [ObjectFetchTimedOutError](https://discuss.ray.io/t/objectfetchtimedouterror/22618)

<div class="topic-metadata">

**Author:** [@Rags](https://discuss.ray.io/u/Rags)\
**Replies:** 1\
**Last updated:** [June 10, 2025, 12:53am UTC](https://discuss.ray.io/t/objectfetchtimedouterror/22618 "2025-06-10T00:53:50Z")

</div>

1. Severity of the issue: (select one) High: Completely blocks me. 2. Environment: Ray version: 2.23 Python version: 3.10 OS: Cloud/Infrastructure: managed k8s Other libs/tools (if relevant): 3. What happened vs. w…

---

## [Ray data \`ReadParquet-\>SplitBlocks(2)\` shows failure even though the entire ray job is successful](https://discuss.ray.io/t/ray-data-readparquet-splitblocks-2-shows-failure-even-though-the-entire-ray-job-is-successful/22567)

<div class="topic-metadata">

**Author:** [@brian-tang](https://discuss.ray.io/u/brian-tang)\
**Replies:** 1\
**Last updated:** [June 3, 2025, 11:22pm UTC](https://discuss.ray.io/t/ray-data-readparquet-splitblocks-2-shows-failure-even-though-the-entire-ray-job-is-successful/22567 "2025-06-03T23:22:52Z")

</div>

Medium: Significantly affects my productivity but can find a workaround. 2. Environment: Ray version: 2.41.0 Python version: 3.11 OS: AL2 Cloud/Infrastructure: AWS EKS 1.30 Other libs/tools (if relevant): 3. What hap…

---

## [How to use the same set of actors in multiple non-adjacent processing steps](https://discuss.ray.io/t/how-to-use-the-same-set-of-actors-in-multiple-non-adjacent-processing-steps/22580)

<div class="topic-metadata">

**Author:** [@mk6](https://discuss.ray.io/u/mk6)\
**Replies:** 0\
**Last updated:** [May 30, 2025, 7:06am UTC](https://discuss.ray.io/t/how-to-use-the-same-set-of-actors-in-multiple-non-adjacent-processing-steps/22580 "2025-05-30T07:06:01Z")

</div>

I’m working on a Ray Data pipeline where I need to utilize a pool of Ray actors (specifically, LLMs, each requiring a GPU) in multiple, non-adjacent processing steps. My pipeline looks like this: LLM Generation: Use t…

---

## [Why UDF time is larger than Remote wall time with concurrency=1?](https://discuss.ray.io/t/why-udf-time-is-larger-than-remote-wall-time-with-concurrency-1/22576)

<div class="topic-metadata">

**Author:** [@kostochkod](https://discuss.ray.io/u/kostochkod)\
**Replies:** 0\
**Last updated:** [May 29, 2025, 10:22am UTC](https://discuss.ray.io/t/why-udf-time-is-larger-than-remote-wall-time-with-concurrency-1/22576 "2025-05-29T10:22:01Z")

</div>

1. Severity of the issue: (select one) None: I’m just curious or want clarification. I try to understand Dataset.stats() output and curious how UDF time is calculated. From documentation: UDF time: The UDF time is …

---

## [Does map\_batches avoid saturating the inference engine?](https://discuss.ray.io/t/does-map-batches-avoid-saturating-the-inference-engine/22538)

<div class="topic-metadata">

**Author:** [@shevateng0](https://discuss.ray.io/u/shevateng0)\
**Replies:** 1\
**Last updated:** [May 25, 2025, 11:30pm UTC](https://discuss.ray.io/t/does-map-batches-avoid-saturating-the-inference-engine/22538 "2025-05-25T23:30:12Z")

</div>

I am experimenting with Ray Data for batch inference. Ray runs in standalone mode as I only use one GPU node. map\_batches() is used to perform batch inference. I want to know if batch inference via Ray will slow down th…

[Next page](https://discuss.ray.io/c/ray-data/15.md?page=1)
