# Offline RL

**URL:** https://discuss.ray.io/c/rllib/offline-rl/23.md

[Latest](https://discuss.ray.io/latest.md) · [Categories](https://discuss.ray.io/categories.md) · [Tags](https://discuss.ray.io/tags.md)

---

## [About the Offline RL category](https://discuss.ray.io/t/about-the-offline-rl-category/7742)

<div class="topic-metadata">

**Author:** [@christy](https://discuss.ray.io/u/christy)\
**Replies:** 0

</div>

---

## [Ray: Resource request cannot be scheduled — how to check CPU usage or actor resource allocation](https://discuss.ray.io/t/ray-resource-request-cannot-be-scheduled-how-to-check-cpu-usage-or-actor-resource-allocation/23264)

<div class="topic-metadata">

**Author:** [@alanyuwenche](https://discuss.ray.io/u/alanyuwenche)\
**Replies:** 1\
**Last updated:** [October 23, 2025, 4:29am UTC](https://discuss.ray.io/t/ray-resource-request-cannot-be-scheduled-how-to-check-cpu-usage-or-actor-resource-allocation/23264 "2025-10-23T04:29:00Z")

</div>

Hello there, I tried training a BC algorithm using offline data and enabled the RL module in the algorithm configuration. I ran the code on Google Colab, which only provides 2 CPUs, and encountered the following error: …

---

## [Hybrid Offline learning and PPO?](https://discuss.ray.io/t/hybrid-offline-learning-and-ppo/10713)

<div class="topic-metadata">

**Author:** [@fruzti](https://discuss.ray.io/u/fruzti)\
**Replies:** 4\
**Last updated:** [April 17, 2025, 5:54am UTC](https://discuss.ray.io/t/hybrid-offline-learning-and-ppo/10713 "2025-04-17T05:54:07Z")

</div>

Hi all, I’ve been looking around and I’m now wondering if it would make sense to combine offline rl with PPO (or another on-line rl algorithm)? I ask because in my application it is posible to have some historical data…

---

## [Load or restore offline algorithm](https://discuss.ray.io/t/load-or-restore-offline-algorithm/21883)

<div class="topic-metadata">

**Author:** [@phil](https://discuss.ray.io/u/phil)\
**Replies:** 2\
**Last updated:** [March 3, 2025, 6:35pm UTC](https://discuss.ray.io/t/load-or-restore-offline-algorithm/21883 "2025-03-03T18:35:24Z")

</div>

Low: It annoys or frustrates me for a moment. When i am training an offline algorithm and saving its checkpoint afterwards it seems not possible to load it and continue to training with offline data. The resulting error…

---

## ["Working with offlien data" tutorial: .read\_parquet loads parquet with observations as strings](https://discuss.ray.io/t/working-with-offlien-data-tutorial-read-parquet-loads-parquet-with-observations-as-strings/21893)

<div class="topic-metadata">

**Author:** [@DennisRTUB](https://discuss.ray.io/u/DennisRTUB)\
**Replies:** 0\
**Last updated:** [February 23, 2025, 7:43pm UTC](https://discuss.ray.io/t/working-with-offlien-data-tutorial-read-parquet-loads-parquet-with-observations-as-strings/21893 "2025-02-23T19:43:42Z")

</div>

I execute the following code from section “Converting tabular data to RLlib’s episode format” (link: Working with offline data — Ray 2.42.1) in the user guide " Working with offline data": from ray.rllib.core.rl\_module.…

---

## [Minimum requirement of offline data for MARWIL](https://discuss.ray.io/t/minimum-requirement-of-offline-data-for-marwil/15772)

<div class="topic-metadata">

**Author:** [@jdong9604](https://discuss.ray.io/u/jdong9604)\
**Replies:** 0\
**Last updated:** [September 11, 2024, 9:23am UTC](https://discuss.ray.io/t/minimum-requirement-of-offline-data-for-marwil/15772 "2024-09-11T09:23:07Z")

</div>

Hello, I’m very new at ray rllib. Now, I’m doing my personal project about offline RL. What I’m planning is extracting offline dataset from global optimization solvers and let MARWIL imitates its policy. After check b…

---

## [Offline rl training with custom action masking model and episodic offline data](https://discuss.ray.io/t/offline-rl-training-with-custom-action-masking-model-and-episodic-offline-data/14085)

<div class="topic-metadata">

**Author:** [@ethanhan](https://discuss.ray.io/u/ethanhan)\
**Replies:** 0\
**Last updated:** [March 20, 2024, 2:56pm UTC](https://discuss.ray.io/t/offline-rl-training-with-custom-action-masking-model-and-episodic-offline-data/14085 "2024-03-20T14:56:08Z")

</div>

Hello, I have a question regarding offline rl training with custom action masking model I have a historical dataset that I have already written into SampleBatch format and saved in json files. In my case, each line is a…

---

## [Poor performance of offline algorithms tuned examples](https://discuss.ray.io/t/poor-performance-of-offline-algorithms-tuned-examples/13871)

<div class="topic-metadata">

**Author:** [@Morphlng](https://discuss.ray.io/u/Morphlng)\
**Replies:** 0\
**Last updated:** [March 1, 2024, 7:04am UTC](https://discuss.ray.io/t/poor-performance-of-offline-algorithms-tuned-examples/13871 "2024-03-01T07:04:00Z")

</div>

Description I’m testing the offline algorithms RLlib provided, especially those using D4RL dataset. However, I’ve found that nearly all tuned\_examples provided in CQL algorithm perform really poor. For example: Hopper-…

---

## [Pre-train a model with baseline policy](https://discuss.ray.io/t/pre-train-a-model-with-baseline-policy/13645)

<div class="topic-metadata">

**Author:** [@hermmanhender](https://discuss.ray.io/u/hermmanhender)\
**Replies:** 0\
**Last updated:** [February 6, 2024, 10:57am UTC](https://discuss.ray.io/t/pre-train-a-model-with-baseline-policy/13645 "2024-02-06T10:57:21Z")

</div>

How severe does this issue affect your experience of using Ray? Low: It annoys or frustrates me for a moment. Hi! I trained a DQN Algorithm to obtain a optimal policy to control actions in a custom environment. For …

---

## [Offline reinforcement learning without environment](https://discuss.ray.io/t/offline-reinforcement-learning-without-environment/9496)

<div class="topic-metadata">

**Author:** [@fst](https://discuss.ray.io/u/fst)\
**Replies:** 3\
**Last updated:** [November 29, 2023, 1:36pm UTC](https://discuss.ray.io/t/offline-reinforcement-learning-without-environment/9496 "2023-11-29T13:36:26Z")

</div>

I have offline data that contains training patterns with: recordings of inputs from a physical machine actions of a behavior policy (=a person) operating the machine rewards for the performance of that behavior policy …

---

## [Offline RL with DQN, PPO, etc](https://discuss.ray.io/t/offline-rl-with-dqn-ppo-etc/12716)

<div class="topic-metadata">

**Author:** [@kris](https://discuss.ray.io/u/kris)\
**Replies:** 0\
**Last updated:** [November 5, 2023, 6:22pm UTC](https://discuss.ray.io/t/offline-rl-with-dqn-ppo-etc/12716 "2023-11-05T18:22:33Z")

</div>

Are there any examples of using non-standard offline RL algorithms using both offline data for training and evaluation? Second, many of the algorithms supported by ray have disappeared with 2.8.0, where can we find docum…

---

## [RL based recommender system with no simulator](https://discuss.ray.io/t/rl-based-recommender-system-with-no-simulator/12504)

<div class="topic-metadata">

**Author:** [@8rigo8](https://discuss.ray.io/u/8rigo8)\
**Replies:** 0\
**Last updated:** [October 18, 2023, 3:52pm UTC](https://discuss.ray.io/t/rl-based-recommender-system-with-no-simulator/12504 "2023-10-18T15:52:35Z")

</div>

Hi! How could I train a model (lets say SlateQ) if I don’t have a simulator environment to do online training? I’d plug the model on my website and I could register observations/actions into a replay buffer, but the re…

---

## [Error when using offline data (.json) for validation](https://discuss.ray.io/t/error-when-using-offline-data-json-for-validation/12311)

<div class="topic-metadata">

**Author:** [@kris](https://discuss.ray.io/u/kris)\
**Replies:** 0\
**Last updated:** [October 1, 2023, 10:02pm UTC](https://discuss.ray.io/t/error-when-using-offline-data-json-for-validation/12311 "2023-10-01T22:02:20Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. I have been trying to train an algorithm using offline data from a .json file while also doing off policy evaluat…

---

## [Offline Evaluation from json](https://discuss.ray.io/t/offline-evaluation-from-json/12220)

<div class="topic-metadata">

**Author:** [@kris](https://discuss.ray.io/u/kris)\
**Replies:** 0\
**Last updated:** [September 22, 2023, 10:04pm UTC](https://discuss.ray.io/t/offline-evaluation-from-json/12220 "2023-09-22T22:04:09Z")

</div>

Using offline data stored in an .json for evaluation using importance sampling and weighted importance sampling. However, I run into an value error: “ValueError: eps\_id 1000 was already passed to the peek function. Make …

---

## [Offline RL passing reward data from .json into environment](https://discuss.ray.io/t/offline-rl-passing-reward-data-from-json-into-environment/12113)

<div class="topic-metadata">

**Author:** [@kris](https://discuss.ray.io/u/kris)\
**Replies:** 3\
**Last updated:** [September 19, 2023, 11:05am UTC](https://discuss.ray.io/t/offline-rl-passing-reward-data-from-json-into-environment/12113 "2023-09-19T11:05:10Z")

</div>

I am trying to build an environment for offline RL that uses custom data. I have followed here (Working With Offline Data — Ray 2.6.3) for creating the jsonwriter for converting external experiences. And I have also been…

---

## [Offline data example](https://discuss.ray.io/t/offline-data-example/9721)

<div class="topic-metadata">

**Author:** [@asdfg](https://discuss.ray.io/u/asdfg)\
**Replies:** 4\
**Last updated:** [April 14, 2023, 9:54pm UTC](https://discuss.ray.io/t/offline-data-example/9721 "2023-04-14T21:54:23Z")

</div>

Hello, Is there a complete example of how to use Offline data using Tensorflow? The example shown below only seems to work with PyTorch. Also are there any examples using any of the offline RL algorithms such as BC or C…

---

## [Rllib with offline RL - epochs](https://discuss.ray.io/t/rllib-with-offline-rl-epochs/9631)

<div class="topic-metadata">

**Author:** [@joshml](https://discuss.ray.io/u/joshml)\
**Replies:** 1\
**Last updated:** [April 13, 2023, 10:43pm UTC](https://discuss.ray.io/t/rllib-with-offline-rl-epochs/9631 "2023-04-13T22:43:15Z")

</div>

I am new to rllib so maybe missing something very obvious. I am performing offline RL using an offline dataset from json files only. I wondered whether there is a more concise way to specify training over the entire data…

---

## [Training on multiple environment](https://discuss.ray.io/t/training-on-multiple-environment/9289)

<div class="topic-metadata">

**Author:** [@Erfan\_Asaadi](https://discuss.ray.io/u/Erfan_Asaadi)\
**Replies:** 2\
**Last updated:** [February 14, 2023, 1:05am UTC](https://discuss.ray.io/t/training-on-multiple-environment/9289 "2023-02-14T01:05:58Z")

</div>

Hi, I was wondering if anybody has a suggestion/comment on the following problem: I am trying to train an agent in multiple environment at the same time. In the other word, at the end of the episode I a need to switch …
