# RLlib

**URL:** https://discuss.ray.io/c/rllib/7.md

[Latest](https://discuss.ray.io/latest.md) · [Categories](https://discuss.ray.io/categories.md) · [Tags](https://discuss.ray.io/tags.md)

---

## [How to pretrain a model with behavior cloning](https://discuss.ray.io/t/how-to-pretrain-a-model-with-behavior-cloning/278)

<div class="topic-metadata">

**Author:** [@BrunoBSM](https://discuss.ray.io/u/BrunoBSM)\
**Replies:** 14\
**Last updated:** [December 5, 2023, 3:28pm UTC](https://discuss.ray.io/t/how-to-pretrain-a-model-with-behavior-cloning/278 "2023-12-05T15:28:27Z")

</div>

I asked this question on Ray’s Slack channel and now I am transcribing the Thread here so it’s available to everyone. Thanks to @sven1977 and @rliaw for answering my question. Bruno Brandao: Does anyone know how can I…

---

## [Reproducible training - setting seeds for all workers / environments](https://discuss.ray.io/t/reproducible-training-setting-seeds-for-all-workers-environments/1051)

<div class="topic-metadata">

**Author:** [@Lauritowal](https://discuss.ray.io/u/Lauritowal)\
**Replies:** 18\
**Last updated:** [May 6, 2023, 11:46pm UTC](https://discuss.ray.io/t/reproducible-training-setting-seeds-for-all-workers-environments/1051 "2023-05-06T23:46:33Z")

</div>

Hi! How can I best set the seed of my environments while training an RL agent? I found the following answer on stack overflow Obtaining different set of configs across multiple calls in ray tune - Stack Overflow and he…

---

## [RLLIB not working with Tune with sample batch input](https://discuss.ray.io/t/rllib-not-working-with-tune-with-sample-batch-input/7134)

<div class="topic-metadata">

**Author:** [@Jason\_Weinberg](https://discuss.ray.io/u/Jason_Weinberg)\
**Replies:** 25\
**Last updated:** [October 4, 2022, 1:58pm UTC](https://discuss.ray.io/t/rllib-not-working-with-tune-with-sample-batch-input/7134 "2022-10-04T13:58:54Z")

</div>

My issue is that RLLIB is not properly logging metrics back to tune when training a model using sample batches. My config has no ENV parameter set, it only uses input sample batches. When metrics come back to tune the ep…

---

## [ValueError: Expected parameter logits (...) to satisfy the constraint IndependentConstraint(Real(), 1)](https://discuss.ray.io/t/valueerror-expected-parameter-logits-to-satisfy-the-constraint-independentconstraint-real-1/4108)

<div class="topic-metadata">

**Author:** [@LukasNothhelfer](https://discuss.ray.io/u/LukasNothhelfer)\
**Replies:** 38\
**Last updated:** [October 14, 2024, 10:14pm UTC](https://discuss.ray.io/t/valueerror-expected-parameter-logits-to-satisfy-the-constraint-independentconstraint-real-1/4108 "2024-10-14T22:14:17Z")

</div>

Today I got a very strange problem with Ray RLlib (ray version 2.0.0dev0). The code that I am running all the time is throwing an error since today but it worked all the time. I get this error: File "/opt/conda/lib/pyth…

---

## [\_winapi.CreateProcess(executable, args, FileNotFoundError: \[WinError 2\]](https://discuss.ray.io/t/winapi-createprocess-executable-args-filenotfounderror-winerror-2/5943)

<div class="topic-metadata">

**Author:** [@Denys\_Ashikhin](https://discuss.ray.io/u/Denys_Ashikhin)\
**Replies:** 9\
**Last updated:** [May 14, 2022, 4:08pm UTC](https://discuss.ray.io/t/winapi-createprocess-executable-args-filenotfounderror-winerror-2/5943 "2022-05-14T16:08:59Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. Hi all, I am getting the following error: policy\_server.py -ip=localhost -checkpoint=brawl Traceback (most rec…

---

## [RLlib Office Hours - Now open for signup](https://discuss.ray.io/t/rllib-office-hours-now-open-for-signup/3153)

<div class="topic-metadata">

**Author:** [@sven1977](https://discuss.ray.io/u/sven1977)\
**Replies:** 13\
**Last updated:** [July 1, 2025, 6:41pm UTC](https://discuss.ray.io/t/rllib-office-hours-now-open-for-signup/3153 "2025-07-01T18:41:52Z")

</div>

Hey everyone, we would like to start offering “office hours” for RLlib, which will basically be 30min video chats with @sven1977 to debug/unblock issues together. If you are interested, please fill out this form here (~…

---

## [\[RLlib\] Impossible actions](https://discuss.ray.io/t/rllib-impossible-actions/890)

<div class="topic-metadata">

**Author:** [@Jogima-cyber](https://discuss.ray.io/u/Jogima-cyber)\
**Replies:** 12\
**Last updated:** [May 11, 2022, 12:19pm UTC](https://discuss.ray.io/t/rllib-impossible-actions/890 "2022-05-11T12:19:28Z")

</div>

Hi there, I’d like to set in a multienv impossible actions. I’ve read the doc (https://docs.ray.io/en/master/rllib-models.html#variable-length-parametric-action-spaces) but I don’t understand the purposes of avail\_action…

---

## [Ray for Rapberry Pi, is possible?](https://discuss.ray.io/t/ray-for-rapberry-pi-is-possible/451)

<div class="topic-metadata">

**Author:** [@Peter\_Pirog](https://discuss.ray.io/u/Peter_Pirog)\
**Replies:** 30\
**Last updated:** [February 23, 2021, 9:04pm UTC](https://discuss.ray.io/t/ray-for-rapberry-pi-is-possible/451 "2021-02-23T21:04:36Z")

</div>

I would like to create regulator for measurement system based on Raspberry Pi controllers. I wonder if is possible to use Ubuntu PC as trainer and raspberry Pi 3 or 4 as workers. I tested that its possible to install te…

---

## [Multi-Agent Training with Different Algorithms](https://discuss.ray.io/t/multi-agent-training-with-different-algorithms/6503)

<div class="topic-metadata">

**Author:** [@mgerstgrasser](https://discuss.ray.io/u/mgerstgrasser)\
**Replies:** 24\
**Last updated:** [October 11, 2022, 2:46pm UTC](https://discuss.ray.io/t/multi-agent-training-with-different-algorithms/6503 "2022-10-11T14:46:03Z")

</div>

How severe does this issue affect your experience of using Ray? None: Just asking a question out of curiosity (Severity is both none and high - at the moment I’m just curious, but depending on the answer this might c…

---

## [Observation\_space not provided in PolicySpec](https://discuss.ray.io/t/observation-space-not-provided-in-policyspec/6501)

<div class="topic-metadata">

**Author:** [@evo11x](https://discuss.ray.io/u/evo11x)\
**Replies:** 21\
**Last updated:** [February 7, 2023, 4:14pm UTC](https://discuss.ray.io/t/observation-space-not-provided-in-policyspec/6501 "2023-02-07T16:14:42Z")

</div>

I don’t know why I am getting this error when running Tune with IMPALA with single agent custom env, if I run the trainer without tune, it runs for a few minutes then it crashes. ValueError: observation\_space not provid…

---

## [\[RLlib\] Visualise custom environment](https://discuss.ray.io/t/rllib-visualise-custom-environment/778)

<div class="topic-metadata">

**Author:** [@vakker00](https://discuss.ray.io/u/vakker00)\
**Replies:** 18\
**Last updated:** [March 30, 2021, 9:55am UTC](https://discuss.ray.io/t/rllib-visualise-custom-environment/778 "2021-03-30T09:55:57Z")

</div>

I’m trying to record the observations from a custom env. I implemented the render method for my environment that just returns an RGB array. If I set monitor: True then Gym complains that: WARN: Trying to monitor an env…

---

## [Issues reproducing stable-baselines3 PPO performance with rllib](https://discuss.ray.io/t/issues-reproducing-stable-baselines3-ppo-performance-with-rllib/3028)

<div class="topic-metadata">

**Author:** [@mjlbach](https://discuss.ray.io/u/mjlbach)\
**Replies:** 14\
**Last updated:** [March 16, 2022, 4:17pm UTC](https://discuss.ray.io/t/issues-reproducing-stable-baselines3-ppo-performance-with-rllib/3028 "2022-03-16T16:17:16Z")

</div>

Hi all, SVL has recently launched a new challenge for embodied, multi-task learning in home environments called BEHAVIOR, as part of this we are recommending users start with ray or stable-baselines3 to get quickly spun…

---

## [Resume=True fails without useful error message](https://discuss.ray.io/t/resume-true-fails-without-useful-error-message/7506)

<div class="topic-metadata">

**Author:** [@hridayns](https://discuss.ray.io/u/hridayns)\
**Replies:** 31\
**Last updated:** [September 26, 2022, 4:27pm UTC](https://discuss.ray.io/t/resume-true-fails-without-useful-error-message/7506 "2022-09-26T16:27:48Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. I’m so close to finishing my training (Ray 2.0.0) and I ran out of disk space (apparently), so I cleared some out…

---

## [Help debugging a memory leak in rllib](https://discuss.ray.io/t/help-debugging-a-memory-leak-in-rllib/2100)

<div class="topic-metadata">

**Author:** [@Bam4d](https://discuss.ray.io/u/Bam4d)\
**Replies:** 21\
**Last updated:** [September 25, 2022, 7:07am UTC](https://discuss.ray.io/t/help-debugging-a-memory-leak-in-rllib/2100 "2022-09-25T07:07:08Z")

</div>

I’m trying to debug a very slow memory leak in rllib that occurs when i am using IMPALA + multi-agent. I cannot find any leak using tools like tracemalloc so I dont think the memory issue is in python. ray memory also …

---

## [Board game self-play PPO](https://discuss.ray.io/t/board-game-self-play-ppo/1425)

<div class="topic-metadata">

**Author:** [@ColdFrenzy](https://discuss.ray.io/u/ColdFrenzy)\
**Replies:** 15\
**Last updated:** [May 4, 2021, 9:47am UTC](https://discuss.ray.io/t/board-game-self-play-ppo/1425 "2021-05-04T09:47:47Z")

</div>

Hi, I’ve implemented a multiagent version of connect 4 and i’m trying to train it with PPO through self-play. At each turn the environment returns the observation and reward for the player that will move next. The obse…

---

## [RNN L2 weights regularization](https://discuss.ray.io/t/rnn-l2-weights-regularization/2582)

<div class="topic-metadata">

**Author:** [@mg64ve](https://discuss.ray.io/u/mg64ve)\
**Replies:** 41\
**Last updated:** [July 5, 2021, 12:59pm UTC](https://discuss.ray.io/t/rnn-l2-weights-regularization/2582 "2021-07-05T12:59:36Z")

</div>

Is there any plan for this feature?

---

## [Issue creating custom action mask enviorment](https://discuss.ray.io/t/issue-creating-custom-action-mask-enviorment/4940)

<div class="topic-metadata">

**Author:** [@Ramie\_Yahya](https://discuss.ray.io/u/Ramie_Yahya)\
**Replies:** 14\
**Last updated:** [October 11, 2023, 5:21pm UTC](https://discuss.ray.io/t/issue-creating-custom-action-mask-enviorment/4940 "2023-10-11T17:21:53Z")

</div>

Hi all, I’m trying to set up an action masking environment by following the examples on GitHub. from gym.spaces import Dict from gym import spaces from ray.rllib.models.tf.fcnet import FullyConnectedNetwork from ray.r…

---

## [Use Policy\_Trainer with TensorBoard](https://discuss.ray.io/t/use-policy-trainer-with-tensorboard/4097)

<div class="topic-metadata">

**Author:** [@Denys\_Ashikhin](https://discuss.ray.io/u/Denys_Ashikhin)\
**Replies:** 33\
**Last updated:** [November 13, 2021, 5:48pm UTC](https://discuss.ray.io/t/use-policy-trainer-with-tensorboard/4097 "2021-11-13T17:48:55Z")

</div>

Hi All, I am using a policy client + server for training purposes, however, I can’t figure out how to have tensorboard display any information for the training runs? Is there a parameter I need to pass? Moreover, can I…

---

## [Deploying a learned policy under "explore=False / True"](https://discuss.ray.io/t/deploying-a-learned-policy-under-explore-false-true/4394)

<div class="topic-metadata">

**Author:** [@klausk55](https://discuss.ray.io/u/klausk55)\
**Replies:** 9\
**Last updated:** [March 17, 2022, 2:55pm UTC](https://discuss.ray.io/t/deploying-a-learned-policy-under-explore-false-true/4394 "2022-03-17T14:55:34Z")

</div>

Hey folks, I have problems in understanding the following important note (it’s a comment in the RLlib’s Trainer config section “Evaluation Settings”): IMPORTANT NOTE: Policy gradient algorithms are able to find the op…

---

## [Removing Algorithms from RLlib](https://discuss.ray.io/t/removing-algorithms-from-rllib/6747)

<div class="topic-metadata">

**Author:** [@avnishn](https://discuss.ray.io/u/avnishn)\
**Replies:** 10\
**Last updated:** [July 22, 2022, 8:40am UTC](https://discuss.ray.io/t/removing-algorithms-from-rllib/6747 "2022-07-22T08:40:53Z")

</div>

Hi RLlib Community, The RLlib team is discussing the idea of removing some algorithms from the library so that we can better focus on improving the quality of our code base as a whole. Doing so will reduce our maintenan…

---

## [RLlib, PyTorch and Mac M1 GPUs: No available node types can fulfill resource request](https://discuss.ray.io/t/rllib-pytorch-and-mac-m1-gpus-no-available-node-types-can-fulfill-resource-request/6769)

<div class="topic-metadata">

**Author:** [@robfitzgerald](https://discuss.ray.io/u/robfitzgerald)\
**Replies:** 11\
**Last updated:** [February 29, 2024, 11:57am UTC](https://discuss.ray.io/t/rllib-pytorch-and-mac-m1-gpus-no-available-node-types-can-fulfill-resource-request/6769 "2024-02-29T11:57:04Z")

</div>

How severe does this issue affect your experience of using Ray? Medium: It contributes to significant difficulty to complete my task, but I can work around it. Hello Ray community! A year ago I began experimenting w/…

---

## [Compute\_actions for Trajectory API](https://discuss.ray.io/t/compute-actions-for-trajectory-api/2781)

<div class="topic-metadata">

**Author:** [@mg64ve](https://discuss.ray.io/u/mg64ve)\
**Replies:** 11\
**Last updated:** [February 10, 2022, 1:00am UTC](https://discuss.ray.io/t/compute-actions-for-trajectory-api/2781 "2022-02-10T01:00:13Z")

</div>

Hello, consider the following documentation: https://docs.ray.io/en/master/rllib-training.html#computing-actions There is no mention this does not apply to models using Trajectory API. Now if you consider the followin…

---

## [Issue with custom LSTMs](https://discuss.ray.io/t/issue-with-custom-lstms/5188)

<div class="topic-metadata">

**Author:** [@Dylan\_Miller](https://discuss.ray.io/u/Dylan_Miller)\
**Replies:** 34\
**Last updated:** [February 26, 2023, 5:54pm UTC](https://discuss.ray.io/t/issue-with-custom-lstms/5188 "2023-02-26T17:54:43Z")

</div>

if state\_outs: B = 4 # For RNNs, have B=4, T=\[depends on sample\_batch\_size\] i = 0 while "state\_in\_{}".format(i) in postprocessed\_batch: postprocessed\_batch\["state\_…

---

## [Meaning of episode\_reward\_mean](https://discuss.ray.io/t/meaning-of-episode-reward-mean/3839)

<div class="topic-metadata">

**Author:** [@carlorop](https://discuss.ray.io/u/carlorop)\
**Replies:** 10\
**Last updated:** [September 21, 2023, 2:36pm UTC](https://discuss.ray.io/t/meaning-of-episode-reward-mean/3839 "2023-09-21T14:36:36Z")

</div>

What is the meaning of the episode\_reward\_mean metric? Is it the sum of the reward obtained in each time step of the episode? What is the difference between episode\_reward\_mean and episode\_reward\_min?

---

## [Is any multi discrete action example for PPO or other algorithms?](https://discuss.ray.io/t/is-any-multi-discrete-action-example-for-ppo-or-other-algorithms/4693)

<div class="topic-metadata">

**Author:** [@James\_Liu](https://discuss.ray.io/u/James_Liu)\
**Replies:** 9\
**Last updated:** [January 29, 2023, 7:25pm UTC](https://discuss.ray.io/t/is-any-multi-discrete-action-example-for-ppo-or-other-algorithms/4693 "2023-01-29T19:25:26Z")

</div>

Hi, I cannot find multi discrete action example. May someone point out this for me? Thanks in advance.

---

## [Observation space with multiple input](https://discuss.ray.io/t/observation-space-with-multiple-input/3726)

<div class="topic-metadata">

**Author:** [@deepgravity](https://discuss.ray.io/u/deepgravity)\
**Replies:** 15\
**Last updated:** [December 10, 2021, 1:41am UTC](https://discuss.ray.io/t/observation-space-with-multiple-input/3726 "2021-12-10T01:41:15Z")

</div>

Hi all, I am working on a Multi-Agent task and I would like to use images and also some other information like the location of the agent as the observation space. It seems that in Gym we can do it as it is explained he…

---

## [Unable to restore fully trained checkpoint](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259)

<div class="topic-metadata">

**Author:** [@Xorgress\_Grox](https://discuss.ray.io/u/Xorgress_Grox)\
**Replies:** 19\
**Last updated:** [October 21, 2023, 4:54am UTC](https://discuss.ray.io/t/unable-to-restore-fully-trained-checkpoint/8259 "2023-10-21T04:54:59Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. I’ve finished training with a bunch of algorithms using the Tuner() API and air library and they all have their a…

---

## [Missing 'grad\_gnorm' key in some \`input\_trees\` after some training time](https://discuss.ray.io/t/missing-grad-gnorm-key-in-some-input-trees-after-some-training-time/8553)

<div class="topic-metadata">

**Author:** [@Blubberblub](https://discuss.ray.io/u/Blubberblub)\
**Replies:** 23\
**Last updated:** [January 29, 2023, 11:24am UTC](https://discuss.ray.io/t/missing-grad-gnorm-key-in-some-input-trees-after-some-training-time/8553 "2023-01-29T11:24:35Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. I’m running a self built multi agent env with PPO on a local ray cluster. At the beggining everything works fine …

---

## [How to define fcnet\_hiddens size and number of layers in rllib tune?](https://discuss.ray.io/t/how-to-define-fcnet-hiddens-size-and-number-of-layers-in-rllib-tune/6504)

<div class="topic-metadata">

**Author:** [@Peter\_Pirog](https://discuss.ray.io/u/Peter_Pirog)\
**Replies:** 18\
**Last updated:** [January 19, 2023, 2:39pm UTC](https://discuss.ray.io/t/how-to-define-fcnet-hiddens-size-and-number-of-layers-in-rllib-tune/6504 "2023-01-19T14:39:20Z")

</div>

I wonder how to define in rllib tune layers and neurons in the layers. I would like to do it with two parameters: number of layers 1,2 or 3 neurons in layer 8 to 256 Each layer has the same number of neurons for simp…

---

## [Apply preprocessor in custom model](https://discuss.ray.io/t/apply-preprocessor-in-custom-model/5961)

<div class="topic-metadata">

**Author:** [@fedetask](https://discuss.ray.io/u/fedetask)\
**Replies:** 19\
**Last updated:** [May 13, 2024, 3:36pm UTC](https://discuss.ray.io/t/apply-preprocessor-in-custom-model/5961 "2024-05-13T15:36:21Z")

</div>

My observation is a Dict { 'observation': Dict({ .. dictionary with observations }), 'mask': Box() # Mask for action masking } and I have a custom model 1 class DQNModel(TFModelV2): 2 3 def \_\_init\_\_(self, …

---

## [How do I set GPU affinity of workers](https://discuss.ray.io/t/how-do-i-set-gpu-affinity-of-workers/1755)

<div class="topic-metadata">

**Author:** [@Bam4d](https://discuss.ray.io/u/Bam4d)\
**Replies:** 17\
**Last updated:** [April 23, 2021, 10:05am UTC](https://discuss.ray.io/t/how-do-i-set-gpu-affinity-of-workers/1755 "2021-04-23T10:05:16Z")

</div>

Lets say I am using IMPALA with several workers and I have 4 GPUs. I want to be able to say GPUs 0,1,2 should be shared across the workers evenly so inference/rendering is accellerated by them, But i want GPU 3 to be de…

---

## [RLLib Multiagent: Load only one policy from checkpoint & Compatibility of RLLib/Tune Checkpoints](https://discuss.ray.io/t/rllib-multiagent-load-only-one-policy-from-checkpoint-compatibility-of-rllib-tune-checkpoints/1827)

<div class="topic-metadata">

**Author:** [@Rafael\_Albert](https://discuss.ray.io/u/Rafael_Albert)\
**Replies:** 9\
**Last updated:** [November 24, 2021, 8:49pm UTC](https://discuss.ray.io/t/rllib-multiagent-load-only-one-policy-from-checkpoint-compatibility-of-rllib-tune-checkpoints/1827 "2021-11-24T20:49:57Z")

</div>

I am working in a multiagent setup with 3 agents and I want to use pretrained weights for one (and only one!) of them. In my current workflow, I run my experiments using tune.run() and would prefer to keep it that way. I…

---

## [Compute/display actions from ray.tune](https://discuss.ray.io/t/compute-display-actions-from-ray-tune/1200)

<div class="topic-metadata">

**Author:** [@Carterbouley](https://discuss.ray.io/u/Carterbouley)\
**Replies:** 10\
**Last updated:** [March 30, 2021, 9:54am UTC](https://discuss.ray.io/t/compute-display-actions-from-ray-tune/1200 "2021-03-30T09:54:24Z")

</div>

Hi everyone. I have trained a PPO agend using: tune.run( run\_or\_experiment="PPO", config={ "env": "Battery", "num\_gpus" : 1, "num\_workers": 13, "num\_cpus\_per\_worker": 1, …

---

## [Stacking callback objects \[Solved. Code included.\]](https://discuss.ray.io/t/stacking-callback-objects-solved-code-included/1925)

<div class="topic-metadata">

**Author:** [@RickLan](https://discuss.ray.io/u/RickLan)\
**Replies:** 12\
**Last updated:** [April 30, 2021, 11:29am UTC](https://discuss.ray.io/t/stacking-callback-objects-solved-code-included/1925 "2021-04-30T11:29:16Z")

</div>

Is there an elegant way to stack DefaultCallbacks objects? I have one that mutates sample batch in on\_postprocess\_trajectory() and another that logs custom metrics in on\_episode\_end(). Right now I merge the code into on…

---

## [Very slow gradient descent on remote workers](https://discuss.ray.io/t/very-slow-gradient-descent-on-remote-workers/1278)

<div class="topic-metadata">

**Author:** [@smorad](https://discuss.ray.io/u/smorad)\
**Replies:** 14\
**Last updated:** [June 8, 2021, 7:37am UTC](https://discuss.ray.io/t/very-slow-gradient-descent-on-remote-workers/1278 "2021-06-08T07:37:58Z")

</div>

I am using ray.tune to run ImpalaTrainer trainables. For some reason, gradient descent becomes unbearably slow using remote workers. However, local workers ray.init(local\_mode=true) do not seem to have this problem. The …

---

## [Best way to have custom value state + LSTM](https://discuss.ray.io/t/best-way-to-have-custom-value-state-lstm/2029)

<div class="topic-metadata">

**Author:** [@nathanlct](https://discuss.ray.io/u/nathanlct)\
**Replies:** 9\
**Last updated:** [April 10, 2022, 7:05pm UTC](https://discuss.ray.io/t/best-way-to-have-custom-value-state-lstm/2029 "2022-04-10T19:05:20Z")

</div>

Hi, I’m doing some training using PPO, and I would like the value function to have additional states that the policy doesn’t have. By default, the FullyConnectedNetwork looks like this (14 should be 7 here): I slig…

---

## [Right way to use tuple action space](https://discuss.ray.io/t/right-way-to-use-tuple-action-space/3594)

<div class="topic-metadata">

**Author:** [@Ofir\_Abu](https://discuss.ray.io/u/Ofir_Abu)\
**Replies:** 9\
**Last updated:** [September 24, 2021, 11:30am UTC](https://discuss.ray.io/t/right-way-to-use-tuple-action-space/3594 "2021-09-24T11:30:09Z")

</div>

Hi there, I’m using a custom environment with a tuple (gym space) action space. TL;DR - I’m having trouble about how should I construct the output of the model from the forward function. Specifically: my action space…

---

## [Maximum recommended reward](https://discuss.ray.io/t/maximum-recommended-reward/6728)

<div class="topic-metadata">

**Author:** [@evo11x](https://discuss.ray.io/u/evo11x)\
**Replies:** 18\
**Last updated:** [July 14, 2022, 7:50pm UTC](https://discuss.ray.io/t/maximum-recommended-reward/6728 "2022-07-14T19:50:04Z")

</div>

What is the maximum reward per step which should be used? I have seen some environments with rewards as high as 100, should I use only rewards between 0 and 1 ? Medium: It contributes to significant difficulty to comp…

---

## [Accessing info dicts in postprocessing callback](https://discuss.ray.io/t/accessing-info-dicts-in-postprocessing-callback/118)

<div class="topic-metadata">

**Author:** [@cwerner](https://discuss.ray.io/u/cwerner)\
**Replies:** 10\
**Last updated:** [January 11, 2021, 1:37am UTC](https://discuss.ray.io/t/accessing-info-dicts-in-postprocessing-callback/118 "2021-01-11T01:37:23Z")

</div>

Hello, I recently updated from rllib 0.8.6 to rllib 1.0.1 and also converted from TF to Torch. Previously, I was using the info dict from my custom environment to pass along values that I used in the postprocessing call…

---

## [Custom RNN Model with Examples - why do they fail?](https://discuss.ray.io/t/custom-rnn-model-with-examples-why-do-they-fail/1334)

<div class="topic-metadata">

**Author:** [@Gregory](https://discuss.ray.io/u/Gregory)\
**Replies:** 11\
**Last updated:** [May 5, 2022, 3:06pm UTC](https://discuss.ray.io/t/custom-rnn-model-with-examples-why-do-they-fail/1334 "2022-05-05T15:06:30Z")

</div>

Using the current pip install ray\[debug\] and ray\[rllib\], here is my minimum reproducible example #1 using dictionary observations. This fails with the following error: RuntimeError: Expected hidden\[0\] size (1, 1, 256),…

---

## [Global optima with centralized critic (basic understanding)](https://discuss.ray.io/t/global-optima-with-centralized-critic-basic-understanding/1478)

<div class="topic-metadata">

**Author:** [@CodingBurmer](https://discuss.ray.io/u/CodingBurmer)\
**Replies:** 10\
**Last updated:** [April 10, 2021, 12:25pm UTC](https://discuss.ray.io/t/global-optima-with-centralized-critic-basic-understanding/1478 "2021-04-10T12:25:03Z")

</div>

hi guys, I just spend over two days tying to find out how a centralized critic leads MARL agents to a global optima. So I thought why no asking here and maybe help other MARL beginners. In my test environment there a tw…

---

## [How to log Render to tensorboard?](https://discuss.ray.io/t/how-to-log-render-to-tensorboard/1607)

<div class="topic-metadata">

**Author:** [@Sertingolix](https://discuss.ray.io/u/Sertingolix)\
**Replies:** 9\
**Last updated:** [July 22, 2021, 10:36am UTC](https://discuss.ray.io/t/how-to-log-render-to-tensorboard/1607 "2021-07-22T10:36:39Z")

</div>

Hi there, I’m trying to log the render of my environment to tensorboard. Similar to @smorad proposal in tensorboard render log i created the following callback class VideoCallback(DefaultCallbacks): def on\_episode…

---

## [Reward function not converging during training](https://discuss.ray.io/t/reward-function-not-converging-during-training/6513)

<div class="topic-metadata">

**Author:** [@leo593](https://discuss.ray.io/u/leo593)\
**Replies:** 14\
**Last updated:** [July 11, 2022, 8:49am UTC](https://discuss.ray.io/t/reward-function-not-converging-during-training/6513 "2022-07-11T08:49:39Z")

</div>

How severe does this issue affect your experience of using Ray? Medium: It contributes to significant difficulty to complete my task, but I can work around it. Hi! I am currently working on an automatic train simula…

---

## [PPO trainer eating up memory](https://discuss.ray.io/t/ppo-trainer-eating-up-memory/1437)

<div class="topic-metadata">

**Author:** [@Rory](https://discuss.ray.io/u/Rory)\
**Replies:** 9\
**Last updated:** [April 2, 2021, 5:47pm UTC](https://discuss.ray.io/t/ppo-trainer-eating-up-memory/1437 "2021-04-02T17:47:04Z")

</div>

Hi there, I’m trying to train a PPO agent via self play in my multi-agent env. At the moment it can manage about 320 training iterations before my system runs of memory (16gb (with 16gb of swap).) If I restart the train…

---

## [Get agent ID in multi-agent setting](https://discuss.ray.io/t/get-agent-id-in-multi-agent-setting/3535)

<div class="topic-metadata">

**Author:** [@lucas\_spangher](https://discuss.ray.io/u/lucas_spangher)\
**Replies:** 16\
**Last updated:** [October 5, 2021, 7:53am UTC](https://discuss.ray.io/t/get-agent-id-in-multi-agent-setting/3535 "2021-10-05T07:53:03Z")

</div>

Sorry, this is probably a very basic question. I’m not sure where to retrieve the agent\_ids for the multiple agents created in a multiagent setting, so I can map the policy functions on. Can anyone please point me there? …

---

## [\[Bug\] Env must be one of the supported types: BaseEnv, gym.Env, MultiAgentEnv, VectorEnv, RemoteBaseEnv](https://discuss.ray.io/t/bug-env-must-be-one-of-the-supported-types-baseenv-gym-env-multiagentenv-vectorenv-remotebaseenv/5704)

<div class="topic-metadata">

**Author:** [@ZKBig](https://discuss.ray.io/u/ZKBig)\
**Replies:** 10\
**Last updated:** [March 2, 2023, 7:39pm UTC](https://discuss.ray.io/t/bug-env-must-be-one-of-the-supported-types-baseenv-gym-env-multiagentenv-vectorenv-remotebaseenv/5704 "2023-03-02T19:39:24Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. When testing my custom multi-agent environment using ‘check\_env’ function，The result shows that my environment i…

---

## [Memory Leak when training PPO on a single agent environment](https://discuss.ray.io/t/memory-leak-when-training-ppo-on-a-single-agent-environment/8712)

<div class="topic-metadata">

**Author:** [@MrDracoG](https://discuss.ray.io/u/MrDracoG)\
**Replies:** 15\
**Last updated:** [December 24, 2022, 3:41am UTC](https://discuss.ray.io/t/memory-leak-when-training-ppo-on-a-single-agent-environment/8712 "2022-12-24T03:41:46Z")

</div>

How severe does this issue affect your experience of using Ray? High: It blocks me to complete my task. I have actually been stuck on this memory leak for a while. I was originally using ray tune to setup and run my s…

---

## [RLlib's PolicyServer and external simulator as client](https://discuss.ray.io/t/rllibs-policyserver-and-external-simulator-as-client/1470)

<div class="topic-metadata">

**Author:** [@klausk55](https://discuss.ray.io/u/klausk55)\
**Replies:** 15\
**Last updated:** [April 12, 2021, 6:39pm UTC](https://discuss.ray.io/t/rllibs-policyserver-and-external-simulator-as-client/1470 "2021-04-12T18:39:05Z")

</div>

Hello Ray community, I use RLlib in combination with a custom external simulator. For this purpose, I use a PolicyServer on RLlib’s side and a client on external simulator’s side (HTTP server/client). Now, my problem i…

---

## [Playing the QMIX Two-step game on Ray](https://discuss.ray.io/t/playing-the-qmix-two-step-game-on-ray/4200)

<div class="topic-metadata">

**Author:** [@xeirwn](https://discuss.ray.io/u/xeirwn)\
**Replies:** 11\
**Last updated:** [October 18, 2022, 8:51am UTC](https://discuss.ray.io/t/playing-the-qmix-two-step-game-on-ray/4200 "2022-10-18T08:51:57Z")

</div>

We are trying to expand the code of the Two-step game (which is an example from the QMIX paper) using the Ray framework . The changes we want to apply should extract the best checkpoint from some trial of a tune.run() , …

---

## [Error with torch policy and ray.get\_gpu\_ids on Windows](https://discuss.ray.io/t/error-with-torch-policy-and-ray-get-gpu-ids-on-windows/2711)

<div class="topic-metadata">

**Author:** [@Fabien-Couthouis](https://discuss.ray.io/u/Fabien-Couthouis)\
**Replies:** 9\
**Last updated:** [July 30, 2021, 8:41am UTC](https://discuss.ray.io/t/error-with-torch-policy-and-ray-get-gpu-ids-on-windows/2711 "2021-07-30T08:41:33Z")

</div>

Hello, I have an error on Windows when using PPO with Pytorch: ...\\envs\\rllib-pt\\lib\\site-packages\\ray\\rllib\\policy\\torch\_policy.py", line 155, in \_\_init\_\_ (pid=18860) self.device = self.devices\[0\] (pid=18860) Inde…

[Next page](https://discuss.ray.io/c/rllib/7.md?page=1&per_page=50)
