# Changing add\_time\_dimension logic

**URL:** <https://discuss.ray.io/t/changing-add-time-dimension-logic/11185>\
**Category:** RLlib\
**Created:** [June 28, 2023, 7:16am UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185 "2023-06-28T07:16:03Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![hossein836](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hossein836/32/2559_2.png) [@hossein836](https://discuss.ray.io/u/hossein836)\
**Post date:** [June 28, 2023, 7:16am UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185/1 "2023-06-28T07:16:03Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- None: Just asking a question out of curiosity

The function `add_time_dimension` is responsible for creating training batches in recurrent networks in rllib. It divides the samples based on the `max_seq_len` parameter. For example, if we have 100 samples and `max_seq_len` is set to 30, we will obtain 4 batches, each with a length of 30 (20 samples will be padding). Although the current approach is not incorrect, I believe it might be better to create batches where all samples are equal to `max_seq_len`. In this case, with 100 samples and `max_seq_len` set to 30, the batches would be [100, 30, f], as opposed to [4, 30, f].

There are trade-offs to consider in this approach. While it may make the training process smoother because all steps are treated equally, similar to non-recurrent nets, it would also increase the computational cost of the model.  
I just wanna know your thoughts on this!

---

<div class="post-metadata">

**Author:** ![Rohan138](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rohan138/32/1160_2.png) [@Rohan138](https://discuss.ray.io/u/Rohan138)\
**Post date:** [June 29, 2023, 12:37am UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185/2 "2023-06-29T00:37:02Z")

</div>

How would you implement the RNN in this case? This seems like it would significantly increase the memory and compute required for the model. Do you know if this would improve the learning throughput?

---

<div class="post-metadata">

**Author:** ![hossein836](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hossein836/32/2559_2.png) [@hossein836](https://discuss.ray.io/u/hossein836)\
**Post date:** [June 29, 2023, 5:59am UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185/3 "2023-06-29T05:59:30Z")

</div>

to implement RNN in this case you should change model view requirements to add past observations(equal to `max_seq_len`) to your model and then reshape all observations to [samples, seq\_len, features].  
Yes it significantly increase memory but my question is : does it make model work better? I don’t know if it increase learning throughput or not but intuitively for me it should increase, because you are passing more batches to model (like 100vs4 in my example) so it should worth to use more memory.  
the very important point for me is in rllib implementation the model learns to act as it sees observations sequences from zero to max\_seq\_len, it’s like adding much noise to model, noises also could be beneficial like we add noise in image classification but it is double edge sword, it can also interrupt model learning.  
I want to know is there any problem beside memory in the implementation which I proposed?  
because it’s make more sense to me. thanks

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [June 29, 2023, 4:23pm UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185/4 "2023-06-29T16:23:16Z")

</div>

Hi @hossein836,

If I understand your suggestion correctly, I think you could already do this by setting your training batch size to (100\*max\_seq\_len). You can actually end up with a few more than 100 in the batch dimension if some of the episodes are less than the max\_seq\_len.

---

<div class="post-metadata">

**Author:** ![hossein836](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hossein836/32/2559_2.png) [@hossein836](https://discuss.ray.io/u/hossein836)\
**Post date:** [June 29, 2023, 5:27pm UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185/5 "2023-06-29T17:27:18Z")

</div>

Hi @mannyv  
You are correct but my point was not that.  
I mean this:

 ![IMG_20230629_215114](https://us1.discourse-cdn.com/flex020/uploads/ray/original/2X/8/8df185ce2581ce47268a4e0e60f78ac55f45705c.jpeg)

The second method is what I proposed

---

<div class="post-metadata">

**Author:** ![hossein836](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hossein836/32/2559_2.png) [@hossein836](https://discuss.ray.io/u/hossein836)\
**Post date:** [July 1, 2023, 9:43am UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185/6 "2023-07-01T09:43:32Z")

</div>

“I don’t want to increase the batch size, as it would also increase RAM usage. My method focuses on gaining more knowledge within the same batch size instead of changing it. However, the downside of this approach is that it will increase GPU RAM usage.”

---

<div class="post-metadata">

**Author:** ![kourosh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kourosh/32/2130_2.png) [@kourosh](https://discuss.ray.io/u/kourosh)\
**Post date:** [July 5, 2023, 4:08am UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185/7 "2023-07-05T04:08:30Z")

</div>

Hi @hossein836,

What you want to do is fair but it should not be the default behavior of LSTMs in RLlib. You can achieve custom sample collection via setting a custom trajectory view for your custom model and obtain what you have in mind. I don’t know the exact syntax but I believe it should be possible to achieve it via this API.  
Reference: [Sample Collections and Trajectory Views — Ray 2.5.1](https://docs.ray.io/en/latest/rllib/rllib-sample-collection.html)  
Something like:  
`{"slided_obs": ViewRequirement("obs", shift="-3:-1", used_for_training=True)}`

---

<div class="post-metadata">

**Author:** ![hossein836](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hossein836/32/2559_2.png) [@hossein836](https://discuss.ray.io/u/hossein836)\
**Post date:** [July 5, 2023, 9:17am UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185/8 "2023-07-05T09:17:32Z")

</div>

Damet Garm @kourosh jan 😉  
I agree with your proposed syntax; that should be fine. However, while I appreciate your hard works, from the perspective of reinforcement learning theories, I have these questions:

1- Do you also believe my method can potentially increase learning throughput ?  
2- Do you also believe my method can potentially decrease instability of learning ?  
3- What are the perks of rllib implementation in contrast to mine? Is it just hardware restrictions, or are there some RL theories that I don’t understand (hence you said “it shouldn’t be default”?

---

<div class="post-metadata">

**Author:** ![kourosh](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/kourosh/32/2130_2.png) [@kourosh](https://discuss.ray.io/u/kourosh)\
**Post date:** [July 6, 2023, 4:21am UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185/9 "2023-07-06T04:21:21Z")

</div>

@hossein836 jan,

You are not really increasing throughput this way, you are basically increasing the gradient update intensity. This means that you are use more samples to update the network. Whether that works better or not in practice, really depends on the particular use-case. So the short answer is that you gotta try and see. My hypothesis is that it won’t really help in general, because on-policy methods like PPO need to move on from their bad initial randomness and if you train them with more intensity when the policy is not so much better than random can actually cause convergence issues.

The implementation in RLlib is a simple extension of the basic algorithms and does not introduce these possible un-wanted effects from increasing training intensity.

---

<div class="post-metadata">

**Author:** ![hossein836](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hossein836/32/2559_2.png) [@hossein836](https://discuss.ray.io/u/hossein836)\
**Post date:** [July 6, 2023, 5:22am UTC](https://discuss.ray.io/t/changing-add-time-dimension-logic/11185/10 "2023-07-06T05:22:51Z")

</div>

@kourosh AALI  
I understand your point. I was considering using a buffer to address the convergence issue you mentioned, specifically sampling x experiences out of y episodes. As far as I know, there is currently no buffer implementation for PPO in rllib. However, the OpenAI Five, which is a PPO agent and RNN model, used a buffer. If I don’t find any information on how to use a buffer in PPO, I will create a new topic.  
many thanks 🙂
