# \[RLlib\] Workaround for incorrect initial state shape with custom RNN models?

**URL:** <https://discuss.ray.io/t/rllib-workaround-for-incorrect-initial-state-shape-with-custom-rnn-models/343>\
**Category:** RLlib\
**Created:** [December 30, 2020, 1:06am UTC](https://discuss.ray.io/t/rllib-workaround-for-incorrect-initial-state-shape-with-custom-rnn-models/343 "2020-12-30T01:06:59Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Gregory](https://avatars.discourse-cdn.com/v4/letter/g/e56c9b/32.png) [@Gregory](https://discuss.ray.io/u/Gregory)\
**Post date:** [December 30, 2020, 1:06am UTC](https://discuss.ray.io/t/rllib-workaround-for-incorrect-initial-state-shape-with-custom-rnn-models/343/1 "2020-12-30T01:06:59Z")

</div>

Greetings everyone,

Back on June 21 [issue 9071](http://github.com/ray-project/ray/issues/9071) was opened regarding incorrect initial state shapes when using a custom model in both Tensorflow and Torch. Can I ask the more experienced users here how to correctly set the initial shape (so that it represents the correct batch size)?

Thanks for any tips

---

<div class="post-metadata">

**Author:** ![Gregory](https://avatars.discourse-cdn.com/v4/letter/g/e56c9b/32.png) [@Gregory](https://discuss.ray.io/u/Gregory)\
**Post date:** [December 30, 2020, 1:35am UTC](https://discuss.ray.io/t/rllib-workaround-for-incorrect-initial-state-shape-with-custom-rnn-models/343/2 "2020-12-30T01:35:31Z")

</div>

This is discussed in depth on ray-project/ray/issues/12509 but using 1.1.0 and the nightly 1.2 the challenge is still present. I’ve not been able to communicate with others about this, so if I find a solution I’ll share it.

---

<div class="post-metadata">

**Author:** ![Gregory](https://avatars.discourse-cdn.com/v4/letter/g/e56c9b/32.png) [@Gregory](https://discuss.ray.io/u/Gregory)\
**Post date:** [January 2, 2021, 3:03pm UTC](https://discuss.ray.io/t/rllib-workaround-for-incorrect-initial-state-shape-with-custom-rnn-models/343/3 "2021-01-02T15:03:58Z")

</div>

For anyone else struggling, I believe I have it running on a custom model by using LSTMWrapper(RecurrentNetwork) as a template for a custom model. It will be interesting when we find why it’s happening in the use case mentioned in the GitHub issue.
