# Get\_initial\_state for LSTM custom model without initial FC

**URL:** <https://discuss.ray.io/t/get-initial-state-for-lstm-custom-model-without-initial-fc/4537>\
**Category:** RLlib\
**Created:** [December 29, 2021, 12:59pm UTC](https://discuss.ray.io/t/get-initial-state-for-lstm-custom-model-without-initial-fc/4537 "2021-12-29T12:59:34Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![mg64ve](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mg64ve/32/606_2.png) [@mg64ve](https://discuss.ray.io/u/mg64ve)\
**Post date:** [December 29, 2021, 12:59pm UTC](https://discuss.ray.io/t/get-initial-state-for-lstm-custom-model-without-initial-fc/4537/1 "2021-12-29T12:59:34Z")

</div>

Hi, I have just put the attention net with requirement in one file:

> <https://github.com/mg64ve/ray_tests/blob/main/pytorch/CustomModels/ViewRequirements/custom_model_trajectory.py>

and then I replaced the model with LSTM only without initial fc layers.  
It is in this file:

> <https://github.com/mg64ve/ray_tests/blob/main/pytorch/CustomModels/ViewRequirements/custom_model_trajectory_lstm.py>

the problem I am facing is with the get\_initial\_state function.  
I have tried many options and no one seems to be good:

```
def get_initial_state(self):
    # h = [
    # torch.zeros(self.lstm_size),
    # torch.zeros(self.lstm_size)
    # ]
    # h = [
    # torch.zeros(1, self.lstm_size),
    # torch.zeros(1, self.lstm_size)
    # ]
    h = self.lstm.weight_hh_l0.data.fill_(0)
    return h

```

Can you please give me an hint on this?  
in pure pytorch one of these methods should work but in ray it does not.  
I am getting the following error:

`(PPOTrainer pid=133035) RuntimeError: Expected hidden[0] size (1, 32, 16), got [1, 4, 16]`

---

<div class="post-metadata">

**Author:** ![sven1977](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/sven1977/32/53_2.png) [@sven1977](https://discuss.ray.io/u/sven1977)\
**Post date:** [January 12, 2022, 3:18pm UTC](https://discuss.ray.io/t/get-initial-state-for-lstm-custom-model-without-initial-fc/4537/2 "2022-01-12T15:18:12Z")

</div>

Hey @mg64ve , thanks for the question! Some problems I see in your implementation of `get_initial_state`:

- The return value should always be a list of state tensors, so in your case, a list with one single item, which is the h-state-tensor (you are returning h directly w/o the list).
- You seem to return a state tensor that has the same shape as the weight matrix, but I think you should return a state tensor that has the same shape as the bias vector.

Also, state tensors in your returned list should all be non-batched, but I think you are doing this correctly here.
