# Pre-train a model with baseline policy

**URL:** <https://discuss.ray.io/t/pre-train-a-model-with-baseline-policy/13645>\
**Category:** Offline RL\
**Created:** [February 6, 2024, 10:57am UTC](https://discuss.ray.io/t/pre-train-a-model-with-baseline-policy/13645 "2024-02-06T10:57:20Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![hermmanhender](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/hermmanhender/32/5691_2.png) [@hermmanhender](https://discuss.ray.io/u/hermmanhender)\
**Post date:** [February 6, 2024, 10:57am UTC](https://discuss.ray.io/t/pre-train-a-model-with-baseline-policy/13645/1 "2024-02-06T10:57:21Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- Low: It annoys or frustrates me for a moment.

Hi!

I trained a DQN Algorithm to obtain a optimal policy to control actions in a custom environment. For the evaluation I compared the policy with a conventional rule-based policy. I did not archive better results with the policy learned. For that reason I would like to take adventage of the rule-based policy and pre-train the model with it.

For that I was exploring the documentation and saw two options to reach my objetive.

The first one is [working with offline data](https://docs.ray.io/en/latest/rllib/rllib-offline.html). Especifically with [converting the external experiences to batch format](https://docs.ray.io/en/latest/rllib/rllib-offline.html#example-converting-external-experiences-to-batch-format), where I must to write a similar code of the example for my application. But my environment can be simulated, so I explored other options.

The second one is to implement a custom exploration model and use to train the DQN/PG/Other Algorithm first and obtain in this way the pre-trained model. Regardles, I didn’t find documentation of how to implement a custom exploration model in DQN.

Someone have experience with this task of pre-train a model with a rule-based policy? Any recommendations or examples of how could I implement one of the two methods?

Thank you so much! Regards
