# Log action sequence per episode

**URL:** <https://discuss.ray.io/t/log-action-sequence-per-episode/15160>\
**Category:** RLlib\
**Created:** [July 9, 2024, 8:51pm UTC](https://discuss.ray.io/t/log-action-sequence-per-episode/15160 "2024-07-09T20:51:44Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![PhilippWillms](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/philippwillms/32/3885_2.png) [@PhilippWillms](https://discuss.ray.io/u/PhilippWillms)\
**Post date:** [July 9, 2024, 8:51pm UTC](https://discuss.ray.io/t/log-action-sequence-per-episode/15160/1 "2024-07-09T20:51:44Z")

</div>

**How severe does this issue affect your experience of using Ray?**

- Low: It annoys or frustrates me for a moment.

Dear ray community,

what is the best way to extract the sequence of actions leading to episode\_max\_reward?

My scenario is that I have an environment with episodes of fixed lengths. The order in which certain actions are taken is decisive for the reward. Is there any smart way to log within ray tune or train API the actions taken in an episode, so that afterwards the sequence of actions can be collected?

I am thinking of custom metrics and callbacks. A first try with on\_episode\_step() unfortunately failed, as last\_action\_for() method is not available in class EpisodeV2, which is used by PPO in ray 2.10.0.

---

<div class="post-metadata">

**Author:** ![starkj](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/starkj/32/3266_2.png) [@starkj](https://discuss.ray.io/u/starkj)\
**Post date:** [July 14, 2024, 1:31am UTC](https://discuss.ray.io/t/log-action-sequence-per-episode/15160/2 "2024-07-14T01:31:57Z")

</div>

If you just want the sequence in order to assign a reward in thr context of a single episode, then it can all be handled within the environment. So why not just have step() append the incoming action to a member action list, then the reward function can see it at any time.

---

<div class="post-metadata">

**Author:** ![PhilippWillms](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/philippwillms/32/3885_2.png) [@PhilippWillms](https://discuss.ray.io/u/PhilippWillms)\
**Post date:** [July 15, 2024, 10:09am UTC](https://discuss.ray.io/t/log-action-sequence-per-episode/15160/3 "2024-07-15T10:09:26Z")

</div>

As I said in original post, reward function visibility is one thing. The other thing which comes to my mind now is to improve accessibility for the best solution. Currently, I have my difficulties to access that and corresponding action sequence **after training** via a trainer based on PPOConfig.build().
