# Switching exploration through action subspaces

**URL:** <https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163>\
**Category:** RLlib\
**Created:** [November 7, 2022, 2:10pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163 "2022-11-07T14:10:22Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Saurabh\_Arora](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/saurabh_arora/32/1012_2.png) [@Saurabh\_Arora](https://discuss.ray.io/u/Saurabh_Arora)\
**Post date:** [November 7, 2022, 2:10pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/1 "2022-11-07T14:10:22Z")

</div>

Hey team

I am using a MDP in which a discrete action space can be divided into two sub-spaces, based on some parameters. Can I create an exploration class that forces learner do following?

- in current call for getting action, sample action from first subspace
- take that action
- in next call for getting action, sample action from second subspace
- take that action
- repeat switching between sub-spaces from call to call

@sven1977

---

<div class="post-metadata">

**Author:** ![Saurabh\_Arora](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/saurabh_arora/32/1012_2.png) [@Saurabh\_Arora](https://discuss.ray.io/u/Saurabh_Arora)\
**Post date:** [November 7, 2022, 8:25pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/3 "2022-11-07T20:25:25Z")

</div>

cc: @RickLan , @mannyv , @arturn , @RickDW , @rusu24edward , @gjoliver

---

<div class="post-metadata">

**Author:** ![rusu24edward](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rusu24edward/32/333_2.png) [@rusu24edward](https://discuss.ray.io/u/rusu24edward)\
**Post date:** [November 7, 2022, 10:13pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/4 "2022-11-07T22:13:02Z")

</div>

Not sure how to do this through exploration. However, you can easily do this by casting it as a multi-agent problem. On the RLlib side, you’ll have two policies, and you’ll alternate which policy you use each step. You can drive this alternation in the simulation by switching which “agent” outputs at each step. In the first step, you output data for policy 1, in the next step policy 2, then back to 1, then 2, and so on until the end. The actions that come in at step will be keyed off the policy id, but you can just take and use the value, and it should run the same as your current set up.

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [November 8, 2022, 2:19am UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/5 "2022-11-08T02:19:56Z")

</div>

@Saurabh_Arora I was going to suggest the same setup as @rusu24edward described.

You are going to have trouble with the exploration approach during the learning phase because at that point the exploration is not usually used and all the steps are computed in a batch at the same time so you will have to have some custom bookkeeping to know which branch should be trained and which should not for each sample in the batch.

---

<div class="post-metadata">

**Author:** ![Saurabh\_Arora](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/saurabh_arora/32/1012_2.png) [@Saurabh\_Arora](https://discuss.ray.io/u/Saurabh_Arora)\
**Post date:** [November 8, 2022, 12:57pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/6 "2022-11-08T12:57:47Z")

</div>

Thanks  
@rusu24edward

But, both policies will have same shared action space. So both agents are trained equally towards all actions. Then, why will one agent prefer a specific subspace? Can you share a minimal set of bulleted steps for this idea?

---

<div class="post-metadata">

**Author:** ![Saurabh\_Arora](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/saurabh_arora/32/1012_2.png) [@Saurabh\_Arora](https://discuss.ray.io/u/Saurabh_Arora)\
**Post date:** [November 8, 2022, 1:25pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/7 "2022-11-08T13:25:34Z")

</div>

In parallel, I am looking at exploration option. How do I get values from TensorType timestep and List[TensorType] action\_distribution.inputs?

> ```
> def get_exploration_action(
> self,
> *,
> action_distribution: ActionDistribution,
> timestep: Optional[Union[int, TensorType]] = None,
> explore: bool = True
> ):
> 
> ```

---

<div class="post-metadata">

**Author:** ![Saurabh\_Arora](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/saurabh_arora/32/1012_2.png) [@Saurabh\_Arora](https://discuss.ray.io/u/Saurabh_Arora)\
**Post date:** [November 8, 2022, 2:37pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/8 "2022-11-08T14:37:05Z")

</div>

@mannyv , you mentioned that exploration is not used in learning phase. RL exploration is supposed to happen while learning. I am confused what you meant.Can you please clarify?

---

<div class="post-metadata">

**Author:** ![rusu24edward](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/rusu24edward/32/333_2.png) [@rusu24edward](https://discuss.ray.io/u/rusu24edward)\
**Post date:** [November 8, 2022, 10:19pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/9 "2022-11-08T22:19:35Z")

</div>

> [@Saurabh\_Arora](#):
>
> But, both policies will have same shared action space.

The action space from the first policy should be the first subspace, and the action space from the second policy should be the second subspace.

Here’s some pseudocode for setting this up:

```python
def MySim(MultiAgentEnv):
    def reset(self):
        return {'policy_1': first_obs} # This tells RLlib to generate an action using policy 1

    def step(self, action_dict):
        # action_dict will be keyed of the agent's id
        action = next(iter(action_dict.values())
        # process the action
        if next(iter(action_dict.keys())) == 'policy_1':
            key_off = 'policy_2'
        else
            key_off = 'policy_1'
        return {key_off: next_obs, key_off: reward, key_off: done_status, key_off: info}
        # This tells RLlib to generate an action using whatever policy is key_off

```

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [November 9, 2022, 1:01pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/10 "2022-11-09T13:01:59Z")

</div>

Hi @Saurabh_Arora,

I am not sure exactly which algorithm you are using but in general there are two main phases in a call to algorithm.train().

The first phase is the rollout phase. During this phase, copies of the environment interact with the policy to generate new samples. During this phase the exploration object is used to sample actions. These samples can be deterministic or probabalistic depending on config settings.

The second phase is the policy update phase. During this phase previously collected samples are used to update the policy according to the algorithms loss function. During this phase for most algorithms in rllib actions are not generated and so the exploration object is not used during this step.

---

<div class="post-metadata">

**Author:** ![Saurabh\_Arora](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/saurabh_arora/32/1012_2.png) [@Saurabh\_Arora](https://discuss.ray.io/u/Saurabh_Arora)\
**Post date:** [November 11, 2022, 7:42pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/11 "2022-11-11T19:42:10Z")

</div>

Thanks for a clear explanation @mannyv . I am able to implement exploration class that achieves the training I want for learner. Can you please look into my follow up question here

> [@Using different get\_exploration\_action logic pre and post training](https://discuss.ray.io/t/using-different-get-exploration-action-method-pre-and-post-training/8242):
>
> Hey team , I have created a custom exploration class for problem I am trying to solve. I want to use two different get\_exploration\_action methods for following two parts of my code: PPO training compute\_action during simulation of episodes from learned policy post training. Is there a way to modify config[“exploration\_config”][“type”] after training, or add a custom argument for call to get\_exploration\_action? If neither, Can you please suggest a way to implement such a set up? cc: @sv…

---

<div class="post-metadata">

**Author:** ![Saurabh\_Arora](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/saurabh_arora/32/1012_2.png) [@Saurabh\_Arora](https://discuss.ray.io/u/Saurabh_Arora)\
**Post date:** [November 11, 2022, 7:44pm UTC](https://discuss.ray.io/t/switching-exploration-through-action-subspaces/8163/12 "2022-11-11T19:44:16Z")

</div>

@rusu24edward , if I switch to multi-agent setting, I have to change a lot of code in rest of my codebase that uses learned policy. I think I should first give a shot to exploration. Can you please look into my follow up question here [Using different get\_exploration\_action method pre and post training](https://discuss.ray.io/t/using-different-get-exploration-action-method-pre-and-post-training/8242) ?
