# Roll out CQL policy

**URL:** <https://discuss.ray.io/t/roll-out-cql-policy/4216>\
**Category:** RLlib\
**Created:** [November 23, 2021, 12:49pm UTC](https://discuss.ray.io/t/roll-out-cql-policy/4216 "2021-11-23T12:49:05Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![fksvensson](https://avatars.discourse-cdn.com/v4/letter/f/839c29/32.png) [@fksvensson](https://discuss.ray.io/u/fksvensson)\
**Post date:** [November 23, 2021, 12:49pm UTC](https://discuss.ray.io/t/roll-out-cql-policy/4216/1 "2021-11-23T12:49:05Z")

</div>

Hello!

I am working with offline RL, specifically CQL. I have trained a policy offline using my prefered data. The policy is stared in a checkpoint and that I would like to restore. Since I want to evaluate my policy online on my environment I make some small adjustments to the configuration. The code looks something like this:

```auto
config['env'] = my_env

config['input'] = 'sampler'

trainer = CQLTrainer(config=config)

trainer.restore(checkpoint_path)

```

Running this I get the error:

```auto
ValueError: Unknown offline input! config['input'] must either be list of offline files (json) or a D4RL-specific InputReader specifier (e.g. 'd4rl.hopper-medium-v0').

```

This does not make sense to me. How can I evaluate a policy created by CQL online if it cannot use the sampler input?

I am using ray 1.4.0

---

<div class="post-metadata">

**Author:** ![Lars\_Simon\_Zehnder](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lars_simon_zehnder/32/1185_2.png) [@Lars\_Simon\_Zehnder](https://discuss.ray.io/u/Lars_Simon_Zehnder)\
**Post date:** [November 23, 2021, 5:45pm UTC](https://discuss.ray.io/t/roll-out-cql-policy/4216/2 "2021-11-23T17:45:59Z")

</div>

Hi @fksvensson ,

and welcome to the board. First of all I would like to ask, if you need this specific version of Ray because your algorithm depends on some specific settings that are now deprecated or because you are using an external environment that needs so? Otherwise I would suggest you to install v1.8.0 or at least v1.7.1.

Then, the `input` hyperparameter defines an input to the `Offline API` if you want to train on already collected data. In your case you want to evaluate your algorithm online by collecting data live. In this case the `env` hyperparameter should suffice. So, try to comment out `config[input]='sampler'` and see, if this makes it run.

Hope this helps

---

<div class="post-metadata">

**Author:** ![fksvensson](https://avatars.discourse-cdn.com/v4/letter/f/839c29/32.png) [@fksvensson](https://discuss.ray.io/u/fksvensson)\
**Post date:** [November 24, 2021, 8:29am UTC](https://discuss.ray.io/t/roll-out-cql-policy/4216/3 "2021-11-24T08:29:19Z")

</div>

Hello Lars and thank you for your fast answer.

I started the project with Ray 1.4.0 and I am afraid I will have combatability issues with the rest of my code if I switch version. Do you know if there are any major changes in the offline RL API that would motivate the change?

Even when I do not set a specific input and only declare the env in the config, I get the same issue, since ‘sampler’ is the default input.

```auto
 Unknown offline input! config['input'] must either be list of offline files (json) or a D4RL-specific InputReader specifier (e.g. 'd4rl.hopper-medium-v0').

```

Do you know id there is any oother trick around this?

Best,  
Frida

---

<div class="post-metadata">

**Author:** ![Lars\_Simon\_Zehnder](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lars_simon_zehnder/32/1185_2.png) [@Lars\_Simon\_Zehnder](https://discuss.ray.io/u/Lars_Simon_Zehnder)\
**Post date:** [November 24, 2021, 8:50am UTC](https://discuss.ray.io/t/roll-out-cql-policy/4216/4 "2021-11-24T08:50:37Z")

</div>

Hi @fksvensson ,

in version 1.5.0 there has been an enhancement of the Input API:

```auto
Added new “input API” for customizing offline datasets (shoutout to Julius F.). (#16957)

```

See [[rllib] Enhancements to Input API for customizing offline datasets by juliusfrost · Pull Request #16957 · ray-project/ray · GitHub](https://github.com/ray-project/ray/pull/16957) for the issue that has been included.

You still can create a virtual environment and test out the newest ray version.

Regarding your issue its hard to make remote guesses. Can you show your `config` here?

Best,  
Simon

---

<div class="post-metadata">

**Author:** ![fksvensson](https://avatars.discourse-cdn.com/v4/letter/f/839c29/32.png) [@fksvensson](https://discuss.ray.io/u/fksvensson)\
**Post date:** [November 24, 2021, 2:08pm UTC](https://discuss.ray.io/t/roll-out-cql-policy/4216/5 "2021-11-24T14:08:33Z")

</div>

Hi again!

I am using Pipenv and it seems like I cannot instal anythin \>1.5.2. I did try with that and it is incompatble with the input reader that I have created for the data that I have.

So when i first aked the question, I loaded the config that I have trained the data with, but changed the input to ‘sampler’. This is the config:

```auto
{'Q_model': {'fcnet_activation': 'relu', 'fcnet_hiddens': [256, 256]}, 'bc_iters': 0, 'clip_actions': True, 'env': <class ' __main__.OTADummy'>, 'evaluation_config': {}, 'evaluation_interval': None, 'evaluation_num_workers': 0, 'framework': 'tfe', 'horizon': 200, 'input': 'sampler', 'input_evaluation': [], 'learning_starts': 256, 'metrics_smoothing_episodes': 5, 'n_step': 3, 'no_done_at_end': True, 'normalize_actions': True, 'num_gpus': 0, 'num_workers': 0, 'optimization': {'actor_learning_rate': 0.0003, 'critic_learning_rate': 0.0003, 'entropy_learning_rate': 0.0003}, 'policy_model': {'fcnet_activation': 'relu', 'fcnet_hiddens': [256, 256]}, 'prioritized_replay': False, 'rollout_fragment_length': 1, 'soft_horizon': True, 'target_entropy': 'auto', 'target_network_update_freq': 1, 'tau': 0.005, 'timesteps_per_iteration': 1000, 'train_batch_size': 256}

```

When you suggested commenting out ‘sampler’, I skipped the step of loading the params entirely, started with an empty config and only set the env, resulting in:

```auto
{'env': <class ' __main__.OTADummy'>}

```

Both of these get the mentioned error.

Best,  
Frida

---

<div class="post-metadata">

**Author:** ![Lars\_Simon\_Zehnder](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lars_simon_zehnder/32/1185_2.png) [@Lars\_Simon\_Zehnder](https://discuss.ray.io/u/Lars_Simon_Zehnder)\
**Post date:** [November 24, 2021, 6:47pm UTC](https://discuss.ray.io/t/roll-out-cql-policy/4216/6 "2021-11-24T18:47:51Z")

</div>

Hi @fksvensson ,

I guess that the Python version might be the problem. I would use `pyenv` (if you are on MacOS-X you can use `brew` to install it. Pull Python 3.9 and then use

```auto
pyenv install 3.9.4
mkdir ray-1.8.0 
cd ray-1.8.0
pyenv local 3.9.4
python -m venv -venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install numpy tensorflow
python -m pip install "ray[default]" 

```

and you should be good to go. ThatÄs the setup I use, so I can choose python versions and install with these versions the packages in a virtual env.

Regarding your problem:

1. Have you registered your input reader via `register_input("custom_input", input_creator)`? before requesting it in your `config`?
2. Have you tested your input reader? Does it really read in the samples and returns a `SampleBatch` as needed?
3. How did you write your ouputs? Did you use a specific ouput writer you coded yourself or did you use the default?
4. Do you actually need the input or can you deploy your policy via running the environment in the background?

I actually do not understand why you include all the training parameters when you want to do evaluation. Evaluation can be done quite easily by running

```auto
rllib rollout \
    ~/ray_results/default/DQN_CartPole-v0_0upjmdgr0/checkpoint_1/checkpoint-1 \
    --run DQN --env CartPole-v0 --steps 10000

```

So you use your checkpoint then the policy (`DQN` here) then your `env` and the number of steps you want to evaluate. The workers sample experiences from the environment using the trained policy - there is no need for an input sampler. You can also do custom evaluation as shown in this [example](https://github.com/ray-project/ray/blob/master/rllib/examples/custom_eval.py) here.

If you still need something else, you can post your code here and we can take a look at it.

Best, Simon

---

<div class="post-metadata">

**Author:** ![mannyv](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/mannyv/32/606_2.png) [@mannyv](https://discuss.ray.io/u/mannyv)\
**Post date:** [November 24, 2021, 8:41pm UTC](https://discuss.ray.io/t/roll-out-cql-policy/4216/7 "2021-11-24T20:41:01Z")

</div>

Hi @fksvensson,

CQL was not designed to run with sampled data; only offline datasets. An easy way to do inference with it is to switch to SAC. The models for both should be the same. If you would rather hack CQL to accept a sampler then have a a look at this issue:

> <https://github.com/ray-project/ray/issues/15138>
>
> \### What is the problem?
> Ray version: nightly
> 
> Perhaps the Ray team is aware …of this but just in case they are not.
> CQL as it exists in master right now is not saving data in the replay buffer and so is not training the policy. 
> 
> This is coming from here
> https://github.com/ray-project/ray/blob/64cc092959264cea42099690f13eceb6036e34d8/rllib/agents/cql/cql.py#L63-L71
> 
> https://github.com/ray-project/ray/blob/64cc092959264cea42099690f13eceb6036e34d8/rllib/agents/cql/cql.py#L98-L100
> 
> There are a couple other little things too:
> 1.) This comment is literally the exact opposite. There is a torch implementation but not a tf one
> https://github.com/ray-project/ray/blob/64cc092959264cea42099690f13eceb6036e34d8/rllib/agents/cql/tests/test\_cql.py#L30-L32
> 2.) There is no learn\_on\_batch in the CQL default config so a user cannot add it to config to use one.
> https://github.com/ray-project/ray/blob/64cc092959264cea42099690f13eceb6036e34d8/rllib/agents/cql/cql.py#L113
> 3.) The default framework in the config is "tf" even though there is no implementation for it. Wouldn't it make more sense to make it "torch"?
> 
> \### Reproduction (REQUIRED)
> \`\`\`python
> import ray
> from ray.rllib.agents.cql import cql
> from ray.rllib.agents.cql.cql import NoOpReplayBuffer
> from ray.rllib.agents.dqn.dqn import calculate\_rr\_weights
> from ray.rllib.execution import \\
> ParallelRollouts, \\
> Replay, \\
> TrainOneStep, \\
> UpdateTargetNetwork, \\
> Concurrently, \\
> StandardMetricsReporting, \\
> StoreToReplayBuffer
> from ray.rllib.execution.replay\_buffer import LocalReplayBuffer
> from ray.rllib.policy.policy import LEARNER\_STATS\_KEY
> from ray.rllib.utils import framework\_iterator
> 
> ray.init(local\_mode=True)
> config = cql.CQL\_DEFAULT\_CONFIG.copy()
> config\["num\_workers"\] = 0 # Run locally.
> config\["twin\_q"\] = True
> config\["clip\_actions"\] = False
> config\["normalize\_actions"\] = True
> config\["learning\_starts"\] = 0
> config\["rollout\_fragment\_length"\] = 100
> config\["train\_batch\_size"\] = 10
> config\["framework"\] = "torch"
> num\_iterations = 2
> 
> def before\_learn\_on\_batch(multi\_agent\_batch, workers, config):
> print(f"sample\_batch count is: {multi\_agent\_batch.count}")
> print(f"sample\_batch keys are: {multi\_agent\_batch.policy\_batches.keys()}")
> return multi\_agent\_batch
> 
> def execution\_plan(workers, config):
> if config.get("prioritized\_replay"):
> prio\_args = {
> "prioritized\_replay\_alpha": config\["prioritized\_replay\_alpha"\],
> "prioritized\_replay\_beta": config\["prioritized\_replay\_beta"\],
> "prioritized\_replay\_eps": config\["prioritized\_replay\_eps"\],
> }
> else:
> prio\_args = {}
> 
> local\_replay\_buffer = LocalReplayBuffer(
> num\_shards=1,
> learning\_starts=config\["learning\_starts"\],
> buffer\_size=config\["buffer\_size"\],
> replay\_batch\_size=config\["train\_batch\_size"\],
> replay\_mode=config\["multiagent"\]\["replay\_mode"\],
> replay\_sequence\_length=config.get("replay\_sequence\_length", 1),
> \*\*prio\_args)
> 
> global replay\_buffer
> replay\_buffer = local\_replay\_buffer
> 
> rollouts = ParallelRollouts(workers, mode="bulk\_sync")
> 
> store\_op = rollouts.for\_each(
> NoOpReplayBuffer(local\_buffer=local\_replay\_buffer))
> #store\_op = rollouts.for\_each(
> # StoreToReplayBuffer(local\_buffer=local\_replay\_buffer))
> def update\_prio(item):
> samples, info\_dict = item
> if config.get("prioritized\_replay"):
> prio\_dict = {}
> for policy\_id, info in info\_dict.items():
> td\_error = info.get("td\_error",
> info\[LEARNER\_STATS\_KEY\].get("td\_error"))
> prio\_dict\[policy\_id\] = (samples.policy\_batches\[policy\_id\]
> .data.get("batch\_indexes"), td\_error)
> local\_replay\_buffer.update\_priorities(prio\_dict)
> return info\_dict
> 
> #post\_fn = config.get("before\_learn\_on\_batch") or (lambda b, \*a: b)
> post\_fn = before\_learn\_on\_batch
> replay\_op = Replay(local\_buffer=local\_replay\_buffer) \\
> .for\_each(lambda x: post\_fn(x, workers, config)) \\
> .for\_each(TrainOneStep(workers)) \\
> .for\_each(update\_prio) \\
> .for\_each(UpdateTargetNetwork(
> workers, config\["target\_network\_update\_freq"\]))
> 
> train\_op = Concurrently(
> \[store\_op, replay\_op\],
> mode="round\_robin",
> output\_indexes=\[1\],
> round\_robin\_weights=calculate\_rr\_weights(config))
> 
> return StandardMetricsReporting(train\_op, workers, config)
> 
> 
> trainer = cql.CQLTrainer.with\_updates(execution\_plan=execution\_plan)(config=config, env="MountainCarContinuous-v0")
> for i in range(num\_iterations):
> print(f"Iteration: {i}")
> trainer.train()
> trainer.stop()
> \`\`\`
> 
> Output looks like this:
> \`\`\`
> Iteration: 0
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\[\])
> Iteration: 1
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\[\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\[\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\[\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\[\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\[\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\[\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\[\])
> \`\`\`
> 
> A fix that has sample data to train on is to change store\_op. I have not verified that correct training actually occurs.
> \`\`\`python
> \#store\_op = rollouts.for\_each(
> # NoOpReplayBuffer(local\_buffer=local\_replay\_buffer))
> store\_op = rollouts.for\_each(
> StoreToReplayBuffer(local\_buffer=local\_replay\_buffer))
> \`\`\`
> 
> Output looks like this:
> \`\`\`
> Iteration: 0
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\['default\_policy'\])
> Iteration: 1
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\['default\_policy'\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\['default\_policy'\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\['default\_policy'\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\['default\_policy'\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\['default\_policy'\])
> sample\_batch count is: 10
> sample\_batch keys are: dict\_keys(\['default\_policy'\])
> \`\`\`
> 
> \- \[x \] I have verified my script runs in a clean environment and reproduces the issue.
> \- \[x \] I have verified the issue also occurs with the \[latest wheels\](https://docs.ray.io/en/master/installation.html).

---

<div class="post-metadata">

**Author:** ![Lars\_Simon\_Zehnder](https://sea2.discourse-cdn.com/flex020/user_avatar/discuss.ray.io/lars_simon_zehnder/32/1185_2.png) [@Lars\_Simon\_Zehnder](https://discuss.ray.io/u/Lars_Simon_Zehnder)\
**Post date:** [November 24, 2021, 9:19pm UTC](https://discuss.ray.io/t/roll-out-cql-policy/4216/8 "2021-11-24T21:19:35Z")

</div>

@mannyv ,

interesting insight. So there is actually no chance to evaluate the **already trained** policy of CQL, but by using another trainer? From the viewpoint of a practitioner I cannot completely follow this decision as usually I collect data from an environment with a behavioral policy. Then use offline learning and when I trained my policy I want to evaluate it later also inside of the environment (at least for some interesting hours 😄).

However, as far as I understood, it is only the replay buffer that does not get filled? Can we not simply turn it off as we want to make online evaluation anyway? And the trainer pulls simply samples from the rollout workers?

---

<div class="post-metadata">

**Author:** ![fksvensson](https://avatars.discourse-cdn.com/v4/letter/f/839c29/32.png) [@fksvensson](https://discuss.ray.io/u/fksvensson)\
**Post date:** [November 25, 2021, 12:59pm UTC](https://discuss.ray.io/t/roll-out-cql-policy/4216/9 "2021-11-25T12:59:30Z")

</div>

Thank you both for your comments, I’m really glad to see that offline RL is an engaging subject.

> [@Lars\_Simon\_Zehnder](#):
>
> Python 3.9

According to [https://pypi.org/pypi/ray/1.8.0/json](https://pypi.org/pypi/ray/1.8.0/json) it seems like it should be compatible with python 3.7, but it looks like it does not support my current macOS

```auto
curl -sL https://pypi.org/pypi/ray/1.8.0/json | jq '.releases["1.8.0"][].filename' | grep -i macos
"ray-1.8.0-cp36-cp36m-macosx_10_15_intel.whl"
"ray-1.8.0-cp37-cp37m-macosx_10_15_intel.whl"
"ray-1.8.0-cp38-cp38-macosx_10_15_x86_64.whl"
"ray-1.8.0-cp38-cp38-macosx_11_0_arm64.whl"
"ray-1.8.0-cp39-cp39-macosx_10_15_x86_64.whl"
"ray-1.8.0-cp39-cp39-macosx_11_0_arm64.whl"

```

As I am working on 10\_14 right now. Even if I were to update it, I would update it to 10\_11 on an intel machine, which is not supported. If i do need Ray 1.8.0, I would use a container, but is seems like that might not be my problem right now.

@mannyv, your suggestion does sound interesting, I did try

`  
trainer = SACTrainer(config=config)  
trainer.restore(checkpoint\_path)

`  
with my CQL chekpoint, but it looks like the weights are not matching. Do you think I could make them match by adapting hyperparameters or did I misunderstand you?

I will try the hack and get back to you.

Thanks,  
Frida
