Oracle#

class statrl.settings.bandits.stochastic.anytime.agents._Oracle.Oracle(env)[source]#

Bases: BanditAgent

Baseline that always plays the best arm.

The oracle knows the arm means and so incurs no regret.

Parameters:

env (StochasticBanditEnv) – The environment.

See also

statrl.settings.bandits.stochastic.anytime.agents._Random.Random

The opposite baseline, which never exploits.

Examples

>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit
>>> oracle = Oracle(BernoulliBandit([0.2, 0.9, 0.5]))
>>> oracle.select_arm()
1

Methods

__init__(env)

reset()

Start a new run.

select_arm()

Play the arm with the highest mean.

update(arm, reward)

Ignore the observed reward (the oracle has nothing to learn).

Attributes

policy

The optimal arm, as a one-element list.

property policy#

The optimal arm, as a one-element list.

Written to the experiment logfile by runLargeMulticoreExperiment().

Type:

list of int

reset()[source]#

Start a new run. The oracle keeps no statistics.

select_arm()[source]#

Play the arm with the highest mean.

Returns:

env.optimal_arm.

Return type:

int

update(arm, reward)[source]#

Ignore the observed reward (the oracle has nothing to learn).

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Observed reward.