Oracle#
- class statrl.settings.bandits.stochastic.anytime.agents._Oracle.Oracle(env)[source]#
Bases:
BanditAgentBaseline that always plays the best arm.
The oracle knows the arm means and so incurs no regret.
- Parameters:
env (StochasticBanditEnv) – The environment.
See also
statrl.settings.bandits.stochastic.anytime.agents._Random.RandomThe opposite baseline, which never exploits.
Examples
>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit >>> oracle = Oracle(BernoulliBandit([0.2, 0.9, 0.5])) >>> oracle.select_arm() 1
Methods
__init__(env)reset()Start a new run.
Play the arm with the highest mean.
update(arm, reward)Ignore the observed reward (the oracle has nothing to learn).
Attributes
The optimal arm, as a one-element list.
- property policy#
The optimal arm, as a one-element list.
Written to the experiment logfile by
runLargeMulticoreExperiment().