Oracle#
- class statrl.settings.bandits.stochastic.batch.agents._Oracle.Oracle(env)[source]#
Bases:
BatchBanditAgentBaseline filling every batch with the best arm.
- Parameters:
env (BatchMAB) – The environment, read for its
optimal_arm.
Examples
>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit >>> from statrl.settings.bandits.stochastic.batch.environment import BatchMAB >>> Oracle(BatchMAB(BernoulliBandit([0.2, 0.9]), [3])).batchplay(3) [1, 1, 1]
Methods
__init__(env)batchplay(batchsize)Fill the whole batch with the optimal arm.
batchupdate(batcharm, batchreward)Ignore the batch's rewards (the oracle has nothing to learn).
play()Return the arm with the highest mean.
reset()Start a new run.
update(arm, reward)Ignore the observed reward (the oracle has nothing to learn).
Attributes
The optimal arm, as a one-element list.
- batchupdate(batcharm, batchreward)[source]#
Ignore the batch’s rewards (the oracle has nothing to learn).
- property policy#
The optimal arm, as a one-element list.
Written to the experiment logfile by
runLargeMulticoreExperiment().