Oracle#

class statrl.settings.bandits.stochastic.batch.agents._Oracle.Oracle(env)[source]#

Bases: BatchBanditAgent

Baseline filling every batch with the best arm.

Parameters:

env (BatchMAB) – The environment, read for its optimal_arm.

Examples

>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit
>>> from statrl.settings.bandits.stochastic.batch.environment import BatchMAB
>>> Oracle(BatchMAB(BernoulliBandit([0.2, 0.9]), [3])).batchplay(3)
[1, 1, 1]

Methods

__init__(env)

batchplay(batchsize)

Fill the whole batch with the optimal arm.

batchupdate(batcharm, batchreward)

Ignore the batch's rewards (the oracle has nothing to learn).

play()

Return the arm with the highest mean.

reset()

Start a new run.

update(arm, reward)

Ignore the observed reward (the oracle has nothing to learn).

Attributes

policy

The optimal arm, as a one-element list.

batchplay(batchsize)[source]#

Fill the whole batch with the optimal arm.

Parameters:

batchsize (int) – Number of pulls in this batch.

Returns:

batchsize copies of env.optimal_arm.

Return type:

list of int

batchupdate(batcharm, batchreward)[source]#

Ignore the batch’s rewards (the oracle has nothing to learn).

Parameters:
  • batcharm (list of int) – The arms that were pulled.

  • batchreward (list of float) – The rewards observed for them.

play()[source]#

Return the arm with the highest mean.

Returns:

env.optimal_arm.

Return type:

int

property policy#

The optimal arm, as a one-element list.

Written to the experiment logfile by runLargeMulticoreExperiment().

Type:

list of int

reset()[source]#

Start a new run. The oracle keeps no statistics, so this is a no-op.

update(arm, reward)[source]#

Ignore the observed reward (the oracle has nothing to learn).

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Observed reward.