Random#

class statrl.settings.bandits.stochastic.anytime.agents._Random.Random(env)[source]#

Bases: BanditAgent

Uniform exploration: pull an arm uniformly at random every round.

Parameters:

env (StochasticBanditEnv) – The environment, read only for its number of arms.

See also

statrl.settings.bandits.stochastic.anytime.agents._Oracle.Oracle

The opposite baseline, which always plays the best arm.

Examples

>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit
>>> agent = Random(BernoulliBandit([0.2, 0.9, 0.5]))
>>> agent.reset()
>>> agent.select_arm() in (0, 1, 2)
True

Methods

__init__(env)

reset()

Start a new run.

select_arm()

Draw an arm uniformly at random.

update(arm, reward)

Ignore the observed reward (this agent does not learn).

reset()[source]#

Start a new run. No statistics are kept, so this does nothing.

select_arm()[source]#

Draw an arm uniformly at random.

Returns:

An index drawn uniformly from range(env.number_arms).

Return type:

int

update(arm, reward)[source]#

Ignore the observed reward (this agent does not learn).

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Observed reward.