BanditAgent#

class statrl.settings.bandits.stochastic.knownhorizon.agent.BanditAgent(name, seed=1)[source]#

Bases: ABC

Base class for horizon-aware stochastic bandit agents.

Identical to the anytime protocol except that reset() receives the horizon, letting an agent tune its behaviour to the number of rounds it will play.

Parameters:
  • name (str) – Label used in logfiles and plot legends.

  • seed (int, default=1) – Seed for the agent’s own randomness.

horizon#

Number of rounds of the current run, set by reset().

Type:

int

np_random#

Agent-local generator, available after the first reset().

Type:

numpy.random.Generator

Methods

__init__(name[, seed])

reset(horizon)

Start a new independent run of known length.

select_arm()

Choose the arm to pull next.

update(arm, reward)

Learn from the reward observed for the arm just pulled.

reset(horizon)[source]#

Start a new independent run of known length.

Parameters:

horizon (int) – Number of rounds that will be played. Stored on horizon and free to be used by the selection rule.

abstractmethod select_arm()[source]#

Choose the arm to pull next.

Returns:

Index of the selected arm, in range(env.number_arms).

Return type:

int

update(arm, reward)[source]#

Learn from the reward observed for the arm just pulled.

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Reward observed for that arm.