BanditAgent#

class statrl.settings.bandits.stochastic.anytime.agent.BanditAgent(name, seed=1)[source]#

Bases: ABC

Base class for anytime stochastic bandit agents.

Subclasses implement the protocol shared by every agent in this setting: reset() to start an independent run, select_arm() to choose an arm, and update() to learn from the observed reward.

Parameters:
  • name (str) – Label used in logfiles and plot legends. Must be unique within an experiment, since dump() builds filenames from it.

  • seed (int, default=1) – Seed for the agent’s own randomness. reset() re-derives np_random from it, so replicates are reproducible.

np_random#

Agent-local generator, available after the first reset().

Type:

numpy.random.Generator

See also

statrl.settings.bandits.stochastic.knownhorizon.agent.BanditAgent

The counterpart whose reset receives the horizon.

Notes

Agents are deep-copied once per replicate by multicoreRuns(), so an agent may hold arbitrary state as long as reset() fully reinitializes it.

Methods

__init__(name[, seed])

reset()

Start a new independent run.

select_arm()

Choose the arm to pull next.

update(arm, reward)

Learn from the reward observed for the arm just pulled.

reset()[source]#

Start a new independent run.

Reseeds np_random and clears any statistics accumulated by a previous run.

abstractmethod select_arm()[source]#

Choose the arm to pull next.

Returns:

Index of the selected arm, in range(env.number_arms).

Return type:

int

abstractmethod update(arm, reward)[source]#

Learn from the reward observed for the arm just pulled.

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Reward sampled by step(). Only this realized reward is observed — never the arm’s mean.