BanditInteraction#

class statrl.settings.bandits.stochastic.anytime.interaction.BanditInteraction[source]#

Bases: Interaction

Interaction loop for the anytime stochastic bandit setting.

Drives select_arm -> step -> update for a fixed number of rounds and returns the cumulative expected score.

See also

statrl.settings.bandits.stochastic.knownhorizon.interaction.BanditInteraction

The counterpart that passes the horizon to reset.

Examples

>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit
>>> from statrl.settings.bandits.stochastic.anytime.agents.IMED import IMED
>>> env = BernoulliBandit([0.2, 0.9, 0.5])
>>> scores = BanditInteraction().run(env, IMED(env.number_arms), horizon=100)
>>> scores.shape
(100,)

Methods

__init__()

renderrun(env, learner, horizon)

Run one interaction, printing each round to stdout.

run(env, learner, horizon)

Run one interaction and return its cumulative expected score.

Attributes

plotlabels

Axis labels (x, y) for the regret plots.

property plotlabels#

Axis labels (x, y) for the regret plots.

Type:

tuple of (str, str)

renderrun(env, learner, horizon)[source]#

Run one interaction, printing each round to stdout.

Same loop as run(), but a Textrenderer is attached to the environment and no score is returned. Intended for inspecting short runs by eye, not for benchmarking.

Parameters:
  • env (StochasticBanditEnv) – The bandit instance. Its renderers list is overwritten.

  • learner (BanditAgent) – The agent.

  • horizon (int) – Number of rounds to play.

run(env, learner, horizon)[source]#

Run one interaction and return its cumulative expected score.

Parameters:
  • env (StochasticBanditEnv) – The bandit instance. Reset at the start of the run.

  • learner (BanditAgent) – The agent. Reset at the start of the run, so a single instance can be reused across replicates.

  • horizon (int) – Number of rounds to play.

Returns:

Cumulative sum of the expected rewards of the arms played, i.e. entry t is \(\sum_{s \le t} \mu_{a_s}\). Accumulating means rather than realized rewards removes the reward noise from the regret curve.

Return type:

ndarray of shape (horizon,)