BanditInteraction#
- class statrl.settings.bandits.stochastic.anytime.interaction.BanditInteraction[source]#
Bases:
InteractionInteraction loop for the anytime stochastic bandit setting.
Drives
select_arm->step->updatefor a fixed number of rounds and returns the cumulative expected score.See also
statrl.settings.bandits.stochastic.knownhorizon.interaction.BanditInteractionThe counterpart that passes the horizon to
reset.
Examples
>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit >>> from statrl.settings.bandits.stochastic.anytime.agents.IMED import IMED >>> env = BernoulliBandit([0.2, 0.9, 0.5]) >>> scores = BanditInteraction().run(env, IMED(env.number_arms), horizon=100) >>> scores.shape (100,)
Methods
__init__()renderrun(env, learner, horizon)Run one interaction, printing each round to stdout.
run(env, learner, horizon)Run one interaction and return its cumulative expected score.
Attributes
Axis labels
(x, y)for the regret plots.- renderrun(env, learner, horizon)[source]#
Run one interaction, printing each round to stdout.
Same loop as
run(), but aTextrendereris attached to the environment and no score is returned. Intended for inspecting short runs by eye, not for benchmarking.- Parameters:
env (StochasticBanditEnv) – The bandit instance. Its
rendererslist is overwritten.learner (BanditAgent) – The agent.
horizon (int) – Number of rounds to play.
- run(env, learner, horizon)[source]#
Run one interaction and return its cumulative expected score.
- Parameters:
env (StochasticBanditEnv) – The bandit instance. Reset at the start of the run.
learner (BanditAgent) – The agent. Reset at the start of the run, so a single instance can be reused across replicates.
horizon (int) – Number of rounds to play.
- Returns:
Cumulative sum of the expected rewards of the arms played, i.e. entry
tis \(\sum_{s \le t} \mu_{a_s}\). Accumulating means rather than realized rewards removes the reward noise from the regret curve.- Return type:
ndarray of shape (horizon,)