BanditInteraction#
- class statrl.settings.bandits.stochastic.knownhorizon.interaction.BanditInteraction[source]#
Bases:
InteractionInteraction loop for the known-horizon stochastic bandit setting.
Identical to the anytime loop except that the horizon is passed to
learner.reset(horizon)instead of being withheld.See also
statrl.settings.bandits.stochastic.anytime.interaction.BanditInteractionThe anytime counterpart.
Examples
>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit >>> from statrl.settings.bandits.stochastic.anytime.agents.IMED import IMED >>> from statrl.settings.bandits.stochastic.knownhorizon.wrappers.wrapper_anytime_knownhorizon import ( ... AnytimeToKnownHorizonAgentWrapper) >>> env = BernoulliBandit([0.2, 0.9, 0.5]) >>> agent = AnytimeToKnownHorizonAgentWrapper(IMED(env.number_arms)) >>> BanditInteraction().run(env, agent, horizon=50).shape (50,)
Methods
__init__()renderrun(env, learner, horizon)Run one interaction, printing each round to stdout.
run(env, learner, horizon)Run one interaction and return its cumulative expected score.
Attributes
Axis labels
(x, y)for the regret plots.- renderrun(env, learner, horizon)[source]#
Run one interaction, printing each round to stdout.
- Parameters:
env (StochasticBanditEnv) – The bandit instance. A
Textrendereris appended to itsrenderers, so calling this twice on the same environment prints every round twice.learner (BanditAgent) – The agent.
horizon (int) – Number of rounds to play.
- run(env, learner, horizon)[source]#
Run one interaction and return its cumulative expected score.
- Parameters:
env (StochasticBanditEnv) – The bandit instance; the same class as in the anytime setting, since only the agent interface differs between the two.
learner (BanditAgent) – The agent, reset with
horizonso it can plan against it.horizon (int) – Number of rounds to play.
- Returns:
Cumulative sum of the expected rewards of the arms played.
- Return type:
ndarray of shape (horizon,)