BatchBanditInteraction#

class statrl.settings.bandits.stochastic.batch.interaction.BatchBanditInteraction[source]#

Bases: Interaction

Interaction loop for the batched bandit setting.

Drives batchplay -> step -> batchupdate. One round is one batch, so horizon counts batches rather than pulls.

Examples

>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit
>>> from statrl.settings.bandits.stochastic.batch.environment import BatchMAB
>>> from statrl.settings.bandits.stochastic.batch.agents.BIMED import BIMED
>>> env = BatchMAB(BernoulliBandit([0.2, 0.9, 0.5]), batchsize=[4] * 20)
>>> BatchBanditInteraction().run(env, BIMED(3), horizon=20).shape
(20,)

Methods

__init__()

renderrun(env, learner, horizon)

Run one interaction with rendering enabled.

run(env, learner, horizon)

Run one interaction and return its cumulative expected score.

Attributes

plotlabels

Axis labels (x, y) for the regret plots.

property plotlabels#

Axis labels (x, y) for the regret plots.

The x-axis counts batches, not pulls, so it is labelled by episode.

Type:

tuple of (str, str)

renderrun(env, learner, horizon)[source]#

Run one interaction with rendering enabled.

Parameters:
  • env (BatchMAB) – The batched bandit instance.

  • learner (BatchBanditAgent) – The agent.

  • horizon (int) – Number of batches to play.

run(env, learner, horizon)[source]#

Run one interaction and return its cumulative expected score.

Parameters:
  • env (BatchMAB) – The batched bandit instance.

  • learner (BatchBanditAgent) – The agent.

  • horizon (int) – Number of batches to play. The number of pulls is the sum of the batch sizes over those rounds.

Returns:

Cumulative sum of each batch’s expected score, i.e. the summed true means of the arms pulled in it. Entry t therefore covers every pull up to the end of batch t.

Return type:

ndarray of shape (horizon,)