BatchMAB#

class statrl.settings.bandits.stochastic.batch.environment.BatchMAB(mab, batchsize)[source]#

Bases: StochasticBanditEnv

Wrap a stochastic bandit into a batched one.

Turns any StochasticBanditEnv into an environment whose step() consumes a whole list of arms and returns a list of rewards.

A round is a batch, not a pull, so the horizon of an experiment counts batches.

Parameters:
  • mab (StochasticBanditEnv) – The underlying bandit supplying the arms.

  • batchsize (callable or sequence of int) – Batch schedule: either batchsize(round) -> int, or a sequence indexed by round which falls back to 1 once exhausted. A plain sequence is picklable, and so survives the process pool used by multicoreRuns(); a lambda does not.

round#

Index of the next batch, counted from 0 since the last reset().

Type:

int

Examples

>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit
>>> env = BatchMAB(BernoulliBandit([0.2, 0.9]), batchsize=[2, 4, 8])
>>> info = env.reset()
>>> info["nextbatchsize"]
2
>>> rewards, info = env.step([0, 1])
>>> len(rewards), info["nextbatchsize"]
(2, 4)

Methods

__init__(mab, batchsize)

close()

Release every attached renderer at the end of a rendered run.

expected_reward(arm)

Mean reward of an arm, for regret accounting only.

get_wrapper_attr(name)

Gets the attribute name from the environment.

has_wrapper_attr(name)

Checks if the attribute name exists in the environment.

render([mode])

Forward the last (arm, reward) pair to every attached renderer.

reset([seed, options])

Start a new run and announce the size of the first batch.

set_wrapper_attr(name, value, *[, force])

Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.

step(action)

Pull every arm of one batch and return all their rewards.

Attributes

means

Mean reward of every arm.

metadata

np_random

Returns the environment's internal _np_random that if not set will initialise with a random seed.

np_random_seed

Returns the environment's internal _np_random_seed that if not set will first initialise with a random int as seed.

number_arms

Number of available arms.

optimal_arm

Index of the best arm.

optimal_mean

Mean reward of the best arm, \(\mu^\star\).

render_mode

spec

unwrapped

Returns the base non-wrapped environment.

action_space

observation_space

reset(seed=None, options=None)[source]#

Start a new run and announce the size of the first batch.

Parameters:
  • seed (int, optional) – Seed for the underlying bandit’s generator.

  • options (dict, optional) – Unused; accepted for gymnasium.Env compatibility.

Returns:

{"nextbatchsize": int, "mean": 0}. Unlike the unbatched reset(), an info dict is returned rather than a dummy observation, because the agent cannot act without knowing the batch size.

Return type:

dict

step(action)[source]#

Pull every arm of one batch and return all their rewards.

Parameters:

action (list of int) – Arms to pull, one per slot of the current batch. Its length must equal the nextbatchsize announced by the previous call.

Returns:

  • batchreward (list of float) – Reward of each pull, in the order the arms were given.

  • info (dict) – {"nextbatchsize": int, "mean": float}, where mean is the sum of the true means of the arms pulled — the batch’s expected score, which the interaction loop accumulates for regret.

Raises:

AssertionError – If action does not have exactly nextbatchsize entries.