BatchMAB#
- class statrl.settings.bandits.stochastic.batch.environment.BatchMAB(mab, batchsize)[source]#
Bases:
StochasticBanditEnvWrap a stochastic bandit into a batched one.
Turns any
StochasticBanditEnvinto an environment whosestep()consumes a whole list of arms and returns a list of rewards.A round is a batch, not a pull, so the horizon of an experiment counts batches.
- Parameters:
mab (StochasticBanditEnv) – The underlying bandit supplying the arms.
batchsize (callable or sequence of int) – Batch schedule: either
batchsize(round) -> int, or a sequence indexed by round which falls back to1once exhausted. A plain sequence is picklable, and so survives the process pool used bymulticoreRuns(); a lambda does not.
Examples
>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit >>> env = BatchMAB(BernoulliBandit([0.2, 0.9]), batchsize=[2, 4, 8]) >>> info = env.reset() >>> info["nextbatchsize"] 2 >>> rewards, info = env.step([0, 1]) >>> len(rewards), info["nextbatchsize"] (2, 4)
Methods
__init__(mab, batchsize)close()Release every attached renderer at the end of a rendered run.
expected_reward(arm)Mean reward of an arm, for regret accounting only.
get_wrapper_attr(name)Gets the attribute name from the environment.
has_wrapper_attr(name)Checks if the attribute name exists in the environment.
render([mode])Forward the last
(arm, reward)pair to every attached renderer.reset([seed, options])Start a new run and announce the size of the first batch.
set_wrapper_attr(name, value, *[, force])Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.
step(action)Pull every arm of one batch and return all their rewards.
Attributes
meansMean reward of every arm.
metadatanp_randomReturns the environment's internal
_np_randomthat if not set will initialise with a random seed.np_random_seedReturns the environment's internal
_np_random_seedthat if not set will first initialise with a random int as seed.number_armsNumber of available arms.
optimal_armIndex of the best arm.
optimal_meanMean reward of the best arm, \(\mu^\star\).
render_modespecunwrappedReturns the base non-wrapped environment.
action_spaceobservation_space- reset(seed=None, options=None)[source]#
Start a new run and announce the size of the first batch.
- Parameters:
seed (int, optional) – Seed for the underlying bandit’s generator.
options (dict, optional) – Unused; accepted for
gymnasium.Envcompatibility.
- Returns:
{"nextbatchsize": int, "mean": 0}. Unlike the unbatchedreset(), an info dict is returned rather than a dummy observation, because the agent cannot act without knowing the batch size.- Return type:
- step(action)[source]#
Pull every arm of one batch and return all their rewards.
- Parameters:
action (list of int) – Arms to pull, one per slot of the current batch. Its length must equal the
nextbatchsizeannounced by the previous call.- Returns:
batchreward (list of float) – Reward of each pull, in the order the arms were given.
info (dict) –
{"nextbatchsize": int, "mean": float}, wheremeanis the sum of the true means of the arms pulled — the batch’s expected score, which the interaction loop accumulates for regret.
- Raises:
AssertionError – If
actiondoes not have exactlynextbatchsizeentries.