BatchGaussianBandit#
- class statrl.settings.bandits.stochastic.batch.envs.parametric.BatchGaussianBandit(means, vars, batchschedule='constant', name='BMAB-Gaussian')[source]#
Bases:
BatchMABBatched Gaussian bandit, built from means and a named batch schedule.
- Parameters:
means (array-like of float or str) – Arm means, or a key of
mean_catalogue("simple4","simple6") naming a stock instance.vars (array-like of float) – Variance of each arm.
batchschedule (str, default='constant') – Key of
schedule_catalogue—"constant","linear","quadratic","cubic","exp","doubleexp","abrupt","exotic"— or"baba,<horizon>"to use the BABA grid for that horizon.name (str, default='BMAB-Gaussian') – Accepted but unused: the instance name comes from
BatchMAB, which derives it from the wrapped bandit and the batch sizes.
- Raises:
KeyError – If
batchscheduleor a stringmeansnames no catalogue entry.
Methods
__init__(means, vars[, batchschedule, name])close()Release every attached renderer at the end of a rendered run.
expected_reward(arm)Mean reward of an arm, for regret accounting only.
get_wrapper_attr(name)Gets the attribute name from the environment.
has_wrapper_attr(name)Checks if the attribute name exists in the environment.
render([mode])Forward the last
(arm, reward)pair to every attached renderer.reset([seed, options])Start a new run and announce the size of the first batch.
set_wrapper_attr(name, value, *[, force])Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.
step(action)Pull every arm of one batch and return all their rewards.
Attributes
meansMean reward of every arm.
metadatanp_randomReturns the environment's internal
_np_randomthat if not set will initialise with a random seed.np_random_seedReturns the environment's internal
_np_random_seedthat if not set will first initialise with a random int as seed.number_armsNumber of available arms.
optimal_armIndex of the best arm.
optimal_meanMean reward of the best arm, \(\mu^\star\).
render_modespecunwrappedReturns the base non-wrapped environment.
action_spaceobservation_space