BatchBernoulliBandit#
- class statrl.settings.bandits.stochastic.batch.envs.parametric.BatchBernoulliBandit(means, batchschedule='constant', name='BMAB-Bernoulli')[source]#
Bases:
BatchMABBatched Bernoulli bandit, built from means and a named batch schedule.
- Parameters:
means (array-like of float or str) – Arm success probabilities, or a key of
mean_catalogue.batchschedule (str, default='constant') – Key of
schedule_catalogue, or"baba,<horizon>".name (str, default='BMAB-Bernoulli') – Accepted but unused; see
BatchGaussianBandit.
- Raises:
KeyError – If
batchscheduleor a stringmeansnames no catalogue entry.
Methods
__init__(means[, batchschedule, name])close()Release every attached renderer at the end of a rendered run.
expected_reward(arm)Mean reward of an arm, for regret accounting only.
get_wrapper_attr(name)Gets the attribute name from the environment.
has_wrapper_attr(name)Checks if the attribute name exists in the environment.
render([mode])Forward the last
(arm, reward)pair to every attached renderer.reset([seed, options])Start a new run and announce the size of the first batch.
set_wrapper_attr(name, value, *[, force])Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.
step(action)Pull every arm of one batch and return all their rewards.
Attributes
meansMean reward of every arm.
metadatanp_randomReturns the environment's internal
_np_randomthat if not set will initialise with a random seed.np_random_seedReturns the environment's internal
_np_random_seedthat if not set will first initialise with a random int as seed.number_armsNumber of available arms.
optimal_armIndex of the best arm.
optimal_meanMean reward of the best arm, \(\mu^\star\).
render_modespecunwrappedReturns the base non-wrapped environment.
action_spaceobservation_space