BatchBernoulliBandit#

class statrl.settings.bandits.stochastic.batch.envs.parametric.BatchBernoulliBandit(means, batchschedule='constant', name='BMAB-Bernoulli')[source]#

Bases: BatchMAB

Batched Bernoulli bandit, built from means and a named batch schedule.

Parameters:
  • means (array-like of float or str) – Arm success probabilities, or a key of mean_catalogue.

  • batchschedule (str, default='constant') – Key of schedule_catalogue, or "baba,<horizon>".

  • name (str, default='BMAB-Bernoulli') – Accepted but unused; see BatchGaussianBandit.

Raises:

KeyError – If batchschedule or a string means names no catalogue entry.

Methods

__init__(means[, batchschedule, name])

close()

Release every attached renderer at the end of a rendered run.

expected_reward(arm)

Mean reward of an arm, for regret accounting only.

get_wrapper_attr(name)

Gets the attribute name from the environment.

has_wrapper_attr(name)

Checks if the attribute name exists in the environment.

render([mode])

Forward the last (arm, reward) pair to every attached renderer.

reset([seed, options])

Start a new run and announce the size of the first batch.

set_wrapper_attr(name, value, *[, force])

Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.

step(action)

Pull every arm of one batch and return all their rewards.

Attributes

means

Mean reward of every arm.

metadata

np_random

Returns the environment's internal _np_random that if not set will initialise with a random seed.

np_random_seed

Returns the environment's internal _np_random_seed that if not set will first initialise with a random int as seed.

number_arms

Number of available arms.

optimal_arm

Index of the best arm.

optimal_mean

Mean reward of the best arm, \(\mu^\star\).

render_mode

spec

unwrapped

Returns the base non-wrapped environment.

action_space

observation_space