BatchGaussianBandit#

class statrl.settings.bandits.stochastic.batch.envs.parametric.BatchGaussianBandit(means, vars, batchschedule='constant', name='BMAB-Gaussian')[source]#

Bases: BatchMAB

Batched Gaussian bandit, built from means and a named batch schedule.

Parameters:
  • means (array-like of float or str) – Arm means, or a key of mean_catalogue ("simple4", "simple6") naming a stock instance.

  • vars (array-like of float) – Variance of each arm.

  • batchschedule (str, default='constant') – Key of schedule_catalogue — "constant", "linear", "quadratic", "cubic", "exp", "doubleexp", "abrupt", "exotic" — or "baba,<horizon>" to use the BABA grid for that horizon.

  • name (str, default='BMAB-Gaussian') – Accepted but unused: the instance name comes from BatchMAB, which derives it from the wrapped bandit and the batch sizes.

Raises:

KeyError – If batchschedule or a string means names no catalogue entry.

Methods

__init__(means, vars[, batchschedule, name])

close()

Release every attached renderer at the end of a rendered run.

expected_reward(arm)

Mean reward of an arm, for regret accounting only.

get_wrapper_attr(name)

Gets the attribute name from the environment.

has_wrapper_attr(name)

Checks if the attribute name exists in the environment.

render([mode])

Forward the last (arm, reward) pair to every attached renderer.

reset([seed, options])

Start a new run and announce the size of the first batch.

set_wrapper_attr(name, value, *[, force])

Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.

step(action)

Pull every arm of one batch and return all their rewards.

Attributes

means

Mean reward of every arm.

metadata

np_random

Returns the environment's internal _np_random that if not set will initialise with a random seed.

np_random_seed

Returns the environment's internal _np_random_seed that if not set will first initialise with a random int as seed.

number_arms

Number of available arms.

optimal_arm

Index of the best arm.

optimal_mean

Mean reward of the best arm, \(\mu^\star\).

render_mode

spec

unwrapped

Returns the base non-wrapped environment.

action_space

observation_space