StochasticBanditEnv#
- class statrl.settings.bandits.stochastic.anytime.environment.StochasticBanditEnv(rewarddistributions, name, last=(None, 0.0))[source]#
Bases:
EnvStochastic multi-armed bandit environment.
Each arm carries an independent reward distribution;
step()draws one sample from the chosen arm. The environment is stateless: interactions never modify the underlying distributions.- Parameters:
rewarddistributions (list) – One distribution per arm. Each must expose a
meanattribute and asample()method — seeArm.name (str) – Label used in logfiles, plot titles, and dump filenames.
last (tuple of (int or None, float), default=(None, 0.0)) – Most recent
(arm, reward)pair, consumed by the renderers.
- renderers#
Renderers notified on every
render()call. Empty by default;renderrun()installs aTextrenderer.- Type:
See also
statrl.settings.bandits.stochastic.anytime.envs.parametric.BernoulliBanditFactory for a Bernoulli instance.
statrl.settings.bandits.stochastic.batch.environment.BatchMABWrapper turning any instance into a batched bandit.
Examples
>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit >>> env = BernoulliBandit([0.2, 0.9, 0.5]) >>> env.number_arms 3 >>> env.optimal_arm 1 >>> _ = env.reset(seed=0) >>> reward = env.step(1) >>> reward in (0.0, 1.0) True
Methods
__init__(rewarddistributions, name[, last])close()Release every attached renderer at the end of a rendered run.
expected_reward(arm)Mean reward of an arm, for regret accounting only.
get_wrapper_attr(name)Gets the attribute name from the environment.
has_wrapper_attr(name)Checks if the attribute name exists in the environment.
render([mode])Forward the last
(arm, reward)pair to every attached renderer.reset([seed, options])Start a new run by reseeding the environment.
set_wrapper_attr(name, value, *[, force])Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.
step(arm)Sample one reward from the given arm.
Attributes
Mean reward of every arm.
metadataReturns the environment's internal
_np_randomthat if not set will initialise with a random seed.np_random_seedReturns the environment's internal
_np_random_seedthat if not set will first initialise with a random int as seed.Number of available arms.
Index of the best arm.
Mean reward of the best arm, \(\mu^\star\).
render_modespecunwrappedReturns the base non-wrapped environment.
action_spaceobservation_space- property means#
Mean reward of every arm.
- property number_arms#
Number of available arms.
- property optimal_arm#
Index of the best arm.
Ties are broken by
numpy.argmax(), i.e. the lowest index wins.- Type:
- property optimal_mean#
Mean reward of the best arm, \(\mu^\star\).
Used to define regret.
- Type:
- render(mode='human')[source]#
Forward the last
(arm, reward)pair to every attached renderer.- Parameters:
mode (str, default='human') – Unused; accepted for
gymnasium.Envcompatibility. Output is whatever the objects inrenderersproduce.
- reset(seed=None, options=None)[source]#
Start a new run by reseeding the environment.
- Parameters:
seed (int, optional) – Seed for the environment’s generator.
Nonedraws a fresh one.options (dict, optional) – Unused; accepted for
gymnasium.Envcompatibility.
- Returns:
The constant dummy observation
0— a bandit is stateless.- Return type: