StochasticBanditEnv#

class statrl.settings.bandits.stochastic.anytime.environment.StochasticBanditEnv(rewarddistributions, name, last=(None, 0.0))[source]#

Bases: Env

Stochastic multi-armed bandit environment.

Each arm carries an independent reward distribution; step() draws one sample from the chosen arm. The environment is stateless: interactions never modify the underlying distributions.

Parameters:
  • rewarddistributions (list) – One distribution per arm. Each must expose a mean attribute and a sample() method — see Arm.

  • name (str) – Label used in logfiles, plot titles, and dump filenames.

  • last (tuple of (int or None, float), default=(None, 0.0)) – Most recent (arm, reward) pair, consumed by the renderers.

renderers#

Renderers notified on every render() call. Empty by default; renderrun() installs a Textrenderer.

Type:

list

np_random#

Environment-local generator, available after reset().

Type:

numpy.random.Generator

See also

statrl.settings.bandits.stochastic.anytime.envs.parametric.BernoulliBandit

Factory for a Bernoulli instance.

statrl.settings.bandits.stochastic.batch.environment.BatchMAB

Wrapper turning any instance into a batched bandit.

Examples

>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit
>>> env = BernoulliBandit([0.2, 0.9, 0.5])
>>> env.number_arms
3
>>> env.optimal_arm
1
>>> _ = env.reset(seed=0)
>>> reward = env.step(1)
>>> reward in (0.0, 1.0)
True

Methods

__init__(rewarddistributions, name[, last])

close()

Release every attached renderer at the end of a rendered run.

expected_reward(arm)

Mean reward of an arm, for regret accounting only.

get_wrapper_attr(name)

Gets the attribute name from the environment.

has_wrapper_attr(name)

Checks if the attribute name exists in the environment.

render([mode])

Forward the last (arm, reward) pair to every attached renderer.

reset([seed, options])

Start a new run by reseeding the environment.

set_wrapper_attr(name, value, *[, force])

Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.

step(arm)

Sample one reward from the given arm.

Attributes

means

Mean reward of every arm.

metadata

np_random

Returns the environment's internal _np_random that if not set will initialise with a random seed.

np_random_seed

Returns the environment's internal _np_random_seed that if not set will first initialise with a random int as seed.

number_arms

Number of available arms.

optimal_arm

Index of the best arm.

optimal_mean

Mean reward of the best arm, \(\mu^\star\).

render_mode

spec

unwrapped

Returns the base non-wrapped environment.

action_space

observation_space

close()[source]#

Release every attached renderer at the end of a rendered run.

expected_reward(arm)[source]#

Mean reward of an arm, for regret accounting only.

Parameters:

arm (int) – Index of the arm.

Returns:

That arm’s true mean. Never pass this to a learner.

Return type:

float

property means#

Mean reward of every arm.

property number_arms#

Number of available arms.

property optimal_arm#

Index of the best arm.

Ties are broken by numpy.argmax(), i.e. the lowest index wins.

Type:

int

property optimal_mean#

Mean reward of the best arm, \(\mu^\star\).

Used to define regret.

Type:

float

render(mode='human')[source]#

Forward the last (arm, reward) pair to every attached renderer.

Parameters:

mode (str, default='human') – Unused; accepted for gymnasium.Env compatibility. Output is whatever the objects in renderers produce.

reset(seed=None, options=None)[source]#

Start a new run by reseeding the environment.

Parameters:
  • seed (int, optional) – Seed for the environment’s generator. None draws a fresh one.

  • options (dict, optional) – Unused; accepted for gymnasium.Env compatibility.

Returns:

The constant dummy observation 0 — a bandit is stateless.

Return type:

int

step(arm)[source]#

Sample one reward from the given arm.

Parameters:

arm (int) – Index of the arm to pull, in range(number_arms).

Returns:

An independent draw from that arm’s reward distribution.

Return type:

float