Anytime#
Module: statrl.settings.bandits.stochastic.anytime
In the stochastic setting each arm has a fixed, unknown reward distribution, and pulling an arm draws an independent sample from it. The anytime variant makes no assumption about the horizon: the agent should perform well whenever it is stopped, so it cannot tune its behaviour to a known number of rounds.
The environment#
StochasticBanditEnv is an
abstract, stateless multi-armed bandit environment — interactions never modify the
underlying distributions. A subclass supplies the arms’ reward distributions and
implements:
number_arms— number of arms,means— the true mean of each arm (used for evaluation and to build an oracle), andstep(arm)— draw one reward from the given arm.
Unlike gymnasium.Env.step() this returns the reward alone rather than a
five-tuple: a bandit has no observable state.
It also derives optimal_mean and optimal_arm from means, and supports seeding
and rendering via gymnasium utilities.
The agent#
BanditAgent is the abstract base
class. Every agent carries a name and implements reset(), select_arm(), and
update(arm, reward).
Shipped agents#
The interaction loop#
BanditInteraction
resets the environment and learner, then repeats select_arm → step →
update for horizon rounds, accumulating the arms’ expected rewards into a
running score it returns as an array. Call it as BanditInteraction().run(env,
learner, horizon); renderrun runs the same loop while printing each pull.