Known horizon#
Module: statrl.settings.bandits.stochastic.knownhorizon
This setting is the stochastic bandit problem again, but with the time horizon known in
advance. That extra information lets algorithms budget exploration against a fixed number
of rounds, so the agent’s reset receives the horizon:
class BanditAgent(ABC):
def reset(self, timehorizon): ...
def select_arm(self) -> int: ...
def update(self, arm, reward): ...
The environment#
StochasticBanditEnv
is the anytime environment, re-exported unchanged: knowing the horizon changes what
the agent may do, not what the problem is. Any environment built for the anytime
setting therefore works here as is.
Bridging to the anytime setting#
Because the two stochastic settings share the same environment/agent shape, the
wrappers package adapts an object built for one setting to the other. Four adapters are
provided in
wrapper_anytime_knownhorizon:
Wrapper |
Adapts … |
|---|---|
|
an anytime environment for use in the known-horizon setting. |
|
a known-horizon environment for use in the anytime setting. |
|
an anytime agent so it accepts a horizon at |
|
a known-horizon agent (given a fixed horizon) for anytime use. |
This lets you benchmark an algorithm designed for one setting against the environments and loops of the other without reimplementing it.
The interaction loop#
BanditInteraction runs the same
select_arm → pull → update cycle for horizon rounds and returns the
cumulative-score array.