BanditAgent#
- class statrl.settings.bandits.stochastic.anytime.agent.BanditAgent(name, seed=1)[source]#
Bases:
ABCBase class for anytime stochastic bandit agents.
Subclasses implement the protocol shared by every agent in this setting:
reset()to start an independent run,select_arm()to choose an arm, andupdate()to learn from the observed reward.- Parameters:
See also
statrl.settings.bandits.stochastic.knownhorizon.agent.BanditAgentThe counterpart whose
resetreceives the horizon.
Notes
Agents are deep-copied once per replicate by
multicoreRuns(), so an agent may hold arbitrary state as long asreset()fully reinitializes it.Methods
__init__(name[, seed])reset()Start a new independent run.
Choose the arm to pull next.
update(arm, reward)Learn from the reward observed for the arm just pulled.
- reset()[source]#
Start a new independent run.
Reseeds
np_randomand clears any statistics accumulated by a previous run.