BatchBanditAgent#
- class statrl.settings.bandits.stochastic.batch.agent.BatchBanditAgent(name='BanditAgent', seed=1)[source]#
Bases:
ABCBase class for batched bandit agents.
In the batched setting an agent must commit to a whole block of pulls before seeing any of their rewards.
Two levels of interface are provided.
play()andupdate()are the per-pull rules;batchplay()andbatchupdate()are what the interaction loop actually calls. Subclasses must implement the batch pair, which lets them exploit within-batch structure: an agent may update its index between the pulls of a batch (using only what it knew when the batch began) even though no reward has yet arrived.- Parameters:
See also
statrl.settings.bandits.stochastic.batch.interaction.BatchBanditInteractionThe loop that drives these agents.
Methods
__init__([name, seed])batchplay(batchsize)Commit to the arms of a whole batch, before any reward is seen.
batchupdate(batcharm, batchreward)Learn from all the rewards of a batch at once.
play()Choose a single arm.
reset()Start a new independent run, reseeding the agent's generator.
update(arm, reward)Learn from one
(arm, reward)pair.- abstractmethod batchplay(batchsize)[source]#
Commit to the arms of a whole batch, before any reward is seen.
- Parameters:
batchsize (int) – Number of pulls in this batch, announced by the environment as
info["nextbatchsize"]. It varies between batches under a non-constant schedule.- Returns:
Exactly
batchsizearm indices. The environment asserts the length. The default implementation repeatsplay().- Return type:
- abstractmethod batchupdate(batcharm, batchreward)[source]#
Learn from all the rewards of a batch at once.
- Parameters:
batcharm (list of int) – The arms that were pulled, as returned by
batchplay().batchreward (list of float) – The rewards observed for them, in the same order.
Notes
The default implementation replays the pairs through
update().
- play()[source]#
Choose a single arm.
- Returns:
Index of the selected arm.
- Return type:
- Raises:
NotImplementedError – If not overridden. Agents whose
batchplay()builds a batch from repeated single pulls must implement this; agents that decide a batch as a whole, such asBABA, need not.