BCBnaif#
- class statrl.settings.bandits.stochastic.batch.agents.BCB.BCBnaif(nbArms, bound=1.0)[source]#
Bases:
BatchBanditAgentBCB without the optimistic within-batch count increment.
Differs from
BCBin one respect: the pull counts are left untouched duringbatchplay()and updated only at the end of the batch. Equivalent to drawingbatchsizei.i.d. actions from the current policy and updating afterwards. Kept as the reference point that isolates what the optimistic increment buys.- Parameters:
See also
BCBThe adapted version, with the within-batch increment.
Methods
__init__(nbArms[, bound])batchplay(batchsize)Fill the whole batch with a single posterior draw's winner.
batchupdate(batcharm, batchreward)Fold the batch's rewards into the counts, histories, and means.
play()Draw one posterior mean per arm and play the best.
reset()Clear every statistic and restore the anchored Dirichlet prior.
update(arm, reward)Record one
(arm, reward)pair in the counts and the history.- batchplay(batchsize)[source]#
Fill the whole batch with a single posterior draw’s winner.
- Parameters:
batchsize (int) – Number of pulls in this batch.
- Returns:
batchsizecopies of one arm. UnlikeBCB.batchplay(), no count is incremented here.- Return type:
- batchupdate(batcharm, batchreward)[source]#
Fold the batch’s rewards into the counts, histories, and means.