BCB#
- class statrl.settings.bandits.stochastic.batch.agents.BCB.BCB(nbArms, bound=1.0)[source]#
Bases:
BatchBanditAgentBCB with CVaR = Expectation (adapted batch version).
In each batch, the arm counts are updated sequentially inside the batch (optimistic within-batch exploration), but reward histories (used for the Dirichlet draw) are only updated at the end of the batch via batchupdate.
- Parameters:
Methods
__init__(nbArms[, bound])batchplay(batchsize)Fill the whole batch with a single posterior draw's winner.
batchupdate(batcharm, batchreward)Append the batch's rewards to the arm histories and refresh the means.
play()Draw one posterior mean per arm and play the best.
reset()Clear every statistic and restore the prior.
update(arm, reward)Record one
(arm, reward)pair in the counts and the history.- batchplay(batchsize)[source]#
Fill the whole batch with a single posterior draw’s winner.
- Parameters:
batchsize (int) – Number of pulls in this batch.
- Returns:
batchsizecopies of one arm.- Return type:
Notes
The scores depend only on the reward histories, which do not change during a batch, so every draw within the batch would select the same arm. The winner is computed once instead of
batchsizetimes. Its count is incremented optimistically up front, keepingnbDrawsconsistent with whatbatchupdate()assumes.
- batchupdate(batcharm, batchreward)[source]#
Append the batch’s rewards to the arm histories and refresh the means.
- Parameters:
Notes
nbDrawsis not incremented here:batchplay()already did so optimistically when it committed the batch.
- play()[source]#
Draw one posterior mean per arm and play the best.
- Returns:
Arm with the highest Dirichlet-reweighted mean this draw. Ties are broken uniformly at random.
- Return type: