BCB#

class statrl.settings.bandits.stochastic.batch.agents.BCB.BCB(nbArms, bound=1.0)[source]#

Bases: BatchBanditAgent

BCB with CVaR = Expectation (adapted batch version).

In each batch, the arm counts are updated sequentially inside the batch (optimistic within-batch exploration), but reward histories (used for the Dirichlet draw) are only updated at the end of the batch via batchupdate.

Parameters:
  • nbArms (int)

  • bound (float) – Upper bound B of the reward support. The Dirichlet prior is initialised with a single pseudo-observation at B.

Methods

__init__(nbArms[, bound])

batchplay(batchsize)

Fill the whole batch with a single posterior draw's winner.

batchupdate(batcharm, batchreward)

Append the batch's rewards to the arm histories and refresh the means.

play()

Draw one posterior mean per arm and play the best.

reset()

Clear every statistic and restore the prior.

update(arm, reward)

Record one (arm, reward) pair in the counts and the history.

batchplay(batchsize)[source]#

Fill the whole batch with a single posterior draw’s winner.

Parameters:

batchsize (int) – Number of pulls in this batch.

Returns:

batchsize copies of one arm.

Return type:

list of int

Notes

The scores depend only on the reward histories, which do not change during a batch, so every draw within the batch would select the same arm. The winner is computed once instead of batchsize times. Its count is incremented optimistically up front, keeping nbDraws consistent with what batchupdate() assumes.

batchupdate(batcharm, batchreward)[source]#

Append the batch’s rewards to the arm histories and refresh the means.

Parameters:
  • batcharm (list of int) – The arms that were pulled.

  • batchreward (list of float) – The rewards observed for them.

Notes

nbDraws is not incremented here: batchplay() already did so optimistically when it committed the batch.

play()[source]#

Draw one posterior mean per arm and play the best.

Returns:

Arm with the highest Dirichlet-reweighted mean this draw. Ties are broken uniformly at random.

Return type:

int

reset()[source]#

Clear every statistic and restore the prior.

Each arm’s reward history restarts with a single pseudo-observation at the upper bound B. That optimistic anchor is what drives exploration: an arm with few observations still has appreciable posterior mass near B.

update(arm, reward)[source]#

Record one (arm, reward) pair in the counts and the history.

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Reward observed for it; appended to that arm’s history, which is the empirical measure the Dirichlet draw reweights.