BCBnaif#

class statrl.settings.bandits.stochastic.batch.agents.BCB.BCBnaif(nbArms, bound=1.0)[source]#

Bases: BatchBanditAgent

BCB without the optimistic within-batch count increment.

Differs from BCB in one respect: the pull counts are left untouched during batchplay() and updated only at the end of the batch. Equivalent to drawing batchsize i.i.d. actions from the current policy and updating afterwards. Kept as the reference point that isolates what the optimistic increment buys.

Parameters:
  • nbArms (int) – Number of arms.

  • bound (float, default=1.0) – Upper bound B of the reward support; the Dirichlet prior is anchored on a single pseudo-observation there.

See also

BCB

The adapted version, with the within-batch increment.

Methods

__init__(nbArms[, bound])

batchplay(batchsize)

Fill the whole batch with a single posterior draw's winner.

batchupdate(batcharm, batchreward)

Fold the batch's rewards into the counts, histories, and means.

play()

Draw one posterior mean per arm and play the best.

reset()

Clear every statistic and restore the anchored Dirichlet prior.

update(arm, reward)

Record one (arm, reward) pair in the counts and the history.

batchplay(batchsize)[source]#

Fill the whole batch with a single posterior draw’s winner.

Parameters:

batchsize (int) – Number of pulls in this batch.

Returns:

batchsize copies of one arm. Unlike BCB.batchplay(), no count is incremented here.

Return type:

list of int

batchupdate(batcharm, batchreward)[source]#

Fold the batch’s rewards into the counts, histories, and means.

Parameters:
  • batcharm (list of int) – The arms that were pulled.

  • batchreward (list of float) – The rewards observed for them.

play()[source]#

Draw one posterior mean per arm and play the best.

Returns:

Arm with the highest Dirichlet-reweighted mean this draw.

Return type:

int

reset()[source]#

Clear every statistic and restore the anchored Dirichlet prior.

update(arm, reward)[source]#

Record one (arm, reward) pair in the counts and the history.

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Reward observed for it.