BIMED#

class statrl.settings.bandits.stochastic.batch.agents.BIMED.BIMED(nbArms, bound=1.0, **kwargs)[source]#

Bases: BatchBanditAgent

Batched IMED with within-batch sequential index updates.

Uses the non-parametric index [1] \(I_a = N_a K_{\inf}(\hat{F}_a, \hat{\mu}^\star) + \log N_a\), where \(K_{\inf}\) is computed by KLinf_threshold(). Being non-parametric, it needs no assumption on the reward family beyond a known upper bound.

Within a batch only \(N_a\) moves — the reward histories, and so \(\hat{F}_a\), stay frozen until the batch ends.

Parameters:
  • nbArms (int) – Number of arms.

  • bound (float, default=1.0) – Known upper bound of the reward support, passed to KLinf_threshold(). An understated bound invalidates the divergence.

  • **kwargs (dict) – Ignored; accepted so the batched agents share a constructor signature.

nbDraws#

Running pull count of each arm, updated within a batch.

Type:

ndarray of shape (nbArms,)

nbDraws_start#

Pull count frozen at the last batch boundary.

Type:

ndarray of shape (nbArms,)

rewardHistory#

Rewards observed per arm; the empirical measure \(\hat{F}_a\).

Type:

list of list of float

indexes#

Current index of each arm.

Type:

ndarray of shape (nbArms,)

x_threshold#

Within-batch inflation factor, log(t + B) / log(t), recomputed at the start of each batch. It grows with the batch size, letting a larger batch spread further from the frozen statistics.

Type:

float

See also

statrl.settings.bandits.stochastic.anytime.agents.IMED.IMED

The unbatched, parametric original.

References

Methods

__init__(nbArms[, bound])

batchplay(batchsize)

Commit a batch, refreshing only the pulled arm's index at each step.

batchupdate(batcharm, batchreward)

Fold in the batch's rewards and re-freeze the pull counts.

play()

Pick the minimal-index arm, unless it has outrun its batch budget.

reset()

Clear every statistic before a new independent run.

update(arm, reward)

Record one (arm, reward) pair and recompute every index.

batchplay(batchsize)[source]#

Commit a batch, refreshing only the pulled arm’s index at each step.

The reward histories (and each arm’s empirical distribution \(\hat{F}_a\)) do not change during a batch, so a pull only moves \(N_a\) for the arm chosen. Recomputing that one index instead of all of them brings the cost of a batch down from \(O(B \cdot K \cdot K_{\inf})\) to \(O(B \cdot K_{\inf})\).

Parameters:

batchsize (int) – Number of pulls in this batch.

Returns:

Exactly batchsize arm indices.

Return type:

list of int

batchupdate(batcharm, batchreward)[source]#

Fold in the batch’s rewards and re-freeze the pull counts.

Extends each arm’s reward history, refreshes its empirical mean, and snapshots nbDraws into nbDraws_start to open the next batch.

Parameters:
  • batcharm (list of int) – The arms that were pulled.

  • batchreward (list of float) – The rewards observed for them.

Notes

nbDraws is not incremented here: batchplay() already did so as it committed each pull.

play()[source]#

Pick the minimal-index arm, unless it has outrun its batch budget.

Returns:

The minimal-index arm when it is still under the within-batch budget x_threshold * nbDraws_start, or is itself the empirical leader, or has never been pulled. Otherwise the least-saturated arm by that same ratio, and failing that the empirical leader. The fallbacks stop a single batch from pouring all its pulls into one arm on the strength of statistics that cannot update until the batch ends.

Return type:

int

reset()[source]#

Clear every statistic before a new independent run.

Two pull counters are kept: nbDraws is the running total, while nbDraws_start freezes it at the last batch boundary. The index uses both, which is what lets an arm’s index move within a batch while the empirical distribution it is based on stays fixed.

update(arm, reward)[source]#

Record one (arm, reward) pair and recompute every index.

The single-pull path, used outside batched interaction. Within a batch, batchupdate() is called instead.

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Reward observed for it.