BIMED#
- class statrl.settings.bandits.stochastic.batch.agents.BIMED.BIMED(nbArms, bound=1.0, **kwargs)[source]#
Bases:
BatchBanditAgentBatched IMED with within-batch sequential index updates.
Uses the non-parametric index [1] \(I_a = N_a K_{\inf}(\hat{F}_a, \hat{\mu}^\star) + \log N_a\), where \(K_{\inf}\) is computed by
KLinf_threshold(). Being non-parametric, it needs no assumption on the reward family beyond a known upper bound.Within a batch only \(N_a\) moves — the reward histories, and so \(\hat{F}_a\), stay frozen until the batch ends.
- Parameters:
nbArms (int) – Number of arms.
bound (float, default=1.0) – Known upper bound of the reward support, passed to
KLinf_threshold(). An understated bound invalidates the divergence.**kwargs (dict) – Ignored; accepted so the batched agents share a constructor signature.
- nbDraws#
Running pull count of each arm, updated within a batch.
- Type:
ndarray of shape (nbArms,)
- nbDraws_start#
Pull count frozen at the last batch boundary.
- Type:
ndarray of shape (nbArms,)
- rewardHistory#
Rewards observed per arm; the empirical measure \(\hat{F}_a\).
- indexes#
Current index of each arm.
- Type:
ndarray of shape (nbArms,)
- x_threshold#
Within-batch inflation factor,
log(t + B) / log(t), recomputed at the start of each batch. It grows with the batch size, letting a larger batch spread further from the frozen statistics.- Type:
See also
statrl.settings.bandits.stochastic.anytime.agents.IMED.IMEDThe unbatched, parametric original.
References
Methods
__init__(nbArms[, bound])batchplay(batchsize)Commit a batch, refreshing only the pulled arm's index at each step.
batchupdate(batcharm, batchreward)Fold in the batch's rewards and re-freeze the pull counts.
play()Pick the minimal-index arm, unless it has outrun its batch budget.
reset()Clear every statistic before a new independent run.
update(arm, reward)Record one
(arm, reward)pair and recompute every index.- batchplay(batchsize)[source]#
Commit a batch, refreshing only the pulled arm’s index at each step.
The reward histories (and each arm’s empirical distribution \(\hat{F}_a\)) do not change during a batch, so a pull only moves \(N_a\) for the arm chosen. Recomputing that one index instead of all of them brings the cost of a batch down from \(O(B \cdot K \cdot K_{\inf})\) to \(O(B \cdot K_{\inf})\).
- batchupdate(batcharm, batchreward)[source]#
Fold in the batch’s rewards and re-freeze the pull counts.
Extends each arm’s reward history, refreshes its empirical mean, and snapshots
nbDrawsintonbDraws_startto open the next batch.- Parameters:
Notes
nbDrawsis not incremented here:batchplay()already did so as it committed each pull.
- play()[source]#
Pick the minimal-index arm, unless it has outrun its batch budget.
- Returns:
The minimal-index arm when it is still under the within-batch budget
x_threshold * nbDraws_start, or is itself the empirical leader, or has never been pulled. Otherwise the least-saturated arm by that same ratio, and failing that the empirical leader. The fallbacks stop a single batch from pouring all its pulls into one arm on the strength of statistics that cannot update until the batch ends.- Return type:
- reset()[source]#
Clear every statistic before a new independent run.
Two pull counters are kept:
nbDrawsis the running total, whilenbDraws_startfreezes it at the last batch boundary. The index uses both, which is what lets an arm’s index move within a batch while the empirical distribution it is based on stays fixed.
- update(arm, reward)[source]#
Record one
(arm, reward)pair and recompute every index.The single-pull path, used outside batched interaction. Within a batch,
batchupdate()is called instead.