RandomBernoulliBandit#

statrl.settings.bandits.stochastic.anytime.envs.parametric.RandomBernoulliBandit(Delta, K, name='MAB-RandomBernoulli')[source]#

Draw a random Bernoulli instance with a prescribed optimality gap.

Useful for studying how regret scales with the gap: the difficulty of a bandit instance is governed by \(\Delta\), so sweeping it while holding K fixed isolates that dependence.

Parameters:
  • Delta (float) – Gap between the best and second-best arm. The remaining arms get means drawn uniformly below the second-best one.

  • K (int) – Number of arms.

  • name (str, default='MAB-RandomBernoulli') – Prefix of the instance name.

Returns:

A Bernoulli bandit whose two leading means differ by exactly Delta.

Return type:

StochasticBanditEnv