KLinf_threshold#

statrl.settings.utils.KLinf_threshold(reward_history, mean_threshold, upper_bound=1.0, custom_optim=True)[source]#

Non-parametric divergence \(K_{\inf}\) between an empirical distribution and the set of distributions with mean above a threshold, evaluated using the concave dual.

Parameters:
  • reward_history (array-like of float) – Rewards observed so far for one arm; they define the empirical measure \(\hat{F}\).

  • mean_threshold (float) – Threshold mean \(\mu^*\), in practice the best empirical mean across arms.

  • upper_bound (float, default=1.0) – Known upper bound B of the reward support.

  • custom_optim (bool, default=True) – If True, first test whether the dual maximum is attained at a boundary (the common case, roughly a 2x speed-up) and otherwise solve jac = 0 with Brent’s method. On non-convergence, and when False, fall back to scipy.optimize.minimize_scalar().

Returns:

The value of \(K_{\inf}\), or inf if the optimizer failed — which makes the arm ineligible for this round.

Return type:

float

See also

klBern, klGauss, used

Examples

>>> bool(KLinf_threshold([0.5, 0.5, 0.5], 0.5) < 1e-6)
True