ALFLearner#

class statrl.settings.bandits.adversarial.lipschitz.agents.ALF.ALFLearner(action_space, epsilon, eta, horizon, metric=None, sampling='argmax')[source]#

Bases: Agent

Adversarial Lipschitz Forecaster

Learns in a continuous Lipschitz bandit/online optimization setting by reducing to a finite expert set, following Maillard and Munos [1].

Parameters:
  • action_space (object) – Continuous domain. Only Box-like spaces, exposing low and high, are supported.

  • epsilon (float) – Cover resolution. The grid uses max(2, int(1 / epsilon)) points per dimension, so the expert count grows exponentially with the dimension of the space.

  • eta (float) – Learning rate of the exponential weights.

  • horizon (int) – Time horizon \(T\). Stored for tuning epsilon and eta; this implementation does not tune them for you.

  • metric (object, optional) – Metric for building the cover. Unused by the current Box grid.

  • sampling ({'argmax', 'sample'}, default='argmax') – Whether to play the highest-weight cover point deterministically or draw one from the weight distribution.

actions#

The cover points, i.e. the expert set.

Type:

ndarray of shape (n_actions, n_dims)

n_actions#

Number of cover points.

Type:

int

weights#

Current expert weights, normalized to sum to one.

Type:

ndarray of shape (n_actions,)

time#

Number of updates performed so far.

Type:

int

Raises:

NotImplementedError – If action_space exposes no low / high; a general metric cover is not implemented.

References

Methods

__init__(action_space, epsilon, eta, horizon)

select_arm([observation])

Pick a cover point according to the current expert weights.

update(action, reward[, observation])

Apply the Hedge update to the expert that was played.

select_arm(observation=None)[source]#

Pick a cover point according to the current expert weights.

Parameters:

observation (object, optional) – Ignored; ALF plays from its weights alone.

Returns:

The chosen cover point. Its index is remembered so that the next update() credits the right expert.

Return type:

ndarray

update(action, reward, observation=None)[source]#

Apply the Hedge update to the expert that was played.

Multiplies that expert’s weight by \(e^{\eta r}\) and renormalizes.

Parameters:
  • action (ndarray) – The action played. Ignored — the expert index recorded by select_arm() is used instead.

  • reward (float) – Observed reward \(f_t(x_t)\).

  • observation (object, optional) – Ignored.