Adversarial Lipschitz online optimization#
Module: statrl.settings.bandits.adversarial.lipschitz
This setting leaves the finite-arm world behind. The action space is a continuous metric space, and at each round an adversary chooses a reward function \(f_t(x)\) (assumed Lipschitz). The learner picks an action \(x_t\), then observes the reward \(f_t(x_t)\). Because the reward functions are adversarial rather than drawn from a fixed distribution, the algorithms differ from the stochastic settings.
The environment#
LipschitzAdversarialEnv
is a gymnasium-like environment constructed from:
action_space— the continuous domain (e.g. aBox),reward_function_sequence— a callable mapping a roundtto its reward functionf_t(x),horizon— the number of roundsT, andobservation_fn— optional; produces the observation at each round (trivial in the pure bandit case).
step(action) returns the gymnasium 5-tuple (observation, reward, terminated,
truncated, info).
The agent#
The base Agent defines
select_arm(observation) (return an action) and an optional update(action,
reward, observation=None).
ALF — Adversarial Lipschitz Forecaster#
ALFLearner reduces the
continuous problem to a finite one: it builds an ε-net (a discrete cover) of the action
space, then runs an exponential-weights / Hedge update over those points
[Maillard2010]. Constructor parameters:
epsilon— discretization resolution of the cover,eta— learning rate for the exponential weights,horizon— used to tuneepsilon/eta,metric— optional metric for building the cover, andsampling—"argmax"(play the best expert) or"sample"(draw from the weight distribution).
The interaction loop#
run_lipschitz_online_learning()
drives play → step → update over the horizon and returns the reward series as
an array.