LipschitzAdversarialEnv#
- class statrl.settings.bandits.adversarial.lipschitz.environment.LipschitzAdversarialEnv(action_space, reward_function_sequence, horizon, observation_fn=None)[source]#
Bases:
objectGymnasium-like environment for adversarial Lipschitz online optimization.
The environment exposes a sequence of reward functions f_t(x), chosen by an adversary (or precomputed generator).
- Each step:
action = learner.select_arm(observation) reward = f_t(action)
Methods
__init__(action_space, ...[, observation_fn])reset([seed])Rewind to round 0 and draw the first reward function.
step(action)Executes one round of interaction.
- reset(seed=None)[source]#
Rewind to round 0 and draw the first reward function.
- Parameters:
seed (int, optional) – If given, seeds the global
numpy.randomstate, which also affects any other code drawing from it. PassNoneto leave it untouched.- Returns:
obs (object) – Observation at round 0, or
Nonein the pure bandit case.info (dict) – Empty, for
gymnasium.Envcompatibility.