LipschitzAdversarialEnv#

class statrl.settings.bandits.adversarial.lipschitz.environment.LipschitzAdversarialEnv(action_space, reward_function_sequence, horizon, observation_fn=None)[source]#

Bases: object

Gymnasium-like environment for adversarial Lipschitz online optimization.

The environment exposes a sequence of reward functions f_t(x), chosen by an adversary (or precomputed generator).

Each step:

action = learner.select_arm(observation) reward = f_t(action)

Methods

__init__(action_space, ...[, observation_fn])

reset([seed])

Rewind to round 0 and draw the first reward function.

step(action)

Executes one round of interaction.

reset(seed=None)[source]#

Rewind to round 0 and draw the first reward function.

Parameters:

seed (int, optional) – If given, seeds the global numpy.random state, which also affects any other code drawing from it. Pass None to leave it untouched.

Returns:

  • obs (object) – Observation at round 0, or None in the pure bandit case.

  • info (dict) – Empty, for gymnasium.Env compatibility.

step(action)[source]#

Executes one round of interaction.

Parameters:

action – x_t chosen by learner

Returns:

  • obs (next observation)

  • reward (float)

  • terminated (bool)

  • truncated (bool)

  • info (dict)