Agent#

class statrl.settings.bandits.adversarial.lipschitz.agent.Agent[source]#

Bases: object

Base class for agents in the adversarial Lipschitz setting.

The protocol differs from the stochastic bandit one in two ways: actions are points of a continuous metric space rather than arm indices, and select_arm() receives an observation, since the adversary’s choice at round t may be partly revealed before the action is committed.

Methods

__init__()

select_arm(observation)

Choose the action \(x_t\) to play this round.

update(action, reward[, observation])

Learn from the reward observed for the action just played.

select_arm(observation)[source]#

Choose the action \(x_t\) to play this round.

Parameters:

observation (object) – Observation for the current round, or None in the pure bandit case where the environment reveals nothing in advance.

Returns:

A point of the action space, typically an ndarray. The interaction loop clips it to the bounds of a Box action space.

Return type:

object

Raises:

NotImplementedError – Always, in the base class; subclasses must override.

update(action, reward, observation=None)[source]#

Learn from the reward observed for the action just played.

Optional: the default does nothing, so a fixed-action baseline needs to implement select_arm() only.

Parameters:
  • action (object) – The action that was played.

  • reward (float) – Value \(f_t(x_t)\) returned by the adversary. Only this scalar is observed — never the whole function \(f_t\).

  • observation (object, optional) – Observation the action was chosen from.