Agent#
- class statrl.settings.bandits.adversarial.lipschitz.agent.Agent[source]#
Bases:
objectBase class for agents in the adversarial Lipschitz setting.
The protocol differs from the stochastic bandit one in two ways: actions are points of a continuous metric space rather than arm indices, and
select_arm()receives an observation, since the adversary’s choice at roundtmay be partly revealed before the action is committed.Methods
__init__()select_arm(observation)Choose the action \(x_t\) to play this round.
update(action, reward[, observation])Learn from the reward observed for the action just played.
- select_arm(observation)[source]#
Choose the action \(x_t\) to play this round.
- Parameters:
observation (object) – Observation for the current round, or
Nonein the pure bandit case where the environment reveals nothing in advance.- Returns:
A point of the action space, typically an
ndarray. The interaction loop clips it to the bounds of aBoxaction space.- Return type:
- Raises:
NotImplementedError – Always, in the base class; subclasses must override.
- update(action, reward, observation=None)[source]#
Learn from the reward observed for the action just played.
Optional: the default does nothing, so a fixed-action baseline needs to implement
select_arm()only.