KnownHorizonToAnytimeAgentWrapper#

class statrl.settings.bandits.stochastic.knownhorizon.wrappers.wrapper_anytime_knownhorizon.KnownHorizonToAnytimeAgentWrapper(learner, horizon, name)[source]#

Bases: BanditAgent

Run a horizon-aware agent in the anytime setting, at a fixed horizon.

An anytime agent is never told the horizon, so the wrapper must supply one up front and reuse it for every run. The wrapped agent then plans against horizon regardless of how long the interaction actually lasts.

Parameters:
  • learner (statrl.settings.bandits.stochastic.knownhorizon.agent.BanditAgent) – The horizon-aware agent to adapt.

  • horizon (int) – Horizon announced to the wrapped agent on every reset. Set it to the horizon of the experiment you intend to run.

  • name (str) – Label for the wrapped agent. Unlike the converse wrapper, no suffix is added, so pick a name that records the fixed horizon.

See also

AnytimeToKnownHorizonAgentWrapper

The converse, lossless adaptation.

Methods

__init__(learner, horizon, name)

reset()

Reset the wrapped agent, announcing the fixed horizon.

select_arm()

Delegate the choice to the wrapped agent.

update(arm, reward)

Forward the observation to the wrapped agent.

Attributes

policy

Policy of the wrapped agent.

property policy#

Policy of the wrapped agent.

Raises:

AttributeError – If the wrapped agent has no policy; only oracles define one.

reset()[source]#

Reset the wrapped agent, announcing the fixed horizon.

select_arm()[source]#

Delegate the choice to the wrapped agent.

Returns:

Index of the selected arm.

Return type:

int

update(arm, reward)[source]#

Forward the observation to the wrapped agent.

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Reward observed for that arm.