AnytimeToKnownHorizonAgentWrapper#

class statrl.settings.bandits.stochastic.knownhorizon.wrappers.wrapper_anytime_knownhorizon.AnytimeToKnownHorizonAgentWrapper(learner)[source]#

Bases: BanditAgent

Run an anytime agent inside the known-horizon setting.

An anytime agent simply ignores the horizon, so the wrapper drops it.

Parameters:

learner (statrl.settings.bandits.stochastic.anytime.agent.BanditAgent) – The anytime agent to adapt.

See also

KnownHorizonToAnytimeAgentWrapper

The converse adaptation.

Examples

>>> from statrl.settings.bandits.stochastic.anytime.agents.IMED import IMED
>>> AnytimeToKnownHorizonAgentWrapper(IMED(3)).name
'IMED-anytime'

Methods

__init__(learner)

reset(horizon)

Reset the wrapped agent, discarding the horizon.

select_arm()

Delegate the choice to the wrapped agent.

update(arm, reward)

Forward the observation to the wrapped agent.

Attributes

policy

Policy of the wrapped agent.

property policy#

Policy of the wrapped agent.

Raises:

AttributeError – If the wrapped agent has no policy. Only oracles define one; the experiment harness reads it to log the optimal policy.

reset(horizon)[source]#

Reset the wrapped agent, discarding the horizon.

Parameters:

horizon (int) – Ignored: an anytime agent must not depend on it.

select_arm()[source]#

Delegate the choice to the wrapped agent.

Returns:

Index of the selected arm.

Return type:

int

update(arm, reward)[source]#

Forward the observation to the wrapped agent.

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Reward observed for that arm.