Opti_swimmer#

class statrl.settings.markovdecisionprocess.discrete_nostructure.agents._Oracle.Opti_swimmer(env)[source]#

Bases: MDPAgent

Hand-coded oracle for RiverSwim: always swim right.

Skips value iteration by encoding the known optimal policy directly.

Parameters:

env (DiscreteMDP) – The RiverSwim instance.

Methods

__init__(env)

play(state)

Always take action 0, "swim right".

reset(inistate)

Start a new run.

update(state, action, reward, observation)

Ignore the transition (the oracle has nothing to learn).

play(state)[source]#

Always take action 0, “swim right”.

Parameters:

state (int) – Current state; ignored.

Returns:

Always 0.

Return type:

int

reset(inistate)[source]#

Start a new run. The policy is fixed, so nothing is cleared.

Parameters:

inistate (int) – Initial state; ignored.

update(state, action, reward, observation)[source]#

Ignore the transition (the oracle has nothing to learn).

Parameters:
  • state (int) – State the action was taken in.

  • action (int) – Action taken.

  • reward (float) – Reward observed.

  • observation (int) – State reached.