Random#

class statrl.settings.markovdecisionprocess.discrete_nostructure.agents._Random.Random(env)[source]#

Bases: MDPAgent

Uniform exploration: take a random action in every state.

Parameters:

env (DiscreteMDP) – The environment, whose action_space supplies the draws.

Methods

__init__(env)

play(state)

Draw an action uniformly at random, ignoring the state.

reset(inistate)

Start a new run.

update(state, action, reward, observation)

Ignore the transition — this agent does not learn.

play(state)[source]#

Draw an action uniformly at random, ignoring the state.

Parameters:

state (int) – Current state; ignored.

Returns:

An action drawn from the environment’s action space.

Return type:

int

reset(inistate)[source]#

Start a new run. No statistics are kept, so this does nothing.

Parameters:

inistate (int) – Initial state; ignored.

update(state, action, reward, observation)[source]#

Ignore the transition — this agent does not learn.

Parameters:
  • state (int) – State the action was taken in.

  • action (int) – Action taken.

  • reward (float) – Reward observed.

  • observation (int) – State reached.