Human#

class statrl.settings.markovdecisionprocess.discrete_nostructure.agents.Human.Human(env)[source]#

Bases: MDPAgent

Interactive agent that asks the user for each action.

Prompts on stdin at every step and blocks until a valid action name is entered.

Parameters:

env (object) – A wrapped environment; env.env must expose nameActions, so a raw DiscreteMDP will not do.

Methods

__init__(env)

play(state)

Print the state and block until the user names an action.

reset(inistate)

Start a new run.

update(state, action, reward, observation)

Ignore the transition, the user is the policy.

play(state)[source]#

Print the state and block until the user names an action.

Parameters:

state (int) – Current state, shown to the user.

Returns:

Index of the action the user chose.

Return type:

int

reset(inistate)[source]#

Start a new run. Nothing is kept between runs.

Parameters:

inistate (int) – Initial state; ignored.

update(state, action, reward, observation)[source]#

Ignore the transition, the user is the policy.

Parameters:
  • state (int) – State the action was taken in.

  • action (int) – Action taken.

  • reward (float) – Reward observed.

  • observation (int) – State reached.