MDPAgent#

class statrl.settings.markovdecisionprocess.discrete_nostructure.agent.MDPAgent(nS, nA, name='Agent', seed=None)[source]#

Bases: object

Base class for agents in the discrete MDP setting.

The protocol differs from the bandit one in that both play() and update() are state-aware: an action is chosen for a state, and learning must account for where the process went next.

Parameters:
  • nS (int) – Number of states.

  • nA (int) – Number of actions.

  • name (str, default='Agent') – Label used in logfiles and plot legends.

  • seed (int, optional) – Seed for the agent’s own randomness.

np_random#

Agent-local generator, available after the first reset().

Type:

numpy.random.Generator

Methods

__init__(nS, nA[, name, seed])

play(state)

Choose an action for the given state.

reset(inistate)

Start a new independent run from a given initial state.

update(state, action, reward, observation)

Learn from one transition.

play(state)[source]#

Choose an action for the given state.

Parameters:

state (int) – Current state.

Returns:

The chosen action. The default picks uniformly at random.

Return type:

int

reset(inistate)[source]#

Start a new independent run from a given initial state.

Parameters:

inistate (int) – State the environment was reset to. Agents that track a current state need it; the default only reseeds.

update(state, action, reward, observation)[source]#

Learn from one transition.

Parameters:
  • state (int) – State the action was taken in.

  • action (int) – Action taken.

  • reward (float) – Reward observed.

  • observation (int) – State reached. Unlike a bandit, this is part of the feedback — learning the transition kernel is half the problem.