MDPInteraction#

class statrl.settings.markovdecisionprocess.discrete_nostructure.interaction.MDPInteraction[source]#

Bases: Interaction

Interaction loop for the discrete MDP setting.

Drives play(state) -> step(action) -> update(state, action, reward, next_state) for a fixed number of rounds.

On a terminal transition the environment is reset and the run continues, so horizon counts steps, not episodes.

Methods

__init__()

renderrun(env, learner, horizon)

Run one interaction, printing each step to stdout.

run(env, learner, horizon)

Run one interaction and return its cumulative expected score.

Attributes

plotlabels

Axis labels (x, y) for the regret plots.

property plotlabels#

Axis labels (x, y) for the regret plots.

Type:

tuple of (str, str)

renderrun(env, learner, horizon)[source]#

Run one interaction, printing each step to stdout.

Parameters:
  • env (DiscreteMDP) – The MDP instance. Its renderers list is overwritten with a text renderer.

  • learner (MDPAgent) – The agent.

  • horizon (int) – Number of steps to play.

run(env, learner, horizon)[source]#

Run one interaction and return its cumulative expected score.

Parameters:
  • env (DiscreteMDP) – The MDP instance. Reset at the start, and again after any terminal transition.

  • learner (MDPAgent) – The agent, reset with the initial state.

  • horizon (int) – Number of steps to play.

Returns:

Cumulative sum of the expected rewards of the visited state-action pairs.

Return type:

ndarray of shape (horizon,)