MDPInteraction#
- class statrl.settings.markovdecisionprocess.discrete_nostructure.interaction.MDPInteraction[source]#
Bases:
InteractionInteraction loop for the discrete MDP setting.
Drives
play(state)->step(action)->update(state, action, reward, next_state)for a fixed number of rounds.On a terminal transition the environment is reset and the run continues, so
horizoncounts steps, not episodes.Methods
__init__()renderrun(env, learner, horizon)Run one interaction, printing each step to stdout.
run(env, learner, horizon)Run one interaction and return its cumulative expected score.
Attributes
Axis labels
(x, y)for the regret plots.- renderrun(env, learner, horizon)[source]#
Run one interaction, printing each step to stdout.
- Parameters:
env (DiscreteMDP) – The MDP instance. Its
rendererslist is overwritten with a text renderer.learner (MDPAgent) – The agent.
horizon (int) – Number of steps to play.
- run(env, learner, horizon)[source]#
Run one interaction and return its cumulative expected score.
- Parameters:
env (DiscreteMDP) – The MDP instance. Reset at the start, and again after any terminal transition.
learner (MDPAgent) – The agent, reset with the initial state.
horizon (int) – Number of steps to play.
- Returns:
Cumulative sum of the expected rewards of the visited state-action pairs.
- Return type:
ndarray of shape (horizon,)