Markov decision processes#

Read more in the user guide.

Protocol#

agent.MDPAgent

Base class for agents in the discrete MDP setting.

environment.DiscreteMDP

Finite Markov decision process with no exploitable structure.

interaction.MDPInteraction

Interaction loop for the discrete MDP setting.

Agents#

agents.IMED_RL.IMEDRL

Indexed Minimum Empirical Divergence Reinforcement Learning (IMED-RL).

agents.PSRL.PSRL

Posterior Sampling Reinforcement Learning.

agents._Oracle.build_opti

Build the oracle for an environment.

agents._Oracle.Opti_controller

Oracle that solves the MDP exactly and follows its optimal policy.

agents._Oracle.Opti_swimmer

Hand-coded oracle for RiverSwim: always swim right.

agents._Random.Random

Uniform exploration: take a random action in every state.

agents.Human.Human

Interactive agent that asks the user for each action.

agents.Human.keyboard_waitfor

Block on stdin until the user types one of the allowed strings.

Environments#

envs.riverswim.RiverSwim

The RiverSwim hard-exploration benchmark MDP.

envs.riverswim.ErgodicRiverSwim

RiverSwim variant in which swimming left may still drift right.

envs.randomMDP.RandomMDP

Randomly generated finite MDP with sparse transitions and rewards.

Renderers#

renderers.textRenderer.TextRenderer

Print each MDP step to stdout as a colourized row of states.