Markov decision processes#
Read more in the user guide.
Protocol#
Base class for agents in the discrete MDP setting. |
|
Finite Markov decision process with no exploitable structure. |
|
Interaction loop for the discrete MDP setting. |
Agents#
Indexed Minimum Empirical Divergence Reinforcement Learning (IMED-RL). |
|
Posterior Sampling Reinforcement Learning. |
|
Build the oracle for an environment. |
|
Oracle that solves the MDP exactly and follows its optimal policy. |
|
Hand-coded oracle for RiverSwim: always swim right. |
|
Uniform exploration: take a random action in every state. |
|
Interactive agent that asks the user for each action. |
|
Block on stdin until the user types one of the allowed strings. |
Environments#
The RiverSwim hard-exploration benchmark MDP. |
|
RiverSwim variant in which swimming left may still drift right. |
|
Randomly generated finite MDP with sparse transitions and rewards. |
Renderers#
Print each MDP step to stdout as a colourized row of states. |