RiverSwim#
- class statrl.settings.markovdecisionprocess.discrete_nostructure.envs.riverswim.RiverSwim(nbStates, rightProbaright=0.6, rightProbaLeft=0.05, rewardL=0.1, rewardR=0.99, name='RiverSwim')[source]#
Bases:
DiscreteMDPThe RiverSwim hard-exploration benchmark MDP.
The agent starts at the left bank of a river of
nbStatesstates. Going left always succeeds and pays a small reward at the leftmost state. Going right pays a large reward at the rightmost state but succeeds only with probabilityrightProbaright, and may even drift back. An agent must therefore give up a certain small reward for many steps to reach an uncertain large one.- Parameters:
nbStates (int) – Number of states, i.e. the length of the river. The longer it is, the harder exploration becomes.
rightProbaright (float, default=0.6) – Probability that swimming right moves right.
rightProbaLeft (float, default=0.05) – Probability that swimming right drifts left instead.
rewardL (float, default=0.1) – Reward for going left at the leftmost state.
rewardR (float, default=0.99) – Reward for going right at the rightmost state.
name (str, default='RiverSwim') – Label used in logfiles, plot titles, and dump filenames.
See also
ErgodicRiverSwimA variant where every state stays reachable under both actions.
Examples
>>> env = RiverSwim(6) >>> env.nS, env.nA (6, 2) >>> env.getMeanReward(5, 0) # large reward at the far bank 0.99
Methods
__init__(nbStates[, rightProbaright, ...])change_rendermode(rendermode)Set the render mode and mark the renderer as needing re-initialization.
close()Release every attached renderer at the end of a rendered run.
expected_reward(state, arm)Mean reward of a state-action pair.
getMeanReward(s, a)Mean reward of a state-action pair.
getTransition(s, a)Transition distribution of a state-action pair, as a dense vector.
get_wrapper_attr(name)Gets the attribute name from the environment.
has_wrapper_attr(name)Checks if the attribute name exists in the environment.
render([mode])Forward the last
(state, action, reward)to every attached renderer.reset([seed, options])Start a new episode by drawing a state from the initial distribution.
seed([seed])Seed the environment's generator.
set_wrapper_attr(name, value, *[, force])Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.
step(a)Take an action: draw the next state and a reward.
Attributes
metadatanp_randomReturns the environment's internal
_np_randomthat if not set will initialise with a random seed.np_random_seedReturns the environment's internal
_np_random_seedthat if not set will first initialise with a random int as seed.render_modespecunwrappedReturns the base non-wrapped environment.
action_spaceobservation_space