ErgodicRiverSwim#
- class statrl.settings.markovdecisionprocess.discrete_nostructure.envs.riverswim.ErgodicRiverSwim(nbStates, rightProbaright=0.6, rightProbaLeft=0.05, rewardL=0.1, rewardR=1.0, ergodic=0.001, name='RiverSwim')[source]#
Bases:
DiscreteMDPRiverSwim variant in which swimming left may still drift right.
Adds a small
ergodicleak to the left action, so every state remains reachable under every policy. Algorithms whose guarantees assume an ergodic MDP need this variant rather than plainRiverSwim.- Parameters:
nbStates (int) – Number of states.
rightProbaright (float, default=0.6) – Probability that swimming right moves right.
rightProbaLeft (float, default=0.05) – Probability that swimming right drifts left instead.
rewardL (float, default=0.1) – Reward for going left at the leftmost state.
rewardR (float, default=1.0) – Reward for going right at the rightmost state.
ergodic (float, default=0.001) – Leak probability on the left action; half of it moves right and the rest stays put. Larger values make exploration easier and the instance less discriminating.
name (str, default='RiverSwim') – Label used in logfiles and figures. Shares
RiverSwim’s default, so pass a distinct name when comparing the two in one results folder.
See also
RiverSwimThe non-ergodic original.
Methods
__init__(nbStates[, rightProbaright, ...])change_rendermode(rendermode)Set the render mode and mark the renderer as needing re-initialization.
close()Release every attached renderer at the end of a rendered run.
expected_reward(state, arm)Mean reward of a state-action pair.
getMeanReward(s, a)Mean reward of a state-action pair.
getTransition(s, a)Transition distribution of a state-action pair, as a dense vector.
get_wrapper_attr(name)Gets the attribute name from the environment.
has_wrapper_attr(name)Checks if the attribute name exists in the environment.
render([mode])Forward the last
(state, action, reward)to every attached renderer.reset([seed, options])Start a new episode by drawing a state from the initial distribution.
seed([seed])Seed the environment's generator.
set_wrapper_attr(name, value, *[, force])Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.
step(a)Take an action: draw the next state and a reward.
Attributes
metadatanp_randomReturns the environment's internal
_np_randomthat if not set will initialise with a random seed.np_random_seedReturns the environment's internal
_np_random_seedthat if not set will first initialise with a random int as seed.render_modespecunwrappedReturns the base non-wrapped environment.
action_spaceobservation_space