ErgodicRiverSwim#

class statrl.settings.markovdecisionprocess.discrete_nostructure.envs.riverswim.ErgodicRiverSwim(nbStates, rightProbaright=0.6, rightProbaLeft=0.05, rewardL=0.1, rewardR=1.0, ergodic=0.001, name='RiverSwim')[source]#

Bases: DiscreteMDP

RiverSwim variant in which swimming left may still drift right.

Adds a small ergodic leak to the left action, so every state remains reachable under every policy. Algorithms whose guarantees assume an ergodic MDP need this variant rather than plain RiverSwim.

Parameters:
  • nbStates (int) – Number of states.

  • rightProbaright (float, default=0.6) – Probability that swimming right moves right.

  • rightProbaLeft (float, default=0.05) – Probability that swimming right drifts left instead.

  • rewardL (float, default=0.1) – Reward for going left at the leftmost state.

  • rewardR (float, default=1.0) – Reward for going right at the rightmost state.

  • ergodic (float, default=0.001) – Leak probability on the left action; half of it moves right and the rest stays put. Larger values make exploration easier and the instance less discriminating.

  • name (str, default='RiverSwim') – Label used in logfiles and figures. Shares RiverSwim’s default, so pass a distinct name when comparing the two in one results folder.

See also

RiverSwim

The non-ergodic original.

Methods

__init__(nbStates[, rightProbaright, ...])

change_rendermode(rendermode)

Set the render mode and mark the renderer as needing re-initialization.

close()

Release every attached renderer at the end of a rendered run.

expected_reward(state, arm)

Mean reward of a state-action pair.

getMeanReward(s, a)

Mean reward of a state-action pair.

getTransition(s, a)

Transition distribution of a state-action pair, as a dense vector.

get_wrapper_attr(name)

Gets the attribute name from the environment.

has_wrapper_attr(name)

Checks if the attribute name exists in the environment.

render([mode])

Forward the last (state, action, reward) to every attached renderer.

reset([seed, options])

Start a new episode by drawing a state from the initial distribution.

seed([seed])

Seed the environment's generator.

set_wrapper_attr(name, value, *[, force])

Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.

step(a)

Take an action: draw the next state and a reward.

Attributes

metadata

np_random

Returns the environment's internal _np_random that if not set will initialise with a random seed.

np_random_seed

Returns the environment's internal _np_random_seed that if not set will first initialise with a random int as seed.

render_mode

spec

unwrapped

Returns the base non-wrapped environment.

action_space

observation_space