RiverSwim#

class statrl.settings.markovdecisionprocess.discrete_nostructure.envs.riverswim.RiverSwim(nbStates, rightProbaright=0.6, rightProbaLeft=0.05, rewardL=0.1, rewardR=0.99, name='RiverSwim')[source]#

Bases: DiscreteMDP

The RiverSwim hard-exploration benchmark MDP.

The agent starts at the left bank of a river of nbStates states. Going left always succeeds and pays a small reward at the leftmost state. Going right pays a large reward at the rightmost state but succeeds only with probability rightProbaright, and may even drift back. An agent must therefore give up a certain small reward for many steps to reach an uncertain large one.

Parameters:
  • nbStates (int) – Number of states, i.e. the length of the river. The longer it is, the harder exploration becomes.

  • rightProbaright (float, default=0.6) – Probability that swimming right moves right.

  • rightProbaLeft (float, default=0.05) – Probability that swimming right drifts left instead.

  • rewardL (float, default=0.1) – Reward for going left at the leftmost state.

  • rewardR (float, default=0.99) – Reward for going right at the rightmost state.

  • name (str, default='RiverSwim') – Label used in logfiles, plot titles, and dump filenames.

nameActions#

["R", "L"] — action 0 is right, action 1 is left.

Type:

list of str

See also

ErgodicRiverSwim

A variant where every state stays reachable under both actions.

Examples

>>> env = RiverSwim(6)
>>> env.nS, env.nA
(6, 2)
>>> env.getMeanReward(5, 0)      # large reward at the far bank
0.99

Methods

__init__(nbStates[, rightProbaright, ...])

change_rendermode(rendermode)

Set the render mode and mark the renderer as needing re-initialization.

close()

Release every attached renderer at the end of a rendered run.

expected_reward(state, arm)

Mean reward of a state-action pair.

getMeanReward(s, a)

Mean reward of a state-action pair.

getTransition(s, a)

Transition distribution of a state-action pair, as a dense vector.

get_wrapper_attr(name)

Gets the attribute name from the environment.

has_wrapper_attr(name)

Checks if the attribute name exists in the environment.

render([mode])

Forward the last (state, action, reward) to every attached renderer.

reset([seed, options])

Start a new episode by drawing a state from the initial distribution.

seed([seed])

Seed the environment's generator.

set_wrapper_attr(name, value, *[, force])

Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.

step(a)

Take an action: draw the next state and a reward.

Attributes

metadata

np_random

Returns the environment's internal _np_random that if not set will initialise with a random seed.

np_random_seed

Returns the environment's internal _np_random_seed that if not set will first initialise with a random int as seed.

render_mode

spec

unwrapped

Returns the base non-wrapped environment.

action_space

observation_space