RandomMDP#

class statrl.settings.markovdecisionprocess.discrete_nostructure.envs.randomMDP.RandomMDP(nbStates, nbActions, maxProportionSupportTransition=0.5, maxProportionSupportReward=0.1, maxProportionSupportStart=0.2, minNonZeroProbability=0.2, minNonZeroReward=0.3, rewardStd=0.5, ergodic=0.0, seed=None, name='RandomMDP')[source]#

Bases: DiscreteMDP

Randomly generated finite MDP with sparse transitions and rewards.

Draws a transition kernel and reward function at random under sparsity constraints, so most state-action pairs pay nothing and each reaches only a few successors.

Parameters:
  • nbStates (int) – Number of states.

  • nbActions (int) – Number of actions.

  • maxProportionSupportTransition (float, default=0.5) – Probability that a given successor gets non-zero mass.

  • maxProportionSupportReward (float, default=0.1) – Probability that a state-action pair is rewarding at all.

  • maxProportionSupportStart (float, default=0.2) – Probability that a state is in the support of the initial distribution.

  • minNonZeroProbability (float, default=0.2) – Floor applied to non-zero probabilities.

  • minNonZeroReward (float, default=0.3) – Floor applied to non-zero mean rewards.

  • rewardStd (float, default=0.5) – Reward standard deviation.

  • ergodic (float, default=0.0) – Floor applied to every transition probability. Leave at 0 for a sparse kernel, or raise it to guarantee an ergodic instance.

  • seed (int, optional) – Seed for the generation.

  • name (str, default='RandomMDP') – Label used in logfiles, plot titles, and dump filenames.

Notes

If the draw yields an all-zero reward function, one random pair is forced to be rewarding, otherwise the instance would be degenerate and every policy optimal.

Methods

__init__(nbStates, nbActions[, ...])

change_rendermode(rendermode)

Set the render mode and mark the renderer as needing re-initialization.

close()

Release every attached renderer at the end of a rendered run.

expected_reward(state, arm)

Mean reward of a state-action pair.

getMeanReward(s, a)

Mean reward of a state-action pair.

getTransition(s, a)

Transition distribution of a state-action pair, as a dense vector.

get_wrapper_attr(name)

Gets the attribute name from the environment.

has_wrapper_attr(name)

Checks if the attribute name exists in the environment.

render([mode])

Forward the last (state, action, reward) to every attached renderer.

reset([seed, options])

Start a new episode by drawing a state from the initial distribution.

reshapeDistribution(distribution, p)

Renormalize a distribution so every non-zero entry is at least p.

seed([seed])

Seed the environment's generator.

set_wrapper_attr(name, value, *[, force])

Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.

sparserand([p, min, max])

Draw a value in [min, max] with probability p, else zero.

step(a)

Take an action: draw the next state and a reward.

Attributes

metadata

np_random

Returns the environment's internal _np_random that if not set will initialise with a random seed.

np_random_seed

Returns the environment's internal _np_random_seed that if not set will first initialise with a random int as seed.

render_mode

spec

unwrapped

Returns the base non-wrapped environment.

action_space

observation_space

reshapeDistribution(distribution, p)[source]#

Renormalize a distribution so every non-zero entry is at least p.

Zeroes entries below p, then repeatedly tops up random entries until the mass is restored, and finally renormalizes. This keeps the generated kernel free of vanishingly small probabilities that would be unlearnable in any realistic horizon.

Parameters:
  • distribution (array-like of float) – The distribution to reshape.

  • p (float) – Minimum mass for a non-zero entry.

Returns:

A distribution summing to one whose non-zero entries are all at least p.

Return type:

list of float

sparserand(p=0.5, min=0.0, max=1.0)[source]#

Draw a value in [min, max] with probability p, else zero.

Parameters:
  • p (float, default=0.5) – Probability of a non-zero draw.

  • min (float, default=0.0, 1.0) – Bounds of the uniform draw when it is non-zero.

  • max (float, default=0.0, 1.0) – Bounds of the uniform draw when it is non-zero.

Returns:

A uniform draw in [min, max], or 0.

Return type:

float