DiscreteMDP#

class statrl.settings.markovdecisionprocess.discrete_nostructure.environment.DiscreteMDP(nS, nA, P, R, isd, nameActions=[], seed=None, name='DiscreteMDP')[source]#

Bases: Env

Finite Markov decision process with no exploitable structure.

Parameters:
  • nS (int) – Number of states.

  • nA (int) – Number of actions.

  • P (dict) – Transition kernel as a dict of dicts, with P[s][a] == [(probability, nextstate, done), ...]. The probabilities of each P[s][a] must sum to one.

  • R (dict) – Reward distributions, with R[s][a] exposing rvs() and mean(), a Dirac for a deterministic reward, or any frozen scipy.stats distribution.

  • isd (array-like of float, shape (nS,)) – Initial state distribution.

  • nameActions (list of str, default=[]) – Human-readable action labels, used by the renderers.

  • seed (int, optional) – Seed for the environment’s generator.

  • name (str, default='DiscreteMDP') – Label used in logfiles, plot titles, and dump filenames.

s#

Current state.

Type:

int

last#

Most recent (state, action, reward), consumed by the renderers.

Type:

tuple

renderers#

Renderers notified on every render() call.

Type:

list

reward_range#

(0, 1), assumed rather than derived from R.

Type:

tuple

Examples

>>> from statrl.settings.markovdecisionprocess.discrete_nostructure.envs.riverswim import RiverSwim
>>> env = RiverSwim(5)
>>> env.nS, env.nA
(5, 2)
>>> state, info = env.reset()
>>> state, reward, done, truncated, info = env.step(0)

Methods

__init__(nS, nA, P, R, isd[, nameActions, ...])

change_rendermode(rendermode)

Set the render mode and mark the renderer as needing re-initialization.

close()

Release every attached renderer at the end of a rendered run.

expected_reward(state, arm)

Mean reward of a state-action pair.

getMeanReward(s, a)

Mean reward of a state-action pair.

getTransition(s, a)

Transition distribution of a state-action pair, as a dense vector.

get_wrapper_attr(name)

Gets the attribute name from the environment.

has_wrapper_attr(name)

Checks if the attribute name exists in the environment.

render([mode])

Forward the last (state, action, reward) to every attached renderer.

reset([seed, options])

Start a new episode by drawing a state from the initial distribution.

seed([seed])

Seed the environment's generator.

set_wrapper_attr(name, value, *[, force])

Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.

step(a)

Take an action: draw the next state and a reward.

Attributes

metadata

np_random

Returns the environment's internal _np_random that if not set will initialise with a random seed.

np_random_seed

Returns the environment's internal _np_random_seed that if not set will first initialise with a random int as seed.

render_mode

spec

unwrapped

Returns the base non-wrapped environment.

action_space

observation_space

change_rendermode(rendermode)[source]#

Set the render mode and mark the renderer as needing re-initialization.

Parameters:

rendermode (str) – The new render mode.

close()[source]#

Release every attached renderer at the end of a rendered run.

expected_reward(state, arm)[source]#

Mean reward of a state-action pair.

Parameters:
  • state (int) – The state.

  • arm (int) – The action taken in it.

Returns:

The true mean reward. Never pass this to a learner.

Return type:

float

getMeanReward(s, a)[source]#

Mean reward of a state-action pair.

Parameters:
  • s (int) – The state.

  • a (int) – The action.

Returns:

The mean of R[s][a]. Read by the oracle; not available to a learning agent.

Return type:

float

getTransition(s, a)[source]#

Transition distribution of a state-action pair, as a dense vector.

Parameters:
  • s (int) – The state.

  • a (int) – The action.

Returns:

Probability of reaching each state. Read by the oracle to solve the MDP; not available to a learning agent.

Return type:

ndarray of shape (nS,)

render(mode='human')[source]#

Forward the last (state, action, reward) to every attached renderer.

Parameters:

mode (str, default='human') – Unused; accepted for gymnasium.Env compatibility.

reset(seed=None, options=None)[source]#

Start a new episode by drawing a state from the initial distribution.

Parameters:
  • seed (int, optional) – Seed for the environment’s generator.

  • options (dict, optional) – Unused; accepted for gymnasium.Env compatibility.

Returns:

  • state (int) – The initial state.

  • info (dict) – {"mean": 0}, matching the shape of what step() returns.

seed(seed=None)[source]#

Seed the environment’s generator.

Parameters:

seed (int, optional) – Seed to use. None draws a fresh one.

Returns:

The seed actually used, in a one-element list.

Return type:

list of int

step(a)[source]#

Take an action: draw the next state and a reward.

Parameters:

a (int) – Action to take in the current state.

Returns:

  • state (int) – The new state.

  • reward (float) – Reward drawn from R[s][a].

  • done (bool) – Whether the transition was terminal.

  • truncated (bool) – Always False; this setting has no time limit.

  • info (dict) – {"mean": float}, the mean reward of the pair. Returned for regret accounting only and must not be given to the learner.