DiscreteMDP#
- class statrl.settings.markovdecisionprocess.discrete_nostructure.environment.DiscreteMDP(nS, nA, P, R, isd, nameActions=[], seed=None, name='DiscreteMDP')[source]#
Bases:
EnvFinite Markov decision process with no exploitable structure.
- Parameters:
nS (int) – Number of states.
nA (int) – Number of actions.
P (dict) – Transition kernel as a dict of dicts, with
P[s][a] == [(probability, nextstate, done), ...]. The probabilities of eachP[s][a]must sum to one.R (dict) – Reward distributions, with
R[s][a]exposingrvs()andmean(), aDiracfor a deterministic reward, or any frozenscipy.statsdistribution.isd (array-like of float, shape (nS,)) – Initial state distribution.
nameActions (list of str, default=[]) – Human-readable action labels, used by the renderers.
seed (int, optional) – Seed for the environment’s generator.
name (str, default='DiscreteMDP') – Label used in logfiles, plot titles, and dump filenames.
See also
statrl.settings.markovdecisionprocess.discrete_nostructure.envs.riverswim.RiverSwimThe standard hard-exploration instance.
statrl.settings.markovdecisionprocess.discrete_nostructure.envs.randomMDP.RandomMDPRandomly generated instances.
Examples
>>> from statrl.settings.markovdecisionprocess.discrete_nostructure.envs.riverswim import RiverSwim >>> env = RiverSwim(5) >>> env.nS, env.nA (5, 2) >>> state, info = env.reset() >>> state, reward, done, truncated, info = env.step(0)
Methods
__init__(nS, nA, P, R, isd[, nameActions, ...])change_rendermode(rendermode)Set the render mode and mark the renderer as needing re-initialization.
close()Release every attached renderer at the end of a rendered run.
expected_reward(state, arm)Mean reward of a state-action pair.
getMeanReward(s, a)Mean reward of a state-action pair.
getTransition(s, a)Transition distribution of a state-action pair, as a dense vector.
get_wrapper_attr(name)Gets the attribute name from the environment.
has_wrapper_attr(name)Checks if the attribute name exists in the environment.
render([mode])Forward the last
(state, action, reward)to every attached renderer.reset([seed, options])Start a new episode by drawing a state from the initial distribution.
seed([seed])Seed the environment's generator.
set_wrapper_attr(name, value, *[, force])Sets the attribute name on the environment with value, see Wrapper.set_wrapper_attr for more info.
step(a)Take an action: draw the next state and a reward.
Attributes
metadatanp_randomReturns the environment's internal
_np_randomthat if not set will initialise with a random seed.np_random_seedReturns the environment's internal
_np_random_seedthat if not set will first initialise with a random int as seed.render_modespecunwrappedReturns the base non-wrapped environment.
action_spaceobservation_space- change_rendermode(rendermode)[source]#
Set the render mode and mark the renderer as needing re-initialization.
- Parameters:
rendermode (str) – The new render mode.
- render(mode='human')[source]#
Forward the last
(state, action, reward)to every attached renderer.- Parameters:
mode (str, default='human') – Unused; accepted for
gymnasium.Envcompatibility.
- reset(seed=None, options=None)[source]#
Start a new episode by drawing a state from the initial distribution.
- Parameters:
seed (int, optional) – Seed for the environment’s generator.
options (dict, optional) – Unused; accepted for
gymnasium.Envcompatibility.
- Returns:
state (int) – The initial state.
info (dict) –
{"mean": 0}, matching the shape of whatstep()returns.
- step(a)[source]#
Take an action: draw the next state and a reward.
- Parameters:
a (int) – Action to take in the current state.
- Returns:
state (int) – The new state.
reward (float) – Reward drawn from
R[s][a].done (bool) – Whether the transition was terminal.
truncated (bool) – Always False; this setting has no time limit.
info (dict) –
{"mean": float}, the mean reward of the pair. Returned for regret accounting only and must not be given to the learner.