Opti_controller#

class statrl.settings.markovdecisionprocess.discrete_nostructure.agents._Oracle.Opti_controller(env, epsilon=0.001, max_iter=100)[source]#

Bases: MDPAgent

Oracle that solves the MDP exactly and follows its optimal policy.

Reads the true transitions and mean rewards from the environment at construction and runs value iteration on them, so it plays optimally from the first step.

Parameters:
  • env (DiscreteMDP) – The environment. Must expose getTransition and getMeanReward; otherwise extractRewardsAndTransitions is tried as a fallback.

  • nS (int) – Numbers of states and actions.

  • nA (int) – Numbers of states and actions.

  • epsilon (float, default=0.001) – Stopping precision recorded on the instance. The value iteration run at construction uses a much tighter 1e-7 regardless, so the policy is effectively exact.

  • max_iter (int, default=100) – Iteration cap recorded on the instance; the construction-time run uses 100000.

policy#

Optimal stochastic policy, uniform over tied optimal actions.

Type:

ndarray of shape (nS, nA)

u#

Bias function from value iteration.

Type:

ndarray of shape (nS,)

Methods

VI([epsilon, max_iter])

Solve the true MDP by value iteration and store the greedy policy.

__init__(env[, epsilon, max_iter])

extractRewardsAndTransitions(s, a)

Reader for the transitions and mean reward of a pair.

play(state)

Sample an action from the optimal policy for this state.

reset(inistate)

Start a new run.

update(state, action, reward, observation)

Ignore the transition (the oracle has nothing to learn).

VI(epsilon=0.01, max_iter=1000)[source]#

Solve the true MDP by value iteration and store the greedy policy.

Parameters:
  • epsilon (float, default=0.01) – Stopping threshold on the span of successive bias differences.

  • max_iter (int, default=1000) – Iteration cap. On reaching it the current iterate is kept and a non-convergence warning is printed.

extractRewardsAndTransitions(s, a)[source]#

Reader for the transitions and mean reward of a pair.

Parameters:
  • s (int) – The state.

  • a (int) – The action.

Returns:

  • transition (ndarray of shape (nS,)) – Probability of reaching each state.

  • reward (float) – Mean reward of the pair.

play(state)[source]#

Sample an action from the optimal policy for this state.

Parameters:

state (int) – Current state.

Returns:

An optimal action, drawn uniformly among ties.

Return type:

int

reset(inistate)[source]#

Start a new run. The policy is fixed, so nothing is cleared.

Parameters:

inistate (int) – Initial state; ignored.

update(state, action, reward, observation)[source]#

Ignore the transition (the oracle has nothing to learn).

Parameters:
  • state (int) – State the action was taken in.

  • action (int) – Action taken.

  • reward (float) – Reward observed.

  • observation (int) – State reached.