Opti_controller#
- class statrl.settings.markovdecisionprocess.discrete_nostructure.agents._Oracle.Opti_controller(env, epsilon=0.001, max_iter=100)[source]#
Bases:
MDPAgentOracle that solves the MDP exactly and follows its optimal policy.
Reads the true transitions and mean rewards from the environment at construction and runs value iteration on them, so it plays optimally from the first step.
- Parameters:
env (DiscreteMDP) – The environment. Must expose
getTransitionandgetMeanReward; otherwiseextractRewardsAndTransitionsis tried as a fallback.nS (int) – Numbers of states and actions.
nA (int) – Numbers of states and actions.
epsilon (float, default=0.001) – Stopping precision recorded on the instance. The value iteration run at construction uses a much tighter
1e-7regardless, so the policy is effectively exact.max_iter (int, default=100) – Iteration cap recorded on the instance; the construction-time run uses 100000.
- policy#
Optimal stochastic policy, uniform over tied optimal actions.
Methods
VI([epsilon, max_iter])Solve the true MDP by value iteration and store the greedy policy.
__init__(env[, epsilon, max_iter])Reader for the transitions and mean reward of a pair.
play(state)Sample an action from the optimal policy for this state.
reset(inistate)Start a new run.
update(state, action, reward, observation)Ignore the transition (the oracle has nothing to learn).
- VI(epsilon=0.01, max_iter=1000)[source]#
Solve the true MDP by value iteration and store the greedy policy.