IMEDRL#
- class statrl.settings.markovdecisionprocess.discrete_nostructure.agents.IMED_RL.IMEDRL(nbr_states, nbr_actions, name='IMED-RL', max_iter=3000, epsilon=0.001, max_reward=1)[source]#
Bases:
MDPAgentIndexed Minimum Empirical Divergence Reinforcement Learning (IMED-RL).
IMED-RL is a model-based reinforcement learning algorithm designed for finite ergodic Markov Decision Processes. The algorithm combines:
empirical estimation of rewards and transition probabilities;
dynamic programming through Value Iteration;
information-theoretic exploration based on the IMED principle.
At every decision step, IMED-RL:
Maintains an empirical MDP from observed transitions and rewards.
Computes the bias function of the empirical optimal policy.
Evaluates an IMED exploration index for every available action.
Selects the action minimizing this index.
To reduce computational complexity, the algorithm also maintains a skeleton of sufficiently explored actions for every state. Value Iteration is performed only over this reduced action set.
- Parameters:
nbr_states (int) – Number of states of the MDP.
nbr_actions (int) – Number of available actions in every state.
name (str, optional) – Name of the learner.
max_iter (int, default=3000) – Maximum number of iterations allowed during Value Iteration.
epsilon (float, default=1e-3) – Stopping tolerance used during Value Iteration.
max_reward (float, default=1) – Known upper bound on instantaneous rewards.
- dirac#
Identity matrix whose rows correspond to Dirac distributions over successor states. It allows transition updates to be performed using simple running averages.
- state_action_pulls#
Number of observations collected for every state-action pair.
- rewards#
Empirical mean reward function.
The array is initialized to 0.5, corresponding to an optimistic uninformed estimate before any observation is collected.
- transitions#
Empirical transition probabilities.
Each transition model is initially uniform over all successor states.
- all_selected#
Boolean indicator specifying whether every action has been sampled at least once in each state.
- Type:
ndarray of shape (nS,)
- phi#
Current estimate of the optimal bias function of the empirical MDP.
- Type:
ndarray of shape (nS,)
- skeleton#
Reduced action set used during planning.
For every state, the skeleton contains only sufficiently sampled actions, thereby reducing the computational cost of Value Iteration.
- index#
Temporary storage of the IMED indices computed for the current state.
- Type:
ndarray of shape (nA,)
- rewards_distributions#
Empirical reward distributions associated with every state-action pair.
Instead of storing only empirical means, IMED-RL keeps the complete empirical distribution of observed rewards because the multinomial IMED index requires the full reward support.
- Type:
Methods
__init__(nbr_states, nbr_actions[, name, ...])Construct an IMED-RL learner.
multinomial_imed(state)Compute the IMED-RL index of every action in a state.
play(state)Choose an action: force-explore until every action is tried, then index.
reset(state)Clear every statistic before a new independent run.
update(state, action, reward, observation)Update the model and refresh the skeleton.
Runs average-reward value iteration over the empirical rewards and transitions, restricted at each state to its skeleton (the actions sampled often enough to be trusted).
- multinomial_imed(state)[source]#
Compute the IMED-RL index of every action in a state.
The index is \(N_{s,a} K_{\inf} + \log N_{s,a}\), exactly the bandit IMED form, but the divergence is taken over the joint distribution of reward plus bias of the next state rather than the reward alone. That joint view is what carries the MDP’s long-run structure into a bandit-style index.
Actions already achieving the best value get the degenerate index \(\log N_{s,a}\).
- Parameters:
state (int) – State whose actions are indexed. Results are written to
self.index.
- reset(state)[source]#
Clear every statistic before a new independent run.
- Parameters:
state (int) – State the environment was reset to.
- update(state, action, reward, observation)[source]#
Update the model and refresh the skeleton.
Updates the running reward mean and transition row for the pair, adds the reward, and recomputes the state’s skeleton (the actions pulled at least \(\log(N_{\max})^2\) times).
- value_iteration()[source]#
Runs average-reward value iteration over the empirical rewards and transitions, restricted at each state to its skeleton (the actions sampled often enough to be trusted).
Stops when successive iterates differ by less than
epsilon, or aftermax_iterationsweeps. Updatesphiin place, normalized to have minimum zero.