IMEDRL#

class statrl.settings.markovdecisionprocess.discrete_nostructure.agents.IMED_RL.IMEDRL(nbr_states, nbr_actions, name='IMED-RL', max_iter=3000, epsilon=0.001, max_reward=1)[source]#

Bases: MDPAgent

Indexed Minimum Empirical Divergence Reinforcement Learning (IMED-RL).

IMED-RL is a model-based reinforcement learning algorithm designed for finite ergodic Markov Decision Processes. The algorithm combines:

  • empirical estimation of rewards and transition probabilities;

  • dynamic programming through Value Iteration;

  • information-theoretic exploration based on the IMED principle.

At every decision step, IMED-RL:

  1. Maintains an empirical MDP from observed transitions and rewards.

  2. Computes the bias function of the empirical optimal policy.

  3. Evaluates an IMED exploration index for every available action.

  4. Selects the action minimizing this index.

To reduce computational complexity, the algorithm also maintains a skeleton of sufficiently explored actions for every state. Value Iteration is performed only over this reduced action set.

Parameters:
  • nbr_states (int) – Number of states of the MDP.

  • nbr_actions (int) – Number of available actions in every state.

  • name (str, optional) – Name of the learner.

  • max_iter (int, default=3000) – Maximum number of iterations allowed during Value Iteration.

  • epsilon (float, default=1e-3) – Stopping tolerance used during Value Iteration.

  • max_reward (float, default=1) – Known upper bound on instantaneous rewards.

nS#

Number of states.

Type:

int

nA#

Number of actions.

Type:

int

dirac#

Identity matrix whose rows correspond to Dirac distributions over successor states. It allows transition updates to be performed using simple running averages.

Type:

ndarray of shape (nS, nS)

actions#

Array containing all action identifiers.

Type:

ndarray of shape (nA,)

max_iteration#

Maximum number of Value Iteration updates.

Type:

int

epsilon#

Convergence tolerance for Value Iteration.

Type:

float

max_reward#

Known upper bound on rewards.

Type:

float

state_action_pulls#

Number of observations collected for every state-action pair.

Type:

ndarray of shape (nS, nA)

state_visits#

Number of visits to every state.

Type:

ndarray of shape (nS,)

rewards#

Empirical mean reward function.

The array is initialized to 0.5, corresponding to an optimistic uninformed estimate before any observation is collected.

Type:

ndarray of shape (nS, nA)

transitions#

Empirical transition probabilities.

Each transition model is initially uniform over all successor states.

Type:

ndarray of shape (nS, nA, nS)

all_selected#

Boolean indicator specifying whether every action has been sampled at least once in each state.

Type:

ndarray of shape (nS,)

phi#

Current estimate of the optimal bias function of the empirical MDP.

Type:

ndarray of shape (nS,)

skeleton#

Reduced action set used during planning.

For every state, the skeleton contains only sufficiently sampled actions, thereby reducing the computational cost of Value Iteration.

Type:

dict[int, ndarray]

index#

Temporary storage of the IMED indices computed for the current state.

Type:

ndarray of shape (nA,)

rewards_distributions#

Empirical reward distributions associated with every state-action pair.

Instead of storing only empirical means, IMED-RL keeps the complete empirical distribution of observed rewards because the multinomial IMED index requires the full reward support.

Type:

dict

s#

Current state of the learner.

Type:

int or None

Methods

__init__(nbr_states, nbr_actions[, name, ...])

Construct an IMED-RL learner.

multinomial_imed(state)

Compute the IMED-RL index of every action in a state.

play(state)

Choose an action: force-explore until every action is tried, then index.

reset(state)

Clear every statistic before a new independent run.

update(state, action, reward, observation)

Update the model and refresh the skeleton.

value_iteration()

Runs average-reward value iteration over the empirical rewards and transitions, restricted at each state to its skeleton (the actions sampled often enough to be trusted).

multinomial_imed(state)[source]#

Compute the IMED-RL index of every action in a state.

The index is \(N_{s,a} K_{\inf} + \log N_{s,a}\), exactly the bandit IMED form, but the divergence is taken over the joint distribution of reward plus bias of the next state rather than the reward alone. That joint view is what carries the MDP’s long-run structure into a bandit-style index.

Actions already achieving the best value get the degenerate index \(\log N_{s,a}\).

Parameters:

state (int) – State whose actions are indexed. Results are written to self.index.

play(state)[source]#

Choose an action: force-explore until every action is tried, then index.

Parameters:

state (int) – Current state.

Returns:

The least-pulled action while some action of this state has never been taken, afterwards the minimal-index action, following a fresh value iteration.

Return type:

int

reset(state)[source]#

Clear every statistic before a new independent run.

Parameters:

state (int) – State the environment was reset to.

update(state, action, reward, observation)[source]#

Update the model and refresh the skeleton.

Updates the running reward mean and transition row for the pair, adds the reward, and recomputes the state’s skeleton (the actions pulled at least \(\log(N_{\max})^2\) times).

Parameters:
  • state (int) – State the action was taken in.

  • action (int) – Action taken.

  • reward (float) – Observed reward.

  • observation (int) – State reached; also becomes the agent’s current state.

value_iteration()[source]#

Runs average-reward value iteration over the empirical rewards and transitions, restricted at each state to its skeleton (the actions sampled often enough to be trusted).

Stops when successive iterates differ by less than epsilon, or after max_iteration sweeps. Updates phi in place, normalized to have minimum zero.