User guide#

statrl organises reinforcement-learning problems as a taxonomy of settings. Each setting is a self-contained package under statrl.settings defining three things:

  • an environment — the problem (arms and reward distributions, or states, actions, and a transition kernel),

  • an agent — the learning algorithm, and

  • an interaction loop — the function running an agent against an environment for a fixed horizon and recording performance.

New here? Read Core concepts first: it explains the protocol every setting shares, and how regret is defined.

Choosing a setting#

If your problem is…

Use

Reference algorithm

Fixed unknown reward distributions, unknown horizon

Anytime

IMED

The same, but the horizon is known in advance

Known horizon

IMED, via a wrapper

Actions committed in blocks, feedback delayed

Batched bandits

BIMED

A continuous action space with adversarial rewards

Adversarial Lipschitz online optimization

ALF

Actions that move a state

Markov decision processes

IMED-RL, PSRL

Settings#

Running experiments#