User guide#
statrl organises reinforcement-learning problems as a taxonomy of
settings. Each setting is a self-contained package under statrl.settings
defining three things:
an environment — the problem (arms and reward distributions, or states, actions, and a transition kernel),
an agent — the learning algorithm, and
an interaction loop — the function running an agent against an environment for a fixed horizon and recording performance.
New here? Read Core concepts first: it explains the protocol every setting shares, and how regret is defined.
Choosing a setting#
If your problem is… |
Use |
Reference algorithm |
|---|---|---|
Fixed unknown reward distributions, unknown horizon |
IMED |
|
The same, but the horizon is known in advance |
IMED, via a wrapper |
|
Actions committed in blocks, feedback delayed |
BIMED |
|
A continuous action space with adversarial rewards |
ALF |
|
Actions that move a state |
IMED-RL, PSRL |