statrl#

The Statistical Reinforcement Learning Toolkit — a research library for bandit and reinforcement-learning algorithms, organised as a taxonomy of settings, each with a matching environment, agent, and interaction loop.

statrl gives you a small, explicit protocol shared across settings: an environment holding the problem, an agent that acts and learns, and an interaction loop that runs the two against each other and returns a cumulative-score time series. On top of that sit reference algorithms (IMED, PSRL, IMED-RL, BIMED, BABA, an adversarial Lipschitz forecaster) and an experiments harness for running many replicates in parallel and plotting regret.

Getting started

Install statrl and measure your first regret curve.

Quickstart
User guide

Narrative walkthrough of each setting and the algorithms it ships.

User guide
API reference

Auto-generated reference for every public module, class, and function.

API reference
Examples

Runnable scripts benchmarking agents and plotting regret.

Examples