Quickstart#
Install statrl (see Installation), then run an agent against a bandit:
>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit
>>> from statrl.settings.bandits.stochastic.anytime.agents.IMED import IMED
>>> from statrl.settings.bandits.stochastic.anytime.interaction import BanditInteraction
>>> from statrl.settings.utils import klBern
>>>
>>> env = BernoulliBandit([0.2, 0.9, 0.5])
>>> agent = IMED(env.number_arms, kullback=klBern)
>>> scores = BanditInteraction().run(env, agent, horizon=2000)
scores is the cumulative expected reward, one entry per round:
>>> scores.shape
(2000,)
Did it learn? Arm 1 is the best, and it should have taken almost every pull:
>>> env.optimal_arm
1
>>> bool(agent.nbDraws[1] > 0.95 * agent.nbDraws.sum())
True
Measuring regret#
Regret compares that score against an oracle that always plays the best arm. Because the loop accumulates expected rewards, the two scores are directly comparable:
>>> from statrl.settings.bandits.stochastic.anytime.agents._Oracle import Oracle
>>>
>>> oracle_scores = BanditInteraction().run(env, Oracle(env), horizon=2000)
>>> regret = oracle_scores - scores
>>> bool(regret[-1] < 40) # IMED's regret grows logarithmically
True
Compare with uniform exploration, whose regret grows linearly:
>>> from statrl.settings.bandits.stochastic.anytime.agents._Random import Random
>>>
>>> random_scores = BanditInteraction().run(env, Random(env), horizon=2000)
>>> bool((oracle_scores - random_scores)[-1] > 500)
True
That gap — tens versus hundreds — is the whole point of the library, and Running experiments turns it into a plot averaged over replicates.
Watching a run#
On a short run, renderrun prints each pull instead of returning a score.
Arms are labelled A, B, C…:
BanditInteraction().renderrun(env, IMED(env.number_arms, klBern), horizon=5)
Environment: MAB-Bernoulli-means-0.2-0.9-0.5
Actions: ABC
------------------------------
(A) r=0.00
(B) r=1.00
(C) r=1.00
(B) r=1.00
(B) r=1.00
------------------------------
The protocol#
Every bandit setting shares the same three-part protocol:
Component |
Responsibility |
|---|---|
Environment |
Holds the arms; |
Agent |
|
Interaction |
|
See Core concepts for how the loop works and how to plug in your own environment or agent.
Where to go next#
Core concepts — the environment / agent / interaction protocol in detail.
User guide — each setting and the algorithms it ships.
Running experiments — many replicates in parallel, with regret plots.
API reference — the full reference.