Core concepts#
Everything in statrl is organised around one idea: a setting is an environment, an agent, and an interaction loop.
The three roles#
Role |
Responsibility |
|---|---|
Environment |
The arms and their reward distributions, or the states, actions, and transition kernel. Samples feedback, and knows the ground truth an agent must not see. |
Agent |
The learning algorithm. Chooses what to do next and updates itself from whatever feedback it observes. |
Interaction |
Runs the two against each other for a fixed horizon and returns the cumulative score used to compute regret. |
The interaction loop#
Each setting’s loop is a subclass of
Interaction. For the anytime stochastic
bandit it is three lines repeated horizon times:
for t in range(horizon):
arm = learner.select_arm() # agent decides
reward = env.step(arm) # environment responds
learner.update(arm, reward) # agent learns
steps_scores[t] = env.expected_reward(arm)
return np.cumsum(steps_scores)
Note that the loop records means, not rewards.** steps_scores[t] is the expected
reward of the arm played, not the reward that was actually observed. The agent
still only ever sees the realized reward. Accumulating means strips the
reward noise out of the score, so the regret curve reflects the quality of the
agent’s choices rather than the luck of its draws.
Regret#
A score is never meaningful on its own, only against what was achievable. So
every experiment also runs an oracle, an agent holding the ground truth:
Oracle
plays the best arm every round.