run_lipschitz_online_learning#
- statrl.settings.bandits.adversarial.lipschitz.interaction.run_lipschitz_online_learning(env, learner)[source]#
Executes the interaction loop between a learner and the adversarial Lipschitz environment.
This is the canonical sequential learning protocol.
- Returns:
rewards – Cumulative reward over the interaction.
- Return type:
nparray