run_lipschitz_online_learning#

statrl.settings.bandits.adversarial.lipschitz.interaction.run_lipschitz_online_learning(env, learner)[source]#

Executes the interaction loop between a learner and the adversarial Lipschitz environment.

This is the canonical sequential learning protocol.

Returns:

rewards – Cumulative reward over the interaction.

Return type:

nparray