Bibliography#
Papers the shipped algorithms and environments come from. Each is also cited in
the References section of the corresponding class docstring.
Algorithms#
Honda, J. and Takemura, A. “Non-asymptotic analysis of a new bandit algorithm for semi-bounded rewards.” Journal of Machine Learning Research, 16(113):3721-3756, 2015.
Introduces IMED and the non-parametric \(K_{\inf}\). Implemented by
IMED,
KLinf_threshold(), and the batched variant
BIMED.
Pesquerel, F. and Maillard, O.-A. “IMED-RL: Regret optimal learning of ergodic Markov decision processes.” Advances in Neural Information Processing Systems (NeurIPS), 2022.
Extends the IMED index to ergodic MDPs. Implemented by
IMEDRL.
Osband, I., Russo, D. and Van Roy, B. “(More) efficient reinforcement learning via posterior sampling.” Advances in Neural Information Processing Systems (NeurIPS), 2013.
Posterior sampling for reinforcement learning. Implemented by
PSRL.
Maillard, O.-A. and Munos, R. “Online learning in adversarial Lipschitz environments.” European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD), 305-320, 2010.
Discretization plus exponential weights over a continuous metric space.
Implemented by
ALFLearner.
Jin, T., Tang, J., Xu, P., Huang, K., Xiao, X. and Gu, Q. “Almost optimal anytime algorithm for batched multi-armed bandits.” International Conference on Machine Learning (ICML), 2021.
The five-phase epoch structure of
BABA.
Gautron, R., Maillard, O.-A., Preux, P. and Corbeels, M. “Bandits with bounded CVaR constraints.” 2024.
Non-parametric Thompson sampling with a Dirichlet prior anchored at the
reward bound. Implemented, in the CVaR = Expectation regime, by
BCB.
Environments#
Strehl, A. L. and Littman, M. L. “An analysis of model-based interval estimation for Markov decision processes.” Journal of Computer and System Sciences, 74(8):1309-1331, 2008.
Source of the RiverSwim benchmark, implemented by
RiverSwim.
Sutton, R. S., Precup, D. and Singh, S. “Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning.” Artificial Intelligence, 112(1-2):181-211, 1999.
Source of the four-room gridworld, built by
fourRoomMap().
Background#
Lai, T. L. and Robbins, H. “Asymptotically efficient adaptive allocation rules.” Advances in Applied Mathematics, 6(1):4-22, 1985.
The lower bound that “asymptotically optimal” refers to throughout this documentation.