MDPAgent#
- class statrl.settings.markovdecisionprocess.discrete_nostructure.agent.MDPAgent(nS, nA, name='Agent', seed=None)[source]#
Bases:
objectBase class for agents in the discrete MDP setting.
The protocol differs from the bandit one in that both
play()andupdate()are state-aware: an action is chosen for a state, and learning must account for where the process went next.- Parameters:
See also
statrl.settings.markovdecisionprocess.discrete_nostructure.agents.IMED_RL.IMEDRL,statrl.settings.markovdecisionprocess.discrete_nostructure.agents.PSRL.PSRLMethods
__init__(nS, nA[, name, seed])play(state)Choose an action for the given state.
reset(inistate)Start a new independent run from a given initial state.
update(state, action, reward, observation)Learn from one transition.
- reset(inistate)[source]#
Start a new independent run from a given initial state.
- Parameters:
inistate (int) – State the environment was reset to. Agents that track a current state need it; the default only reseeds.