Interaction#

class statrl.experiments.onerun.Interaction[source]#

Bases: ABC

Base class for the interaction loop of a setting.

Methods

__init__()

renderrun(env, learner, horizon)

Run one interaction with rendering enabled, returning no score.

run(env, learner, horizon)

Run one interaction and return its cumulative expected score.

Attributes

plotlabels

Axis labels (x, y) for the regret plots.

abstract property plotlabels#

Axis labels (x, y) for the regret plots.

Read by runLargeMulticoreExperiment() and forwarded to plotScoreDiffs(). Settings whose rounds are not time steps override it — the batch setting labels its x-axis by episode.

Type:

tuple of (str, str)

abstractmethod renderrun(env, learner, horizon)[source]#

Run one interaction with rendering enabled, returning no score.

Parameters:
  • env (object) – Environment of the setting; the implementation attaches renderers.

  • learner (object) – Agent of the setting.

  • horizon (int) – Number of rounds to play.

abstractmethod run(env, learner, horizon)[source]#

Run one interaction and return its cumulative expected score.

Parameters:
  • env (object) – Environment of the setting. Reset by the implementation.

  • learner (object) – Agent of the setting. Reset by the implementation, so one instance can serve many replicates.

  • horizon (int) – Number of rounds to play.

Returns:

Cumulative expected reward. Implementations must return exactly horizon entries — oneRunWithDump() asserts it.

Return type:

ndarray of shape (horizon,)