runLargeMulticoreExperiment#
- statrl.experiments.massiveruns.runLargeMulticoreExperiment(env, agents, oracle, interact, timeHorizon=1000, nbReplicates=100, root_folder='results/')[source]#
Benchmark several agents on one environment and plot their regret.
For each agent it runs
nbReplicatesindependent interactions in parallel, runs the oracle for the same number, computes regret as the oracle’s cumulative score minus each agent’s, and writes a logfile and regret figures underroot_folder.- Parameters:
env (object) – Environment to benchmark on. Must expose
name; an optionaldisplaynameis used as the figure title when present.agents (list of object) – Agents to compare. Their
nameattributes must be distinct — dump filenames and plot legends are keyed on them, so duplicates silently merge two agents’ results.oracle (object) – Reference agent defining zero regret, and the only one required to expose a
policy(it is written to the logfile). Must belong to the same setting asagents.interact (statrl.experiments.onerun.Interaction) – Interaction loop of the setting, shared by every agent in the run.
timeHorizon (int, default=1000) – Number of rounds per interaction.
nbReplicates (int, default=100) – Number of independent runs per agent. Regret quantiles are taken across these, so a handful of replicates gives a very rough band.
root_folder (str, default='results/') – Output directory, created if absent. Must end with a separator.
- Returns:
Everything is written to disk.
root_folderreceives alogfile_*.txt, oneregret_*pickle per agent, and the figuresRegrets_*.png/.pdfin linear and log-y scale. The intermediateaux_*dumps are deleted on the way out.- Return type:
None
See also
statrl.experiments.parallelruns.multicoreRunsThe parallel layer underneath.
statrl.experiments.analyzeruns.computeScoreDiffsTurns the dumps into regret statistics.
statrl.experiments.plotruns.plotScoreDiffsDraws the figures.
Notes
Cost grows as
(len(agents) + 1) * nbReplicates * timeHorizon. Start small — the defaults already amount to 100 000 rounds per agent.Examples
>>> from statrl.settings.bandits.stochastic.anytime.envs.parametric import BernoulliBandit >>> from statrl.settings.bandits.stochastic.anytime.agents.IMED import IMED >>> from statrl.settings.bandits.stochastic.anytime.agents._Oracle import Oracle >>> from statrl.settings.bandits.stochastic.anytime.agents._Random import Random >>> from statrl.settings.bandits.stochastic.anytime.interaction import BanditInteraction >>> from statrl.settings.utils import klBern >>> env = BernoulliBandit([0.2, 0.9, 0.5]) >>> runLargeMulticoreExperiment( ... env, ... agents=[IMED(env.number_arms, klBern), Random(env)], ... oracle=Oracle(env), ... interact=BanditInteraction(), ... timeHorizon=1000, nbReplicates=50, ... )