computeScoreDiffs#
- statrl.experiments.analyzeruns.computeScoreDiffs(names, dump_scores, timeHorizon, envName, root_folder)[source]#
Turn per-replicate score dumps into regret statistics over time.
Loads every dump, subtracts each agent’s cumulative score from the oracle’s averaged one to obtain regret, and summarizes the replicates by their mean, median, and four quantiles.
- Parameters:
names (list of str) – Agent names, in the same order as
dump_scores. Used to name the per-agent regret pickles.dump_scores (list of list of str) – One list of dump filenames per agent. The last entry must be the oracle’s, and it is what every other entry is compared against — the function has no other way to tell which agent is the reference.
timeHorizon (int) – Number of rounds each run played.
envName (str) – Environment name, used in the output filenames.
root_folder (str) – Directory the regret pickles are written to.
- Returns:
mean, median (list of ndarray) – Per-agent mean and median regret at each sampled time.
quantile1, quantile2, quantile3, quantile4 (list of ndarray) – Per-agent regret quantiles at levels 0.1, 0.25, 0.75, and 0.9. The plots shade 0.1-0.9 and 0.25-0.75 as nested bands.
times (list of int) – Sampled time steps, shared by every returned series.
Notes
Long runs are downsampled to at most ~1000 points (
skip = timeHorizon // 1000), which bounds both plot size and memory. Each returned series therefore haslen(times)entries, nottimeHorizon.The oracle’s score is averaged across its replicates before the subtraction, so the result is regret against mean oracle performance rather than a paired difference per replicate.