Random#

class statrl.settings.bandits.stochastic.batch.agents._Random.Random(env)[source]#

Bases: BatchBanditAgent

Uniform exploration: fill each batch with independent random arms.

Parameters:

env (BatchMAB) – The environment, read only for its number of arms.

Methods

__init__(env)

batchplay(batchsize)

Fill the batch with independent uniform draws.

batchupdate(batcharm, batchreward)

Ignore the batch's rewards (this agent does not learn).

play()

Draw an arm uniformly at random.

reset()

Start a new run.

update(arm, reward)

Ignore the observed reward (this agent does not learn).

batchplay(batchsize)[source]#

Fill the batch with independent uniform draws.

Parameters:

batchsize (int) – Number of pulls in this batch.

Returns:

batchsize independently drawn arm indices.

Return type:

list of int

batchupdate(batcharm, batchreward)[source]#

Ignore the batch’s rewards (this agent does not learn).

Parameters:
  • batcharm (list of int) – The arms that were pulled.

  • batchreward (list of float) – The rewards observed for them.

play()[source]#

Draw an arm uniformly at random.

Returns:

An index drawn uniformly from the arms of the wrapped bandit.

Return type:

int

reset()[source]#

Start a new run. No statistics are kept, so this does nothing.

update(arm, reward)[source]#

Ignore the observed reward (this agent does not learn).

Parameters:
  • arm (int) – Index of the arm that was pulled.

  • reward (float) – Observed reward.