allmax#

statrl.settings.utils.allmax(a)[source]#

Maximum of a together with every index attaining it.

Unlike randmax(), no tie is broken: all maximizers are returned, which lets a caller build a uniform policy over optimal actions.

Parameters:

a (sequence of float) – Values to maximize over.

Returns:

(max_value, indices) where indices lists every position attaining max_value. An empty input returns an empty list.

Return type:

tuple of (float, list of int) or list

Examples

>>> allmax([0.2, 0.5, 0.5, 0.1])
(0.5, [1, 2])