When random is not actually random enough

(ersc.io)

26 points | by steveklabnik 4 hours ago

3 comments

  • tialaramex 1 hour ago
    The thing you actually want is rejection sampling: https://en.wikipedia.org/wiki/Rejection_sampling.

    That Wiki page makes it sound very complicated but for this purpose our implementation can be laughably simple which has the advantage that you know why it works and can maintain it properly with confidence.

    Get suitably large inputs, for example if you're trying to pick integers between 2 and 11 inclusive, a nibble (half a byte) would be fine. Now, is the random input in the range you wanted? If so, you've got your answer. If not, throw this random input away and get more.

    Too many programmers act as though random numbers were a precious resource.

    • Dylan16807 47 minutes ago
      Wow that page really gets lost in the weeds of multiple dimensions.

      And yeah it's just rerolling when your random number is out of range. If you want it as simple as possible, always generate from 0-n, and grab barely enough random bits for n to fit.

  • wilbo 1 hour ago
    I got lost when OP talked about using 10 integers to choose from 3 choices. I think I figured out what was missing in the explanation.

    random_u64() Mod 3 does indeed have a single bucket that is oversized. This overweights one option by about 5×10^-20.

    rand() Itself has only 32767 possible values, so it's also common for a bucket to be overweighted depending on the number of buckets.

    • incompatible 1 hour ago
      Makes you wonder at what point overweighting by about 5×10^-20 is something you'd want to care about.
  • fwlr 47 minutes ago
    I don’t think the “random uint” api is too low-level, or lacks a pit of success - I think you’re just reaching for the wrong api. The problem of “make n bits pseudo randomly set to either 1 or 0” is nearby to your problem of “choose an element according to a probability distribution”, but it’s a separate problem in its own right.

    I think actually this is an argument for language designers to include a “std.choice” in their standard library that consumes random bytes and correctly performs common ergonomic operations like “get one element at random from this collection”.

    (If your standard library tries to make a distinction between “regular random number generators” and “cryptographically secured random number generators”, I think this distinction between “generate random bits” and “make probabilistic choices” is about equally important.)