About weighted random numbers
Drawing lots is ancient, but for most of its history the whole point was to make every outcome equally likely. The Athenians machined fairness into hardware: the kleroterion, a stone slab slotted with citizens' bronze identity tickets and fed by a tube of black and white balls, allotted jurors and officials in the fifth and fourth centuries BCE. Deliberately weighting a draw — making one outcome three times as likely as another — belongs to a much later statistical tradition.
Survey statisticians formalised it. Giving every unit of a population the same chance of selection is wasteful when the units differ enormously in size, so Morris Hansen and William Hurwitz set out, in the early 1940s, the theory of sampling with probability proportional to size: a city of a million residents is selected a thousand times more often than a village of a thousand, and the resulting estimate is reweighted to compensate.
Making a computer draw quickly was the next problem. The obvious method — lay the weights end to end along a line, pick a uniform point, and walk along until you pass it — costs time proportional to the number of outcomes on every single draw. In 1974 the British engineer Alastair J. Walker found something better. His alias method chops the weights into n columns of equal height, each holding at most two outcomes: a primary and an "alias". A draw then needs only a random column plus one comparison against a stored threshold, which is constant time no matter how long the list is. Walker published the refined algorithm in ACM Transactions on Mathematical Software in 1977, and Michael Vose gave a numerically stable linear-time way to build the two tables in 1991. Vose's construction is what this page runs.
The same idea arrived in other fields under other names. John Holland's genetic algorithms, set out in the 1970s, selected parents by fitness-proportionate or "roulette wheel" selection, where a candidate's chance of breeding is proportional to how well it scores.
Key properties
- Each value is drawn with probability equal to its own weight divided by the sum of all the weights, so weights need not add up to 1 or 100.
- Multiplying every weight by the same positive number leaves all the probabilities unchanged: 3:1 and 30:10 are the same distribution.
- A weight of 0 means the value can never be drawn; it keeps its place in the list with probability zero.
- Draws are independent and with replacement, so the same value can come up several times in a row and some values may not appear at all.
- The number of times a value of probability p appears in n draws follows a binomial distribution, with expected value n·p and variance n·p(1−p).
- Walker’s alias method draws in constant time after a one-off setup proportional to the number of values, using one table of thresholds and one of aliases.
- Walking a list of cumulative weights instead costs time proportional to the number of values per draw; binary search over those cumulative weights reduces it to logarithmic time.
- Observed shares converge on the target slowly: the standard error of an estimated proportion is √(p(1−p)/n), so telling a 1% outcome apart from a 2% one reliably takes thousands of draws.
Where they turn up
- Video game loot tables and gacha pulls are weighted draws; Apple’s App Store Review Guidelines require apps that sell loot boxes to disclose the odds of each item type.
- Phased rollouts and A/B tests route traffic by weight — 95% to the current version, 5% to the candidate — which is a weighted draw made once per request.
- Load balancers implement it directly: nginx and HAProxy both accept a per-server "weight" so that larger machines receive proportionally more connections.
- Fitness-proportionate ("roulette wheel") selection in genetic algorithms picks parents in proportion to their score.
- Large household and education surveys select units with probability proportional to size — enrolment, population, floor area — and then reweight the answers so the sample still represents the whole population.
- The word2vec negative-sampling procedure draws noise words from the corpus unigram distribution raised to the power 3/4, a weighted draw chosen to damp the dominance of very common words.
How to use this generator
The generated values appear at the top, with a copy button beside them. To turn them into an image, pick a look from the style presets under Make an image, choose an export size, and download as PNG, JPEG or WebP. Everything is rendered in your browser, so nothing you generate is sent to a server.
The address bar updates as you work, so the link always reproduces exactly what you see — handy for sharing a specific sequence or saving a configuration for later. Use Copy to take the values as plain text, or Export data for CSV, JSON, NDJSON, SQL or XML.
Sources
- Alias method — Wikipedia — CC BY-SA 4.0
- Categorical distribution — Wikipedia — CC BY-SA 4.0
- Fitness proportionate selection — Wikipedia — CC BY-SA 4.0
- Kleroterion — Wikipedia — CC BY-SA 4.0
- Inverse transform sampling — Wikipedia — CC BY-SA 4.0
Historical summaries on this page draw on the openly licensed references listed above. Spotted an error? Tell us and we will fix it.