아래의 배경 설명은 아직 번역되지 않아 영어로 표시됩니다.
가중치가 있는 난수 소개
Drawing lots is ancient, but for most of its history the whole point was to make every outcome equally likely. The Athenians machined fairness into hardware: the kleroterion, a stone slab slotted with citizens' bronze identity tickets and fed by a tube of black and white balls, allotted jurors and officials in the fifth and fourth centuries BCE. Deliberately weighting a draw — making one outcome three times as likely as another — belongs to a much later statistical tradition.
Survey statisticians formalised it. Giving every unit of a population the same chance of selection is wasteful when the units differ enormously in size, so Morris Hansen and William Hurwitz set out, in the early 1940s, the theory of sampling with probability proportional to size: a city of a million residents is selected a thousand times more often than a village of a thousand, and the resulting estimate is reweighted to compensate.
Making a computer draw quickly was the next problem. The obvious method — lay the weights end to end along a line, pick a uniform point, and walk along until you pass it — costs time proportional to the number of outcomes on every single draw. In 1974 the British engineer Alastair J. Walker found something better. His alias method chops the weights into n columns of equal height, each holding at most two outcomes: a primary and an "alias". A draw then needs only a random column plus one comparison against a stored threshold, which is constant time no matter how long the list is. Walker published the refined algorithm in ACM Transactions on Mathematical Software in 1977, and Michael Vose gave a numerically stable linear-time way to build the two tables in 1991. Vose's construction is what this page runs.
The same idea arrived in other fields under other names. John Holland's genetic algorithms, set out in the 1970s, selected parents by fitness-proportionate or "roulette wheel" selection, where a candidate's chance of breeding is proportional to how well it scores.
주요 성질
- Each value is drawn with probability equal to its own weight divided by the sum of all the weights, so weights need not add up to 1 or 100.
- Multiplying every weight by the same positive number leaves all the probabilities unchanged: 3:1 and 30:10 are the same distribution.
- A weight of 0 means the value can never be drawn; it keeps its place in the list with probability zero.
- Draws are independent and with replacement, so the same value can come up several times in a row and some values may not appear at all.
- The number of times a value of probability p appears in n draws follows a binomial distribution, with expected value n·p and variance n·p(1−p).
- Walker’s alias method draws in constant time after a one-off setup proportional to the number of values, using one table of thresholds and one of aliases.
- Walking a list of cumulative weights instead costs time proportional to the number of values per draw; binary search over those cumulative weights reduces it to logarithmic time.
- Observed shares converge on the target slowly: the standard error of an estimated proportion is √(p(1−p)/n), so telling a 1% outcome apart from a 2% one reliably takes thousands of draws.
등장하는 곳
- Video game loot tables and gacha pulls are weighted draws; Apple’s App Store Review Guidelines require apps that sell loot boxes to disclose the odds of each item type.
- Phased rollouts and A/B tests route traffic by weight — 95% to the current version, 5% to the candidate — which is a weighted draw made once per request.
- Load balancers implement it directly: nginx and HAProxy both accept a per-server "weight" so that larger machines receive proportionally more connections.
- Fitness-proportionate ("roulette wheel") selection in genetic algorithms picks parents in proportion to their score.
- Large household and education surveys select units with probability proportional to size — enrolment, population, floor area — and then reweight the answers so the sample still represents the whole population.
- The word2vec negative-sampling procedure draws noise words from the corpus unigram distribution raised to the power 3/4, a weighted draw chosen to damp the dominance of very common words.
이 생성기 사용법
생성된 값은 위쪽에 표시되고 옆에 복사 단추가 있습니다. 이미지로 만들려면 이미지 만들기의 스타일에서 모양을 고르고, 내보내기 크기를 정한 뒤 PNG·JPEG·WebP로 내려받으세요. 모두 브라우저에서 그려지므로 생성한 내용이 서버로 전송되지 않습니다.
작업하는 동안 주소창이 갱신되므로, 링크는 항상 지금 보이는 상태를 그대로 재현합니다. 특정 수열을 공유하거나 설정을 저장해 두기에 좋습니다. 값을 일반 텍스트로 가져가려면 복사를, CSV·JSON·NDJSON·SQL·XML이 필요하면 데이터 내보내기를 사용하세요.
출처
- Alias method — Wikipedia — CC BY-SA 4.0
- Categorical distribution — Wikipedia — CC BY-SA 4.0
- Fitness proportionate selection — Wikipedia — CC BY-SA 4.0
- Kleroterion — Wikipedia — CC BY-SA 4.0
- Inverse transform sampling — Wikipedia — CC BY-SA 4.0
이 페이지의 역사적 설명은 위에 나열한 공개 라이선스 자료를 바탕으로 합니다. 잘못된 내용을 발견하셨나요? 알려주시면 바로잡겠습니다.