an interactive note · samputhy khim ·
a balanced problem set
a random selection inherits the shape of the bank it comes from.
contestkit includes a balancing stage for generated problem sets. the motivation is simple: a collection of good questions does not automatically produce a good contest.
suppose the bank has 24 algebra questions, 16 geometry questions, and eight each in number theory and combinatorics. sampling every question with equal probability gives algebra three times as many opportunities to appear as combinatorics.
draw a set
the demo compares two approaches on that same bank. one draws a random sample. the other reserves equal spaces for all four topics, then fills each topic’s quota with questions closest to your target difficulty. both use the same shuffled bank, so equally suitable questions can vary between samples.
| topic | random | balanced |
|---|---|---|
| algebra24 in bank | ||
| geometry16 in bank | ||
| number theory8 in bank | ||
| combinatorics8 in bank |
random: 4 / 4 / 2 / 2 across the four topics.
balanced: 3 / 3 / 3 / 3.
average difficulty: 2.25 random / 2.83 balanced.
average distance from your target: 0.92 random / 0.17 balanced.
inspect the selected problems
random
- c8 · combinatorics3/5
- a6 · algebra1/5
- c1 · combinatorics1/5
- g16 · geometry1/5
- n8 · number theory3/5
- g3 · geometry3/5
- a7 · algebra2/5
- g13 · geometry3/5
- g9 · geometry4/5
- a17 · algebra2/5
- a2 · algebra2/5
- n7 · number theory2/5
balanced
- a23 · algebra3/5
- a18 · algebra3/5
- a8 · algebra3/5
- g3 · geometry3/5
- g13 · geometry3/5
- g8 · geometry3/5
- n8 · number theory3/5
- n3 · number theory3/5
- n7 · number theory2/5
- c8 · combinatorics3/5
- c3 · combinatorics3/5
- c2 · combinatorics2/5
random is a rule, too
in a random set of 12 questions, the expected algebra count is 12 × 24 / 56, or about 5.14. the expected combinatorics count is about 1.71. any individual sample can differ from those expectations.
try it: draw several samples. the random column changes, while the balanced column stays at three questions per topic. randomness treats the questions equally; a quota controls how much space each topic receives.
difficulty is a second constraint
topic balance alone says nothing about the challenge of a paper. here, each candidate has a difficulty label from 1 to 5. within each topic, the balanced selector sorts candidates by their absolute distance from the target, breaking ties using the shuffled order.
try it: move the difficulty target from 1 to 5, then inspect the selected problems. the random sample stays unchanged until you draw again. the balanced selection changes its questions while keeping its topic counts.
the average difficulty can hide a mixed set: a 1 and a 5 average to 3, just as two 3s do. the second statistic reports average distance from the target. that distinguishes “near the target” from “averages to the target.”
the bank sets the limits
increase the set size to 20 and ask for difficulty 5. each topic must supply five different questions, but some topics have very few hard candidates. the selector has to use more distant labels to meet the quota.
equal topic counts are a design choice for this example. a real contest can specify different weights, a progression in difficulty, curriculum coverage, or limits on similar questions. those requirements should be explicit, and the final paper still needs review.
this demonstration uses invented problem identifiers and difficulty labels. it illustrates the selection logic without reproducing contestkit’s problem bank.
related project: ContestKit — engineering case study.