an interactive note · samputhy khim ·

a balanced problem set

a random selection inherits the shape of the bank it comes from.

contestkit includes a balancing stage for generated problem sets. the motivation is simple: a collection of good questions does not automatically produce a good contest.

suppose the bank has 24 algebra questions, 16 geometry questions, and eight each in number theory and combinatorics. sampling every question with equal probability gives algebra three times as many opportunities to appear as combinatorics.

draw a set

the demo compares two approaches on that same bank. one draws a random sample. the other reserves equal spaces for all four topics, then fills each topic’s quota with questions closest to your target difficulty. both use the same shuffled bank, so equally suitable questions can vary between samples.

01 / build a problem set56 candidates · sample 1
topic distribution · target: 3 problems per topic
topicrandombalanced
algebra24 in bank
4
3
geometry16 in bank
4
3
number theory8 in bank
2
3
combinatorics8 in bank
2
3

random: 4 / 4 / 2 / 2 across the four topics.
balanced: 3 / 3 / 3 / 3.

average difficulty: 2.25 random / 2.83 balanced.
average distance from your target: 0.92 random / 0.17 balanced.

inspect the selected problems

random

  • c8 · combinatorics3/5
  • a6 · algebra1/5
  • c1 · combinatorics1/5
  • g16 · geometry1/5
  • n8 · number theory3/5
  • g3 · geometry3/5
  • a7 · algebra2/5
  • g13 · geometry3/5
  • g9 · geometry4/5
  • a17 · algebra2/5
  • a2 · algebra2/5
  • n7 · number theory2/5

balanced

  • a23 · algebra3/5
  • a18 · algebra3/5
  • a8 · algebra3/5
  • g3 · geometry3/5
  • g13 · geometry3/5
  • g8 · geometry3/5
  • n8 · number theory3/5
  • n3 · number theory3/5
  • n7 · number theory2/5
  • c8 · combinatorics3/5
  • c3 · combinatorics3/5
  • c2 · combinatorics2/5
a synthetic bank with uneven topic counts and illustrative difficulty labels. random selection samples without replacement. balanced selection takes equal topic quotas, choosing the closest available difficulty labels within each topic.

random is a rule, too

in a random set of 12 questions, the expected algebra count is 12 × 24 / 56, or about 5.14. the expected combinatorics count is about 1.71. any individual sample can differ from those expectations.

try it: draw several samples. the random column changes, while the balanced column stays at three questions per topic. randomness treats the questions equally; a quota controls how much space each topic receives.

difficulty is a second constraint

topic balance alone says nothing about the challenge of a paper. here, each candidate has a difficulty label from 1 to 5. within each topic, the balanced selector sorts candidates by their absolute distance from the target, breaking ties using the shuffled order.

try it: move the difficulty target from 1 to 5, then inspect the selected problems. the random sample stays unchanged until you draw again. the balanced selection changes its questions while keeping its topic counts.

the average difficulty can hide a mixed set: a 1 and a 5 average to 3, just as two 3s do. the second statistic reports average distance from the target. that distinguishes “near the target” from “averages to the target.”

the bank sets the limits

increase the set size to 20 and ask for difficulty 5. each topic must supply five different questions, but some topics have very few hard candidates. the selector has to use more distant labels to meet the quota.

equal topic counts are a design choice for this example. a real contest can specify different weights, a progression in difficulty, curriculum coverage, or limits on similar questions. those requirements should be explicit, and the final paper still needs review.

this demonstration uses invented problem identifiers and difficulty labels. it illustrates the selection logic without reproducing contestkit’s problem bank.

related project: ContestKit — engineering case study.