an interactive note · samputhy khim ·

the weight of an upset

the same win can tell a rating system very different things.

ratings were another part of building maths warriors. a useful question is: should every win be worth the same number of points?

beating someone much stronger is a surprising result. beating someone much weaker is closer to what the ratings predicted. an elo-style update measures that difference between expectation and outcome.

start with an expectation

hold your rating at 1400 and move your opponent’s rating below. the curve shows your expected score: 1 for a win, ½ for a draw, and 0 for a loss. an expected score of 50% means half a point per game on average, not necessarily a 50% chance of winning when draws are possible.

01 / change the matchupyour rating: 1400
expected score versus opponent ratingagainst an opponent rated 1600, a player rated 1400 has an expected score of 24.0 percent. expected score decreases as the opponent rating rises.0%50%100%400900140019002400your expected scoreopponent rating
your result
expected score
24.0%
rating change
+24.3
new rating
1424.3

32 × (1 − 0.240) = +24.3

an illustrative logistic elo model, with decimal ratings and a fixed starting rating of 1400. this does not reproduce a particular competition’s rating rules.

the difference is the update

the demo uses this logistic expectation and rating update:

expected = 1 / (1 + 10(opponent − you) / 400)

change = k × (result − expected)

against an equally rated opponent, the expected score is 0.5. with k = 32, a win adds 16 points and a loss subtracts 16. a draw leaves the rating unchanged.

try it: set the opponent to 1800 and keep “win” selected. the expectation falls to about 9.1%, so a win adds about 29.1 points. switch to “loss”: the decrease is only about 2.9 points.

how quickly should a rating move?

the k factor scales the response. doubling it doubles the size of the update; it does not change the expected score. a large value reacts quickly to new results, while a small one moves more cautiously.

try it: select a draw against a stronger opponent. your rating rises, because half a point is more than the model expected. now move the opponent below 1400 and watch the sign reverse.

this is a one-game illustration with a chosen logistic curve. actual rating systems can add rounding, provisional ratings, caps, or different k factors. it is not a specification of the maths warriors implementation.

related reading: fide’s 2024 rating regulations, section 8, which describe expected-score tables and the result-minus-expectation update.

related project: Maths Warriors — engineering case study.