Research

Two organizations I work in, with the same habit: ask a precise question, answer it with reproducible code, and publish the numbers with their limits.

Research & engineering · github.com/simcc-games

SIMCC Games

Simulation studies that give SIMCC evidence when it rules on open questions and designs new games.

Every study plays thousands of seeded games through the real Maths Warriors engine, never a reimplementation from the rulebook. Rule alternatives run in a Python port of the engine, which a parity suite checks against 3,000 golden games move for move. Studies that compare alternatives end with neutral evidence for SIMCC’s ruling and never recommend a rule.

8
studies
3,000
golden games for parity
CC BY 4.0
reports and CSVs
How often the first player wins, when both sides use the same AI
40%50%60%70%80%fairEasy56%Medium65%Hard69%Impossible71%

4,000 games per tier · Wilson 95% intervals · study 3

Strength of the four AI tiers, with two simple heuristics for comparison
1000110012001300easy1000medium1215hard1290impossible1332randombiggest-target

The gaps shrink up the ladder: 214 points from easy to medium, then 76, then 42. A player who always attacks the biggest die it can reach is as strong as medium.

Bradley–Terry fit · 1,000 games per tier pairing, 600 per heuristic · weakest anchored at 1000 · study 6

The first-player advantage turned out to be a race: almost every turn is a capture, so the first player reaches six captures first unless they are forced to skip. In games with no skip, the first player won every time. Whether a player can capture at all depends on how many dice they have left:

Share of target values a player can capture, by number of dice
0%50%100%123456MindStrength

Mind attacks combine two or three dice with + − × ÷. With two dice they reach 13% of targets; with all six, 98%.

20,000 random boards per dice count · study 2

All eight studies

  1. Dice distributionsWhat do the six dice and a whole board produce?
  2. ReachabilityWhich targets can a set of dice capture?
  3. First-player advantageDoes moving first help, and what would the open rule questions change?
  4. Game length and tempoHow long are games, and does an early lead snowball?
  5. The timeout ruleHow often does the clock decide games?
  6. AI tier calibrationAre the AI tiers evenly spaced, and is the top tier safe from a trivial heuristic?
  7. Puzzle bank qualityIs the trainer bank balanced and free of repeats?
  8. The four new gamesAre the new generators solvable, separated by band, and deep enough?

Research tools & platform · github.com/Reason-Sea

ReasonSEA

Open research on how learners in Southeast Asia explain their mathematical reasoning, separately from whether their answer is right.

I built the source-analysis tools and a catalogue of the 64 mathematics items the OECD released from PISA 2012 and 2022, then mapped what the catalogue covers to find gaps before designing original tasks. Alongside it is a team workspace for analysing items, rehearsing draft tasks, and exporting notes.

64
released PISA items catalogued
2012 & 2022
assessment cycles
MIT
open research code
Released PISA mathematics items by content area and process
PISA 2012 · 54 items
FormulateEmployInterpretReasoning
Quantity582n/a
Change and relationships672n/a
Uncertainty and data219n/a
Space and shape372n/a
PISA 2022 · 10 items
FormulateEmployInterpretReasoning
Quantity0210
Change and relationships1001
Uncertainty and data1022
Space and shape0000

The 2022 release has no Space and shape items at all, so task design can’t lean on released examples there. The 2012 framework had no separate Reasoning category, so those cells are not counted as zero.

Reason-Sea/research · data/derived/source-coverage.json · catalogue 2026-09-25.1

Public working draft, September 2026. No learner data has been collected, and the findings describe the item catalogue, not student performance.