Tactical Reroll
⋯

The 5 Sep points change as a test of the ledger (3 Oct 2026)

From reports/tiers/2026-10-03-sep-points-test.md , rendered when the site is built.

design/results-methods.md section 6 item 4 (and 5.5 item 4). Command: node tools/pointstest.js --corpus build/grimstat-corpus --out reports/tiers/2026-10-03-sep-points-test (numbers in the JSON beside this file). The ledger at the step-6 baseline (tools/ledger-benchmarks/2026-10-03-step6-baseline.json), inputs frozen at eb3590a (hash 30f0056e6103, BSData 374f505), untouched. Tournament lists from grimstat-corpus 570b8c6 (CC BY 4.0; published by their players and organisers on MiniHeadQuarters), its top-level files only: the listhammer/ folder is not read, and nothing dated after 30 Sep (the tool stops if it meets one). Our own counts and averages; unit names only.

In plain English

The design

Who is in it

Results

Slopes are the change in the outcome for a typical predicted change (x = 1): inclusion in share of the army's lists (0.01 = one percentage point), lift in finishing percentile.

RunOutcomeUnits (moved)Ledger slopeRobust 95%Event bootstrap 95%Permutation pNaive slope (robust 95%)Sign agreement
Primary (by prices, 5 Jul on)inclusion222 (57)+0.004−0.059 to +0.067−0.110 to +0.4900.91−0.011 (−0.069 to +0.047)34 of 53 (64%), p 0.053
lift84 (20)−0.106−0.230 to +0.018−0.555 to +0.4760.31−0.159 (−0.309 to −0.009)6 of 18, p 0.24
Sides by dateinclusion222 (57)+0.009−0.054 to +0.072−0.101 to +0.8730.79−0.00939 of 53 (74%), p 0.001
lift84 (22)−0.051−0.183 to +0.081−0.439 to +1.3920.57−0.0759 of 20
Every month (June too)inclusion233 (61)+0.001−0.059 to +0.061−0.127 to +0.2900.98−0.00737 of 58 (64%), p 0.048
lift101 (22)−0.035−0.162 to +0.093−0.573 to +0.3210.73−0.0888 of 20
Price-only units (profile unchanged)inclusion191 (26)+0.138−0.026 to +0.302−0.177 to +0.5050.067+0.176 (+0.025 to +0.328), p 0.01714 of 24
lift76 (12)−0.123−0.470 to +0.225−0.728 to +0.6660.67−0.1413 of 10

Sign agreement: of the moved units with a predicted change, those whose outcome moved the predicted way against the unchanged units of their army, with a two-sided binomial p.

The pooled slope. On inclusion, +0.004 per typical predicted change: under half a percentage point, with a robust interval of about ±6 points and a permutation p of 0.91. On placing lift, −0.11 percentile, the wrong sign, interval −0.23 to +0.02, permutation p 0.31. Neither is distinguishable from zero, in any of the four runs.

The sign test is the one result near the line: in the primary run 34 of 53 moved units moved the predicted way on inclusion (p = 0.053), and 39 of 53 when the sides are split by date (p = 0.001). The sign of x is the sign of the price change, so this says players took cheaper units more and dearer ones less, which the price alone predicts. The ledger's contribution is the size of the change, and the slopes show no sign of it.

Price-only units (exploratory: one of four looks, with no correction for that). Where the update changed the price and nothing else, the naive price cut predicts inclusion (slope +0.18 per 17% price cut, robust interval +0.02 to +0.33, permutation p 0.017); the ledger's x is weaker (p 0.067). With 26 units in 7 armies it is a lead for the dataslate test, not a finding.

Placings. No version predicts lift; in the primary run the naive slope even leans negative (cheaper units' lists did slightly worse, relative to the army: its robust interval just excludes zero, its permutation p is 0.18). With 20 moved units and lifts whose own errors are around 0.1 to 0.2 percentile each, this is noise until shown otherwise; it is the reason the dataslate forecast says in advance that its lift test will likely be inconclusive.

What limits this test

What next