design/results-methods.md section 6 item 4 (and 5.5 item 4). Command: node tools/pointstest.js --corpus build/grimstat-corpus --out reports/tiers/2026-10-03-sep-points-test (numbers in the JSON beside this file). The ledger at the step-6 baseline (tools/ledger-benchmarks/2026-10-03-step6-baseline.json), inputs frozen at eb3590a (hash 30f0056e6103, BSData 374f505), untouched. Tournament lists from grimstat-corpus 570b8c6 (CC BY 4.0; published by their players and organisers on MiniHeadQuarters), its top-level files only: the listhammer/ folder is not read, and nothing dated after 30 Sep (the tool stops if it meets one). Our own counts and averages; unit names only.
In plain English
- The question: on 5 Sep, 129 units changed points. The ledger can say, for each, how much its worth per 100 points moved from the price change alone. Did the units it says got better value get taken more often, and did they lift their lists' placings more, than the units of the same army the update left alone?
- The answer: we can't tell. The pooled slope is near zero for inclusion and slightly the wrong way for placings, and neither is distinguishable from zero. The test is small: 83 lists after the update in the armies with a ledger, from 10 events, and over half the moved units it can use are Orks (5 Ork lists after).
- One weak, honest signal: the moved units' inclusion did go the predicted way more often than not (34 of 53, 64%, p = 0.053 against a coin). But that direction is just "cheaper, so taken more", which the price alone predicts; the ledger's size of the change adds nothing measurable.
- For the units whose rules didn't also change (26 of them), cheaper units were taken more: the plain price cut has a slope clear of zero (permutation p = 0.017), the ledger's version a little weaker (p = 0.067). This was one of four looks, so it is a lead, not a finding.
- What it means for the ledger: this test can't confirm or refute it. The 1 Oct dataslate (188 units, three times as many on the units' side) is the real test; its forecast and scoring rule are pre-registered in
2026-10-03-dataslate-forecast.md.
The design
- Which costs a list played. The price written in each list. Per event, we count the lines of units the update moved whose price equals the old or the new cost at the smallest size; the majority decides (results-methods.md section 4.7). All 35 events were decided by their prices, none by date. Four events dated before 5 Sep already show the new costs (they played Games Workshop's release before BSData's update); no event after 5 Sep shows the old ones.
- Window. Events from 5 Jul to 26 Sep (June leaves out: it mixes editions, and the 4 Jul event was a farewell to the old one). Before: 110 clean placed lists from 11 events; after: 83 lists from 10 events, in the armies with a ledger.
- The prediction, x. Each unit's best solo form in the ledger, valued at the step-6 baseline. The ledger's BSData (374f505) is after 5 Sep, so its worth per 100 points is at the new price; the price change alone keeps the form's worth in points and changes only its cost, so the change in worth per 100 is worth × (1 − new ÷ old), positive when the unit got cheaper. Prices at the form's own size where the update lists them, else the smallest size's ratio. Units the update didn't touch are the controls, x = 0; units whose profile alone moved are left out. A typical predicted change (the root mean square over the moved units) is 158 points of worth per 100; x is in those units below.
- The naive predictor, beside it: the log price cut, ln(old ÷ new), which knows nothing of the ledger (typical size 0.17, about a 17% price move).
- The outcomes, per unit, after minus before. Inclusion: the share of its army's lists that field it. Lift: the mean finishing percentile (1 won, 0 last) of its army's lists with it minus without it.
- The fit. Weighted least squares with army fixed effects (each unit against the other units of its own army: the difference in differences), weights 1 ÷ (1/n before + 1/n after) (the army's lists for inclusion; with and without on both sides for lift). Data rule: armies with 5+ lists on each side; units seen in 2+ lists (inclusion) or with 2+ lists with and 2+ without on each side (lift).
- Uncertainty, three ways: a robust (HC1) 95% interval; a bootstrap over events (each side's events resampled, 2,000 draws), which is lopsided with ten events a side; and a permutation test, x shuffled among the units of each army (2,000 shuffles), the cleanest of the three here.
Who is in it
- Units: 57 moved units and 165 controls for inclusion, in 8 armies (the other 8 ledger armies have fewer than 5 lists on a side). The audit's 79 moved units seen on both sides count every army; only the 16 with a ledger have a prediction.
- By army (moved units, inclusion): Orks 31, Astra Militarum 8, Tyranids 6, Death Guard 4, Adeptus Mechanicus 3, Emperor's Children 2, Dark Angels 2, World Eaters 1.
- Orks dominate, and the 5 Sep update rewrote most Ork datasheets as well as their points: 12 Ork lists before, 5 after. The price-only prediction misses those rules changes by construction.
- Lift: 20 moved units and 64 controls clear the 2-and-2 rule.
Results
Slopes are the change in the outcome for a typical predicted change (x = 1): inclusion in share of the army's lists (0.01 = one percentage point), lift in finishing percentile.
| Run | Outcome | Units (moved) | Ledger slope | Robust 95% | Event bootstrap 95% | Permutation p | Naive slope (robust 95%) | Sign agreement |
|---|---|---|---|---|---|---|---|---|
| Primary (by prices, 5 Jul on) | inclusion | 222 (57) | +0.004 | −0.059 to +0.067 | −0.110 to +0.490 | 0.91 | −0.011 (−0.069 to +0.047) | 34 of 53 (64%), p 0.053 |
| lift | 84 (20) | −0.106 | −0.230 to +0.018 | −0.555 to +0.476 | 0.31 | −0.159 (−0.309 to −0.009) | 6 of 18, p 0.24 | |
| Sides by date | inclusion | 222 (57) | +0.009 | −0.054 to +0.072 | −0.101 to +0.873 | 0.79 | −0.009 | 39 of 53 (74%), p 0.001 |
| lift | 84 (22) | −0.051 | −0.183 to +0.081 | −0.439 to +1.392 | 0.57 | −0.075 | 9 of 20 | |
| Every month (June too) | inclusion | 233 (61) | +0.001 | −0.059 to +0.061 | −0.127 to +0.290 | 0.98 | −0.007 | 37 of 58 (64%), p 0.048 |
| lift | 101 (22) | −0.035 | −0.162 to +0.093 | −0.573 to +0.321 | 0.73 | −0.088 | 8 of 20 | |
| Price-only units (profile unchanged) | inclusion | 191 (26) | +0.138 | −0.026 to +0.302 | −0.177 to +0.505 | 0.067 | +0.176 (+0.025 to +0.328), p 0.017 | 14 of 24 |
| lift | 76 (12) | −0.123 | −0.470 to +0.225 | −0.728 to +0.666 | 0.67 | −0.141 | 3 of 10 |
Sign agreement: of the moved units with a predicted change, those whose outcome moved the predicted way against the unchanged units of their army, with a two-sided binomial p.
The pooled slope. On inclusion, +0.004 per typical predicted change: under half a percentage point, with a robust interval of about ±6 points and a permutation p of 0.91. On placing lift, −0.11 percentile, the wrong sign, interval −0.23 to +0.02, permutation p 0.31. Neither is distinguishable from zero, in any of the four runs.
The sign test is the one result near the line: in the primary run 34 of 53 moved units moved the predicted way on inclusion (p = 0.053), and 39 of 53 when the sides are split by date (p = 0.001). The sign of x is the sign of the price change, so this says players took cheaper units more and dearer ones less, which the price alone predicts. The ledger's contribution is the size of the change, and the slopes show no sign of it.
Price-only units (exploratory: one of four looks, with no correction for that). Where the update changed the price and nothing else, the naive price cut predicts inclusion (slope +0.18 per 17% price cut, robust interval +0.02 to +0.33, permutation p 0.017); the ledger's x is weaker (p 0.067). With 26 units in 7 armies it is a lead for the dataslate test, not a finding.
Placings. No version predicts lift; in the primary run the naive slope even leans negative (cheaper units' lists did slightly worse, relative to the army: its robust interval just excludes zero, its permutation p is 0.18). With 20 moved units and lifts whose own errors are around 0.1 to 0.2 percentile each, this is noise until shown otherwise; it is the reason the dataslate forecast says in advance that its lift test will likely be inconclusive.
What limits this test
- Size. 83 lists after, from 10 events; 8 armies with 5+ lists a side. Below results-methods.md section 5.6's bar for a balance-update test (100+ lists after).
- Orks. 31 of the 57 moved units are Orks, measured on 5 lists after the update, and nearly every one changed its rules too. The price-only prediction can't see a rules change.
- Price only. The ledger's own BSData is after the update, so it can't value the old datasheets; a full before-and-after ledger would need a freeze at a pre-5 Sep commit (BSData 46d8cc5) and the effects files of that time.
- Belief, not worth. Inclusion is what players believe and copy (ALSA's lesson, section 3.7), and it lags; three weeks after an update is short.
- The usual limits of this corpus: French and Belgian local events, no players, no games.
What next
- The 1 Oct dataslate forecast (
2026-10-03-dataslate-forecast.md) is the main test: 188 units moved, the scoring rule and its code fixed before October's files arrive. - Re-run this test when the corpus's September is complete (its file ends 26 Sep), and add the dataslate's before side to it.