5 Oct 2026. The review's "hedonic pricing" (reports/tiers/2026-10-04-review.md, recommendation 7). A tool (tools/hedonic.js), this report, its numbers (reports/tiers/2026-10-05-hedonic.json) and a CANDIDATE benchmark (tools/ledger-benchmarks/2026-10-05-hedonic-candidate.json, not confirmed). No ledger behaviour, data file, frozen input or other benchmark changed. Tournament counts are from grimstat-corpus (https://github.com/N041M/grimstat-corpus, commit 570b8c6, CC BY 4.0): army lists published by their players and tournament organisers on MiniHeadQuarters (miniheadquarters.com), July to September 2026, through docs/data/inclusion.json (counts only; its listhammer/ folder is not used). Our own numbers; unit names only.
The short answer
- We tried to teach the ledger's dials what players buy, and the dials can't learn it. Fitting the eight active dials (Kill's weight, Soak, Hold, Score, the melee landing scale and three support multipliers) to which units appear in tournament lists improves the match inside the armies it was fitted on a little, but it does not improve the match on an army it hasn't seen: left out one army at a time, the fitted dials order that army's popular units no better than step 30 does (59.1–59.7% against step 30's 59.6%, at every strength of the pull toward step 30).
- So the honest fit is "stay where you are". The rule fixed in advance (choose the pull by the left-out armies) picks the strongest pull tried but one (λ 0.3), and the dials barely move: Kill 498 → 492, debuff 1.00 → 0.98, the rest within 1%. On the reviewers this candidate is step 30 to within noise: seven lists 73.2% → 73.3% (ρ 0.540 → 0.541); held-out Tyranid units 75.3% → 75.1%; the 14 never-fitted armies 59.7% → 59.6%; 18 units a place closer, 17 a place farther.
- Let the dials move freely and the reviewers disagree more. At the repo's usual pull (λ 0.01, not chosen) the dials move a lot (Kill weight −60%, debuff −32%, command points +24%, Hold +9%, Score +18%), Kill drops from 91% to 79% of a typical unit's Value, and the seven headline lists fall from 73.2% to 70.9% (every Tyranid list down); the 14 armies 59.7% → 59.3%. The pull of popularity on these dials points away from the reviewers on our headline army.
- Are we fitting popularity? Mostly, yes, by the reviewers' own measure. List share alone orders the reviewers right on 70% of pairs over 16 armies (Tyranids 86%), against the ledger's 61%. It beats the ledger in 13 of the 16 (Blood Angels a tie; Space Marines, with 2 lists, and Agents, with 1, the exceptions). The ledger still adds something of its own: with list share held fixed, its order still agrees with the reviewers' consensus (partial ρ 0.33 on Tyranids, 0.18 averaged over 16 armies), and where two units are taken equally often it orders them right 62% of the time. Where list share is wrong, the ledger is at chance (52%).
- What this means. What players buy and what the reviewers grade is mostly something our lines don't measure: the same support, scoring and role units the review named (Neurolictor, Neurotyrant, Gargoyles, Tyrant Guard). Re-weighting Kill, Soak, Hold and Score can't reach them; only a new line can (the mission-card model, role, support). The candidate is recorded for the record; there is no case to adopt it.
1. The target
| Army | Tournament lists | Ledger units (taken at least once) | Pairs 2+ lists apart |
|---|---|---|---|
| Tyranids | 23 | 52 (41) | 1,055 |
| Adeptus Mechanicus | 15 | 34 (28) | 439 |
| Astra Militarum | 15 | 72 (49) | 1,538 |
| Blood Angels | 13 | 20 (14) | 141 |
| Dark Angels | 23 | 93 (56) | 2,600 |
| Death Guard | 16 | 36 (24) | 457 |
| Drukhari | 13 | 37 (22) | 484 |
| Emperor's Children | 22 | 23 (19) | 207 |
| Leagues of Votann | 13 | 22 (21) | 180 |
| Orks | 21 | 55 (35) | 1,139 |
| World Eaters | 12 | 30 (24) | 313 |
| 11 armies | 186 | 474 (333) | 8,553 |
Counts from grimstat-corpus (CC BY 4.0), July–September 2026. Left out for fewer than 12 lists: Space Marines (2), Agents of the Imperium (1), Genestealer Cults (5), Space Wolves (11), Thousand Sons (8).
- Inclusion rate, not share of points spent (decided before any reviewer number was seen): inclusion.json holds counts only (a points share needs each list's units and prices), and a points share grows with a unit's price whenever it is taken, which the ledger's per-point Value divides out on purpose. Inclusion has the opposite bias: cheap fillers and "taxes" are taken often. Both biases are named; neither is corrected.
- A unit never taken counts 0. A chapter's borrowed Space Marine datasheets are counted under the chapter (inclusion.json files them there). Blood Angels' ledger holds only the four borrowed datasheets its panel grades, so 29 borrowed datasheets its players take are outside the comparison; Orks' Lootas [Legends] (5 lists) are not in the ledger.
- Only pairs inside one army, and only pairs 2+ lists apart (a one-list gap is noise, as pairs inside one tier are for the reviewers). Each army's pairs weigh one army in total, so Dark Angels' 2,600 pairs count as much as Blood Angels' 141.
2. The fit
tools/ledgerfit.js's machinery, unchanged except that a panel may now carry a numeric label in place of letters (pairsOfScores): each unit's latent is its ledger Value at its ranking row through the page's compute(); Bradley–Terry on the pairs with a sharpness per army; ridge toward step 30, λ Σ ((θ − θ30) ÷ |θ30|)². Fitted: wKill, wSoak, wHold, wScore, lands, and three rule multipliers chosen before the fit from their carrier counts alone: buff, cp, debuff (the support kinds with 50+ carriers across 7+ of the 11 armies: 76, 56 and 92 units; Synapse, 18 carriers in one army, was left out). Every switch at step 30 (burstKill, threatSoak on); every other knob hibernated at step 30.
λ by leaving one army out, on inclusion only (fit on 10 armies, score the 11th on its inclusion pairs; rule: the largest λ within 0.1 point of the best):
| λ | Held-out inclusion pairwise (mean of 11) | Held-out ρ with list share |
|---|---|---|
| step 30, no fit | 59.6% | 0.229 |
| 0.001 | 59.5% | 0.224 |
| 0.01 | 59.1% | 0.217 |
| 0.03 | 59.4% | 0.222 |
| 0.1 | 59.5% | 0.228 |
| 0.3 (chosen) | 59.7% | 0.230 |
| 1 | 59.5% | 0.229 |
Nothing beats doing nothing by more than 0.1 point. This is the main result: what the dials learn from ten armies' lists doesn't carry to the eleventh.
The fitted dials (SE: bootstrap over units, 20 resamples, at λ 0.3; the λ 0.01 column is for scale, not chosen):
| Knob | Step 30 | Candidate (λ 0.3) | SE | Move | λ 0.01 (not chosen) | Move |
|---|---|---|---|---|---|---|
| r.debuff | 1.000 | 0.980 | 0.028 | −2% | 0.677 | −32% |
| wKill | 498.2 | 492.2 | 3.1 | −1% | 198.1 | −60% |
| r.buff | 1.120 | 1.108 | 0.022 | −1% | 0.896 | −20% |
| r.cp | 1.880 | 1.892 | 0.040 | +1% | 2.326 | +24% |
| wScore | 9.652 | 9.697 | 0.019 | +0% | 11.43 | +18% |
| wHold | 18.89 | 18.98 | 0.048 | +0% | 20.64 | +9% |
| lands | 1.500 | 1.498 | 0.037 | −0% | 1.777 | +18% |
| wSoak | 14.97 | 14.95 | 0.010 | −0% | 13.75 | −8% |
The direction is the one the review guessed (Kill down, Score and Hold up, command points up), and debuffs and buffs down. At λ 0.01 Kill's median share of a unit's Value falls from 91% to 79% (Hold 3% → 7%, Soak 2.5% → 5%, Score 1.6% → 4%). Even then, in-sample inclusion agreement over the 11 armies rises only from 59.6% to 60.3%: the lines can't express most of what players buy.
3. Judged on what it never saw (the reviewers)
No reviewer letter was read by the fit or by the choice of λ (the fitter's problem holds no reviewer panel; letters are read only by --judge and the other tools below, after the candidate was written).
The seven headline lists (tools/ledgerstep.js, step 30 → candidate; the λ 0.01 fit beside it):
| List | Pairwise 30 → cand | ρ 30 → cand | ρ at λ 0.01 |
|---|---|---|---|
| Tyranids Auspex (full) | 82% → 82% | 0.716 → 0.718 | 0.634 |
| Tyranids HivemindHobbies (full) | 81% → 81% | 0.677 → 0.679 | 0.588 |
| Tyranids Into the Hive Mind (full) | 77% → 77% | 0.632 → 0.634 | 0.524 |
| Tyranids Maelstrom | 75% → 75% | 0.633 → 0.631 | 0.555 |
| Tyranids Astrategas | 69% → 69% | 0.434 → 0.439 | 0.413 |
| Space Marines Auspex (full) | 67% → 67% | 0.374 → 0.372 | 0.399 |
| Space Marines Tobias | 62% → 62% | 0.317 → 0.314 | 0.270 |
| Mean | 73.2% → 73.3% | 0.540 → 0.541 | 0.483 (70.9%) |
Held-out Tyranid units (tools/diagnose.js --on, 10 × 5 folds, the same folds): the ledger as it is 75.3% → 75.1% (+0.2 ± 0.6 over step 6, against step 30's +0.4); re-fitted in the folds 74.5% → 74.2%. Inside noise.
The movement view (tools/stepmoves.js): 18 units a place closer, 17 farther, 101 within a place; typical gap 16.9 → 16.9; 50 → 53 of 136 in the reviewers' tier. At λ 0.01: 65 closer, 43 farther, but the typical gap grows 16.9 → 17.0 and the Tyranids' mean gap 9.2 → 10.0 (Gargoyles 47 → 37, Termagants 39 → 22, Biovores 7 → 1 closer; Neurogaunts 32 → 15, Harpy 36 → 20, Maleceptor 18 → 31, Neurotyrant 38 → 47, Neurolictor 28 → 32 farther).
All 16 armies (tools/hedonic.js --judge; the reviewer columns match tools/ledgercarry.js: 16 armies 61.1% → 61.0%, the 14 never-fitted on reviewers 59.7% → 59.6%). List share's agreement with the same reviewers beside it (share from grimstat-corpus, CC BY 4.0; ties in share count half):
| Army | Tournament lists | In the fit | Reviewer lists | Ledger step 30 | Candidate | λ 0.01 | List share | Ledger vs inclusion, 30 → cand |
|---|---|---|---|---|---|---|---|---|
| Tyranids | 23 | yes | 5 | 76.8% | 76.9% | 73.6% | 85.8% | 77.3% → 77.3% |
| Space Marines | 2 | no | 2 | 64.5% | 64.3% | 64.0% | 58.0% | 75.4% → 75.4% |
| Adeptus Mechanicus | 15 | yes | 1 | 73.0% | 73.0% | 66.8% | 74.8% | 65.6% → 65.4% |
| Agents of the Imperium | 1 | no | 1 | 59.6% | 60.5% | 60.1% | 58.3% | – |
| Astra Militarum | 15 | yes | 1 | 66.9% | 66.9% | 65.4% | 72.8% | 59.6% → 59.7% |
| Blood Angels | 13 | yes | 1 | 58.1% | 56.5% | 44.4% | 58.1% | 42.6% → 43.3% |
| Dark Angels | 23 | yes | 1 | 56.3% | 56.6% | 53.8% | 62.6% | 59.4% → 59.5% |
| Death Guard | 16 | yes | 1 | 53.4% | 52.5% | 51.7% | 53.6% | 61.3% → 61.7% |
| Drukhari | 13 | yes | 1 | 57.5% | 57.0% | 58.5% | 77.8% | 55.4% → 55.2% |
| Emperor's Children | 22 | yes | 1 | 66.0% | 66.5% | 69.6% | 76.2% | 58.5% → 58.9% |
| Genestealer Cults | 5 | no | 2 | 64.7% | 64.7% | 66.7% | 76.0% | 61.7% → 61.4% |
| Leagues of Votann | 13 | yes | 2 | 59.2% | 58.7% | 56.9% | 62.8% | 45.6% → 45.6% |
| Orks | 21 | yes | 1 | 63.7% | 63.5% | 60.4% | 64.6% | 61.1% → 61.1% |
| Space Wolves | 11 | no | 1 | 51.4% | 51.4% | 60.0% | 87.5% | 55.1% → 55.1% |
| Thousand Sons | 8 | no | 1 | 38.5% | 38.5% | 42.1% | 70.4% | 44.1% → 44.4% |
| World Eaters | 12 | yes | 3 | 67.8% | 68.3% | 73.6% | 83.9% | 69.0% → 69.6% |
| 16 armies | 61.1% | 61.0% | 60.5% | 70.2% | 59.4% → 59.6% | |||
| the 14 | 59.7% | 59.6% | 59.3% | 69.9% | 56.8% → 57.0% | |||
| in the inclusion fit (11) | 63.5% | 63.3% | 61.3% | 70.3% | 59.6% → 59.7% | |||
| outside it (5) | 55.8% | 55.9% | 58.6% | 70.0% | 59.1% → 59.1% |
The 14 are still never fitted on reviewers, but ten of them are now in the inclusion fit; the five armies outside it (Space Marines, Agents, Genestealer Cults, Space Wolves, Thousand Sons) are the fully unseen check. The λ 0.01 fit does rise there (55.8% → 58.6%, mostly Space Wolves and Thousand Sons, 1 reviewer list each) while it falls on the armies it was fitted on (63.5% → 61.3%): a mixed signal on five thin armies, not a case for it.
4. "Are we fitting popularity?"
Each army's reviewer consensus (mean letter S 5 … D 1 over its lists) against list share and against the ledger (step 30; the candidate is the same to two decimals). Share from grimstat-corpus, CC BY 4.0.
| Army | Pairwise: share | Pairwise: ledger | ρ share | ρ ledger | Partial ρ ledger given share | Ledger right where share is silent (≤ 1 list apart; pairs) | Ledger right where share is wrong (pairs) |
|---|---|---|---|---|---|---|---|
| Tyranids | 85.8% | 76.0% | 0.884 | 0.690 | 0.334 | 65.2% (244) | 55.8% (95) |
| Space Marines | 57.5% | 65.0% | 0.265 | 0.384 | 0.360 | 64.8% (2,928) | – (2) |
| Adeptus Mechanicus | 74.8% | 73.0% | 0.570 | 0.502 | 0.381 | 67.2% (67) | 63.8% (47) |
| Agents of the Imperium | 58.3% | 59.6% | 0.231 | 0.230 | 0.267 | 59.6% (223) | – |
| Astra Militarum | 72.8% | 66.9% | 0.496 | 0.383 | 0.338 | 67.8% (487) | 55.7% (149) |
| Blood Angels | 58.1% | 58.1% | 0.165 | 0.176 | 0.197 | 59.4% (32) | 71.4% (35) |
| Dark Angels | 62.6% | 56.3% | 0.305 | 0.147 | 0.106 | 56.5% (758) | 48.3% (522) |
| Death Guard | 53.6% | 53.4% | 0.070 | 0.073 | 0.056 | 50.4% (137) | 42.6% (136) |
| Drukhari | 77.8% | 57.5% | 0.634 | 0.187 | 0.230 | 47.1% (34) | 75.9% (29) |
| Emperor's Children | 76.2% | 66.0% | 0.577 | 0.355 | 0.283 | 71.9% (32) | 63.3% (30) |
| Genestealer Cults | 75.4% | 65.0% | 0.639 | 0.369 | 0.175 | 53.1% (98) | 50.0% (14) |
| Leagues of Votann | 69.8% | 58.4% | 0.458 | 0.244 | 0.265 | 59.4% (32) | 55.9% (34) |
| Orks | 64.6% | 63.7% | 0.298 | 0.275 | 0.204 | 57.4% (101) | 54.7% (172) |
| Space Wolves | 87.5% | 51.4% | 0.787 | 0.033 | −0.237 | 42.9% (28) | 25.0% (8) |
| Thousand Sons | 70.4% | 38.5% | 0.483 | −0.266 | −0.235 | 37.5% (104) | 40.0% (45) |
| World Eaters | 83.8% | 67.1% | 0.826 | 0.468 | 0.166 | 63.1% (103) | 38.5% (26) |
| Mean of 16 | 70.6% | 61.0% | 0.481 | 0.265 | 0.181 | 62.2% (5,408) | 51.6% (1,344) |
Reading, in plain words:
- List share explains most of the reviewers' order; the ledger about half as much. Mean ρ with the consensus: share 0.48, ledger 0.27. On Tyranids share reaches 0.88, close to the reviewers' agreement with each other.
- The ledger isn't just a noisy copy of popularity. Holding share fixed, it still agrees with the reviewers (partial ρ 0.18 on average, 0.33 on Tyranids, 0.36 on Space Marines, negative only on Space Wolves and Thousand Sons, one reviewer list each); between two equally popular units it picks the reviewers' favourite 62% of the time.
- But where the players and the reviewers agree against the ledger, the ledger has nothing to say (52%, chance, on the 1,344 pairs where share is wrong). Those are the units whose worth isn't in their own dice.
- The ledger agrees with inclusion about as well as it agrees with the reviewers (ρ 0.18 over 16 armies, 0.23 over the 11), and re-weighting its lines can't raise that on an army it hasn't seen (section 2).
5. Doubts and limits
- The ridge's strength was chosen on held-out inclusion, which is nearly flat in λ; the "chosen" λ 0.3 is the rule's tie-break toward step 30, not a sharp optimum. The finding (no transfer) holds at every λ tried.
- Inclusion counts every copy alike: a 60-point tax unit and a centrepiece count the same, and three months of one source (186 lists in the 11 armies) is thin for rare units. Blood Angels' and Space Wolves' ledgers miss most borrowed datasheets.
- The fit is within-army only, so it can't speak to which army is strong; it also can't move a unit whose Value has no line that players pay for (support, list role), which is the point.
- Bootstrap SEs at λ 0.3 are small because the ridge holds the dials; they say the fit is stable, not that it is right.
6. How to repeat
node tools/hedonic.js --threads 3 --boot 20 --write(about 30 minutes on 3 threads: 66 left-out fits, the full fit, 20 resamples);--lambda 0.01for one fit without the λ choice (about 70 s).node tools/hedonic.js --judge --cand tools/ledger-benchmarks/2026-10-05-hedonic-candidate.json(the 16-army tables).node tools/ledgerstep.js --base tools/ledger-benchmarks/2026-10-05-step30-baseline.json --on wKill=492.19,wSoak=14.948,wScore=9.6971,wHold=18.976,lands=1.4977,r.debuff=0.97969,r.buff=1.1078,r.cp=1.8918(and tools/stepmoves.js with the same--on).node tools/ledgercarry.js --on threatSoak,burstKill,wKill=492.19,wSoak=14.948,wScore=9.6971,wHold=18.976,lands=1.4977,r.mortal=1,r.detachment=1.065,r.debuff=0.97969,r.buff=1.1078,r.heal=1.805,r.cp=1.8918 --w 0and tools/diagnose.js--onwith the same values joined by+.