5 Oct 2026. Jordan: "is HIT wrong, SOAK wrong, the whole point system wrong?" Two tests on the a20dbd2c3316 snapshot, Track A (step 34b), the exam's pool (15 armies on their 11th-edition lists). A report only: no ledger value, setting, page or published letter changes. Popularity never enters a letter.
The answer in plain English
Mostly the second: the reviewers rank on something our datasheet maths doesn't see, and no setting of our knobs finds it.
- What players take explains the reviewers far better than our maths does. Add each unit's tournament list share to Track A and agreement across the 15 armies goes from 59.1% to 72.2%. That holds on armies the strength wasn't chosen on (+12.5 points held out, ± 3.2). The same thing with list shares shuffled among each army's units gains nothing (+0.6). Popularity alone, with our Value only breaking its ties, scores 73.0%: better than our maths on 13 of the 15 armies. World Eaters reach 83.8%, inside the reviewers' own 82 to 85. Our Value and list share barely agree (rank correlation 0.18 over the 15).
- Letting every knob loose on World Eaters finds nothing real. With all 136 knobs free (line weights, the rule multipliers, a slope on every datasheet feature, as T-662 did on Tyranids), the fit goes from 67.5% to 94.3%. That is above the reviewers' own agreement. Shuffled letters reach 94.3% too, on all five seeds. So the whole rise is memorising: 136 knobs on 30 units can buy any order. On units the fit never saw, it scores 64.4%, against Track A's 67.3%. Nothing in the datasheets, weighted any way we can weight it, predicts the reviewers' order on new units better than we already do.
- So it isn't one weight that's wrong. Turning hit or soak up or down wouldn't close the gap. If it would, the free fit would carry over to unseen units, and it doesn't. The missing part is the game value players act on: the missions, objectives and screening, which units go together, a detachment's or stratagem's best use, and what's hot. Players act on it when they pick lists, and the reviewers grade it. Our Value models a unit trading dice with targets. Jakhals (S S S from every World Eaters reviewer, 35th percentile in our Value) and the Chaos Rhino are the clearest cases (T-672).
- One caution. Reviewers read tournament results too, so some of popularity's lift is the reviewers and the players sharing a source, not separate proof. It still tells us where the information is: in how the game is played, not in the dice.
What it means for the work: stop hunting for the right weight on Kill or Soak. The gains left are in modelling what a unit does for a list: objectives, screens, delivery and synergy (T-695 army rules, T-700 total worth). We could also say openly, on the public pages, that our letters measure a unit's dice-and-points worth, and show popularity beside them as a separate column. That would be a product question for Jordan; it never feeds a letter.
T-701: how much of the reviewers is popularity
Method. tools/popshare.js through the exam's post hook (T-685): each unit's list share is the share of its army's 11th-edition tournament lists that take it. The source is tools/popularity.js (T-693): 487 lists from 30 events, 27 June to 26 September, grimstat-corpus, CC BY 4.0. Space Marines pool the Codex chapters without a ledger of their own. z is that share standardised over the army's ranked units. The settings:
- Value′ = Value × max(0.05, 1 + p × z), at p 0.05 to 5;
- popularity alone (ties count wrong, as the exam counts them);
- popularity with Track A breaking its ties;
- a 50/50 rank blend.
Held out by army as T-682 to T-684 were: 3 splits, the strength chosen on 10 armies (0 included) and scored on 5. The scrambled control shuffles list share among each army's units (5 seeds) and goes through the same choice. Exam intervals come from tools/exam.js with --post tools/popshare.js (1,000 draws); its numbers equal popshare's to the decimal.
Held out (gain over Track A on the 5 armies left out):
| Family | Held out (picks) | ± SE | Scrambled held out | ± SE |
|---|---|---|---|---|
| Value × (1 + p × z) | +12.54 (p 5; 2; 2) | 3.15 | +0.61 (0.05; 0.05; 0.05) | 0.15 |
| Popularity alone (or Track A) | +7.16 | 4.39 | +0.00 (Track A every fold) | 0.00 |
| Popularity, Track A breaking ties | +13.00 | 3.78 | +0.00 (Track A every fold) | 0.00 |
| 50/50 rank blend | +9.38 | 2.12 | +0.00 (Track A every fold) | 0.00 |
Pooled (in sample) by strength: p 0.1 63.0, p 0.3 66.9, p 0.5 69.2, p 1 71.3, p 2 72.2, p 5 72.0 (Track A 59.1). The curve flattens past p 1: popularity then ranks and Value only orders within it. Scrambled at p 0.3 59.1, at p 2 54.8. Exam at p 2: +13.1 [+8.9, +17.2], 14 of 15 armies better. Scrambled at p 2: −1.6 [−7.1, +3.6]. At p 0.3: +7.8 [+5.4, +10.4]. Blend: +10.2 [+7.0, +13.5], 15 of 15.
Per army (pairwise on the 11th-edition lists; popularity at p 2, the held-out pick):
| Army (panel lists; tournament lists) | Track A | + popularity (p 2) | Scrambled at p 2 | Popularity alone | Popularity, ties by Track A | 50/50 blend | Ceiling (ours; known) |
|---|---|---|---|---|---|---|---|
| Space Marines (2; 33) | 60.3 | 80.0 | 55.1 | 75.1 | 81.5 | 75.1 | 90.5; 91 |
| Adeptus Mechanicus (1; 18) | 71.7 | 76.4 | 56.1 | 71.1 | 75.2 | 78.0 | – |
| Agents of the Imperium (1; 1) | 56.5 | 63.2 | 54.6 | 28.7 | 63.2 | 64.1 | – |
| Astra Militarum (1; 17) | 66.6 | 77.0 | 57.6 | 64.4 | 76.4 | 75.2 | – |
| Blood Angels (1; 14) | 61.3 | 60.5 | 58.1 | 52.4 | 60.5 | 62.9 | – |
| Dark Angels (1; 28) | 53.9 | 59.1 | 52.8 | 54.2 | 60.8 | 59.5 | – |
| Death Guard (1; 19) | 53.4 | 57.1 | 52.6 | 47.3 | 53.4 | 56.1 | – |
| Drukhari (1; 14) | 58.0 | 70.5 | 51.5 | 73.9 | 76.8 | 69.6 | – |
| Emperor's Children (1; 24) | 69.6 | 79.6 | 57.6 | 73.3 | 77.0 | 78.0 | – |
| Genestealer Cults (2; 5) | 56.4 | 70.3 | 53.5 | 69.4 | 74.8 | 66.3 | 84.9 |
| Leagues of Votann (1; 15) | 61.0 | 86.7 | 60.6 | 83.8 | 87.6 | 78.1 | – |
| Orks (1; 27) | 61.7 | 68.5 | 57.4 | 60.6 | 64.7 | 66.3 | – |
| Space Wolves (1; 12) | 50.7 | 85.0 | 52.7 | 87.1 | 87.9 | 74.3 | – |
| Thousand Sons (1; 9) | 38.2 | 66.0 | 44.1 | 66.0 | 71.5 | 56.3 | – |
| World Eaters (3; 16) | 67.5 | 83.2 | 58.2 | 79.3 | 83.8 | 79.4 | 89.3; 82–85 |
| Pooled, 15 | 59.1 | 72.2 | 54.8 | 65.8 | 73.0 | 69.3 | |
| Tyranids (3; 27; fitted on, not pooled) | 82.4 | 89.1 | 61.3 | 85.5 | 89.2 | 89.0 | 92.7; 92.7 |
"Ceiling (ours)": each list against the mean letter of the army's other lists (tools/progress.js ceilingOf, ties in the others left out), where an army has two 11e lists. "Known": the brief's figures (World Eaters' 82–85 is T-672's one-to-one and two-against-one count with ties at half).
Toward the ceiling (the share of the gap from Track A to the reviewers' ceiling that popularity at p 2 closes): Space Marines 65% (of 30.2 points), World Eaters 72% (of 21.7; 15.7 points, past the known 82–85), Genestealer Cults 49% (of 28.5), Tyranids 65% (of 10.2).
Read with care. Each army is one to three lists and 18 to 84 units, so a single army's interval is wide (Blood Angels −0.8 [−27, +27]). The pooled numbers are solid. Agents of the Imperium have one tournament list, so their gain is luck (their scramble gains more). Blood Angels and Death Guard barely move. Votann's Auspex list bands units partly by how much they're played (T-691's caution), so its +25.7 is partly the same signal twice.
T-702: a free fit on World Eaters
Method. tools/wefit.js: rung (c) of T-662 on World Eaters instead of Tyranids, on Track A's model with every knob free, 136 in all:
- the six line weights against Kill's;
- the melee lands scale and the arrival reach;
- the rule kinds on 3+ units;
- the fast dials;
- a slope for each of the 116 datasheet features that vary among the World Eaters (tools/ledgerprofile.js select, near-duplicates folded).
It is searched with tools/crack2core.js search() (surrogate, Latin hypercube, CMA-ES, coordinate ascent; 150,000 evaluations a search, half for the held-out runs). The score is pairwise on letters S to D over the three 11e lists (Red Path, Tactical Sugar, Exalted), ties wrong: the exam's number. The scrambled control moves each unit's row of three letters whole (5 seeds, the same budget).
| World Eaters (136 knobs, 30 graded units) | Track A | Free fit | Gain |
|---|---|---|---|
| Real letters, in sample | 67.5 | 94.3 | +26.7 |
| Scrambled letters, in sample (mean of 5) | 50.1 | 94.3 (94.3–94.4) | +44.2 |
| Fit on two lists, scored on Red Path | 62.7 | 83.8 | +21.1 |
| Fit on two lists, scored on Tactical Sugar | 72.5 | 90.9 | +18.4 |
| Fit on two lists, scored on Exalted | 67.4 | 81.7 | +14.3 |
| Held-out list, mean of 3 | 67.5 | 85.5 | +17.9 |
| Held-out units (5 folds of 6), mean | 67.3 | 64.4 | −2.9 |
| Reviewers: each list against the other two (ours) | 89.3 (87.3, 94.8, 85.8) | ||
| Reviewers: one against another (ours; T-672's ties-half count 82.4, two against one 85.2) | 88.6 |
Read-out. The fit reaches 94.3%, above the reviewers' own agreement. Shuffled letters reach exactly the same, so all of it is memorising. The held-out-list score (85.5%) looks good but isn't evidence. The three reviewers grade the same 30 units much alike, so a fit that memorises the units' order from two lists carries it to the third. Even so, the other two lists' plain consensus scores the third better (89.3% against 85.5%). The real test is units the fit never saw, and there it falls below Track A (64.4% against 67.3%; the folds run 54.7 to 81.9). The same pattern as on Tyranids (T-662: 92.9% real, 91.5% shuffled, held-out units +0.9).
Files
- tools/popshare.js (new): the post hook (
modemult, pop, poptie, blend;p;seedfor the scramble), the measurement and the held-out analysis; output build/popshare.json (git-ignored). - tools/wefit.js (new): the free fit, scrambles, held-out lists and units; output build/wefit.json (git-ignored).
- Read only: tools/popularity.js's reports/tiers/2026-10-05-unit-popularity.json (grimstat-corpus, CC BY 4.0, counts only); the frozen ledger (inputs hash a20dbd2c3316). No ledger, exam, effects or other builder's tool changed.
Run: node --max-old-space-size=2000 tools/popshare.js (about 10 s); node tools/exam.js tools/ledger-benchmarks/2026-10-05-step34b-baseline.json --post tools/popshare.js --post-args mode=mult,p=2 --md; node --max-old-space-size=2000 tools/wefit.js (about 4 minutes, one thread).
Does the 94% carry to other armies?
5 Oct, on the c5e7d85b43f8 snapshot (pinned to d8b5a1bb14). A report: no ledger value, setting or letter moves.
Method. tools/wecarry.js runs the World Eaters free fit again exactly as tools/wefit.js does (the same 136 knobs, search, seeds and budget; it reaches 94.3% again on this snapshot) and the five scrambled-letter fits beside it, and saves every fit's knob values (reports/tiers/2026-10-05-wecarry.json). Each fit is then applied as it is, with no re-fitting, to every other army, scored on that army's own 11th-edition lists (letters S to D, ties wrong: the exam's number; every Track A cell matches tools/exam.js on the step-34b baseline). Each army's model carries the same 116 datasheet features, standardised within the army as on World Eaters, so a slope means the same per SD everywhere; the weights, lands, reach and fast dials carry as they are, and a rule kind's multiplier carries where the army has the kind. Controls: the five scrambled fits carried the same way, and the real fit with its values permuted among its knobs (each value's place in its own range moved to another knob; five seeds).
| Army (11e lists, graded units) | Track A | Real fit carried | Scrambled fits carried, mean (range) | Permuted fit, mean (range) |
|---|---|---|---|---|
| World Eaters (3, 30; fitted on, in sample) | 67.5 | 94.3 (+26.7) | 49.9 (36.1–60.4) | 63.2 (57.0–65.0) |
| Tyranids (3, 52) | 82.4 | 58.4 (−24.1) | 56.7 (41.5–65.3) | 58.4 (50.8–64.3) |
| Space Marines (2, 84) | 60.3 | 51.6 (−8.7) | 51.6 (48.1–55.0) | 52.9 (44.6–58.3) |
| Adeptus Mechanicus (1, 29) | 71.7 | 62.7 (−9.0) | 52.4 (36.6–67.4) | 51.0 (40.1–64.9) |
| Agents of the Imperium (1, 26) | 56.5 | 51.1 (−5.4) | 50.9 (40.4–58.7) | 50.8 (42.2–60.1) |
| Astra Militarum (1, 61) | 66.6 | 55.7 (−10.9) | 54.8 (46.6–65.7) | 50.1 (36.4–66.9) |
| Blood Angels (1, 19) | 61.3 | 34.7 (−26.6) | 56.3 (46.8–62.1) | 47.4 (33.1–59.7) |
| Dark Angels (1, 77) | 53.9 | 49.5 (−4.4) | 50.3 (46.5–55.7) | 51.2 (42.1–58.5) |
| Death Guard (1, 36) | 54.6 | 51.1 (−3.6) | 48.1 (43.5–51.9) | 50.9 (42.2–56.9) |
| Drukhari (1, 23) | 58.0 | 33.8 (−24.2) | 52.6 (45.9–66.2) | 41.4 (30.9–53.6) |
| Emperor's Children (1, 23) | 69.6 | 55.5 (−14.1) | 42.3 (31.9–58.1) | 44.3 (30.4–52.4) |
| Genestealer Cults (2, 24) | 56.4 | 59.7 (+3.3) | 55.6 (51.5–61.9) | 55.5 (39.6–68.6) |
| Leagues of Votann (1, 18) | 61.0 | 42.9 (−18.1) | 60.6 (43.8–82.9) | 39.6 (25.7–48.6) |
| Orks (1, 45) | 61.7 | 49.2 (−12.5) | 53.7 (30.4–63.0) | 58.5 (49.2–73.4) |
| Space Wolves (1, 20) | 50.7 | 45.7 (−5.0) | 54.7 (48.6–60.0) | 49.3 (35.0–57.9) |
| Thousand Sons (1, 28) | 38.2 | 48.5 (+10.4) | 52.0 (47.2–64.7) | 46.7 (35.6–60.8) |
| The 14 other pool armies, mean | 58.6 | 49.4 (−9.2) | 52.6 (49.8–55.6) | 49.3 (42.1–54.5) |
| The 15-army pool, mean (World Eaters in sample) | 59.2 | 52.4 (−6.8) | 52.4 (49.7–55.3) | 50.2 (43.1–55.2) |
In the two mean rows, the scrambled and permuted ranges are the range of the five fits' own means over the armies.
Read-out. It doesn't carry. On the 14 other armies the real fit scores 49.4%: coin-flip level, 9.2 points below Track A, below each of the five scrambled fits' means (49.8 to 55.6), and level with the same values shuffled among the knobs (49.3). It beats Track A on 2 of the 14 (Genestealer Cults +3.3; Thousand Sons +10.4, the army where Track A is below chance) and beats every scrambled fit on none. On the Tyranids it falls from 82.4% to 58.4%, the same as the scrambled fits (56.7) and the permuted one (58.4). Blood Angels and Drukhari fall to about 34%: the fit orders their units the wrong way round more often than not. So the 94.3% learnt nothing that transfers. It is a setting of 136 knobs that sorts 30 World Eaters units, as T-702's held-out units already showed, and is no more use elsewhere than a fit to shuffled letters.
Run: node --max-old-space-size=2500 tools/wecarry.js --md (about 2 minutes, one thread; --from reports/tiers/2026-10-05-wecarry.json reuses the saved fits and skips the searches).