The question. Track A ranks the reviewers' units about 59% right pairwise on the 15 armies it wasn't fitted on. No weighting of its lines does much better, so we looked at which units it gets wrong and what those units have in common. Miss is the reviewers' position minus ours. Both positions are percentiles within the army: the reviewers' from each unit's mean letter, ours from its Value. A positive miss means we underrate the unit.
What we found
- Our ranking is a weak guide to theirs, unit by unit. Our position explains 8% of the variance in the reviewers' position when the fit is held out by army (fitted slope 0.29). The average miss is 27 points, and 149 of the 618 graded units are off by more than 40. A noisy ranking makes the units we place lowest look underrated and the ones we place highest look overrated, whatever they are. So every pattern below is judged by the miss beyond what our position predicts. On the raw miss, "units we rank low" is the top pattern in all 16 armies, and that only restates regression to the mean.
- No single missing ingredient. The best group, psykers, explains 2% of the variance held out. The ten best groups together raise the share explained from 8% to 22%. A lasso on all 340 columns of tools/diagnose.js's table gets 2%.
- The groups that do lean the same way in many armies are about characters and edge cases. We underrate units whose worth reaches other units: psykers, leaders with a squad buff, Epic Heroes, Lone Operatives. We overrate aircraft, overwatch pieces, cheap small units and the older multi-weapon line squads. One check fits the leader finding: at Track A, every one of the 733 units in the 16 armies takes its Value from its solo form. A leader's buff to the squad it joins is never credited to the leader.
- What fixing the top three is worth. Each unit in a group moves by that group's miss beyond our position, and the exam's pairwise is scored again (pooled over 15 armies, Tyranids out):
- From 59.4 to 63.5 in sample. This is the oracle upper bound.
- 63.0 with each group's move learnt on the other armies.
- 58.3 with the search for groups also rerun without the army being scored. It is 61.2 for the top five and 68.4 in sample for the top ten.
- Moving units by the raw miss, as first posed, gives 59.9.
The in-sample gain comes from a few armies: Thousand Sons 37.9 to 62.1 (12 psykers), Drukhari 57.5 to 66.2 and Astra Militarum 66.9 to 74.5. The Tyranids fall from 79.3 to 76.9 and the World Eaters from 67.8 to 64.3. The top-three picks were also unstable when the search was rerun 16 times, each time without one army. Psykers made the top three in 10 of those runs, aircraft in 9, cheap units in 7, Lone Operative and few weapons in 6 each.
Read: the patterns are real and could be fixed, but the most they buy is about 3 to 4 points. The rest of the gap is unit by unit, and the datasheet columns don't capture it. Likely causes: army and detachment fit, competition within the army (a newer datasheet does the same job), and the reviewers' read of the meta. On 13 of the 15 armies the consensus is one reviewer, so their noise is in the miss too. Worth trying first, each as one switch through node tools/exam.js: (1) credit a leader with what its buff adds to its best squad; (2) discount aircraft for their time off the board; (3) credit Lone Operative as protection in Soak. Nothing is adopted here.
| # | Group (units; armies; armies leaning the same way ≥ 5 pts, of those with 2+) | Raw miss | Beyond our position | Our move to fit | Variance explained, held out | Examples (raw miss) | Missing ingredient (our guess) |
|---|---|---|---|---|---|---|---|
| 1 | Psykers (48; 13; 7 of 8) | +17 | +14 | +50 | 2.0% | Navigator (Agents, +90), Librarian (Dark Angels, +81), Sorcerer (Thousand Sons, +77), Librarian (Space Marines, +68), Infernal Master (Thousand Sons, +66) | What a psyker does for the squad it leads and its psychic abilities. The ledger counts its own kill and Soak only. |
| 2 | Aircraft (15; 7; 5 of 6, the Orks the other way) | −11 | −26 | −97 | 1.7% | Voidraven Bomber (−52), Stormhawk Interceptor (−43), Avenger Strike Fighter (−38), Ravenwing Dark Talon (−37) | Board time: aircraft arrive late, can't hold or score and have to keep moving, but the ledger credits a full game's kill. |
| 3 | Cheapest size 70 pts or less (160; 16; 8 of 16) | +4 | −6 | −39 | 1.5% | Locus (−85), Catachan Command Squad (−79), Iron Priest (−76), Lazarus (−74), Tzaangor Enlightened (−73) | Value per point flatters small pieces. Reviewers weigh what a unit does in absolute terms and what it costs in a slot. |
| 4 | Lone Operative (22; 11; 5 of 6) | +16 | +18 | +80 | 1.4% | Lieutenant with Combi-weapon (Dark Angels +92, Space Marines +87), Tech-Priest Enginseer (+61), Inquisitor Kroyle (+56), Wazdakka Gutsmek (+56) | Protection from shooting, which Soak doesn't count, plus cheap action scoring. |
| 5 | Fewest weapons carried, bottom third (202; 16; 9 of 16) | +1 | −5 | −44 | 1.4% | Deff Dread (−80), Foetid Bloat-drone (−78), Firestrike Servo-Turrets (−77), Heldrake (−62) | One-trick units: their kill is one profile at its best target, with no flexibility. |
| 6 | Leaders with a squad buff in leaderfx (94; 16; 10 of 15) | +16 | +9 | +61 | 2.1% | Skitarii Marshal (+80), Wolf Guard Battle Leader (+74), Nazdreg (+73), Sanguinary Priest (+71), Biologus Putrifier (+69) | The buff to the squad it leads. Value is always the solo form, so the leader never gets credit for it. |
| 7 | Kill little of heavy vehicles, bottom third (206; 16; 10 of 16) | +11 | +5 | +34 | 1.4% | Miasmic Malignifier (+83), Venom (+81), Jakhals (+62), Wolf Guard Battle Leader (+74) | We pay too much for anti-tank kill compared with the reviewers, or credit too little to units whose job isn't kill. |
| 8 | Several pistol profiles: older line squads and their officers (51; 13; 8 of 11) | −16 | −10 | −42 | 1.1% | Catachan Command Squad (−79), Iron Priest (−76), Devastator Squad (−73), Hearthkyn Warriors (−73), Death Company Marines (−58) | Competition within the army: a newer datasheet does the job better, but we score each unit on its own. |
| – | Epic Heroes, a check (59; 14; 10 of 14) | +4 | +11 | 1.2% alone | Unique auras and abilities; overlaps with 1, 4 and 6. | ||
| – | Overwatch rules, a check (18; 7; 5 of 6) | −15 | −14 | 0.6% alone | Firestrike Servo-Turrets (−77), Invictor Tactical Warsuit (−40) | Overwatch is valued too highly. | |
| – | Transports / Battleline / Fights First, checks | +15 / −13 / −26 | +1 / +2 / +1 | None beyond our position. Their raw misses are only regression to the mean. |
How it's done: node tools/misses.js --json reports/tiers/2026-10-05-misses.json runs in 16 s. It takes Track A from tools/ledger-benchmarks/2026-10-05-step34b-baseline.json at wCarry 0, through compute() as the exam does, and reproduces tools/exam.js's column A on all 16 armies. It counts 618 graded units, the exam's 597 plus units graded F only. The candidate groups are 1,250 one-condition splits of the table: each flag, each column's top or bottom third within its army, and its pooled quartiles. The model is the reviewers' position on ours plus group indicators. Groups are picked one at a time by held-out squared error, leaving each army out in turn, and a group has to lean the same way in at least 4 armies. The .json holds every unit's miss, the ten groups with per-army means and runners-up, the best 40 single groups, the lasso and the oracle per army.