5 Oct 2026. A measurement, nothing adopted. Snapshot d8b5a1bb14 (hash c5e7d85b43f8), Track A (step 34b) at wCarry 0. Tool: tools/leadercomp.js (an exam post step). Numbers: 2026-10-05-leadercomp.json.
In plain English
Jordan's idea, from his Votann notes: a leader is worth what it adds to the squads worth leading, over the best other leader for that squad. Everyone puts the Kâhl on Hearthguard, so the Grimnyr, the Iron-master and the Champion have nothing worth leading. The same goes for Space Marines' Phobos and Reiver Captains and Lieutenants.
We built that as one game-wide dial and measured it the usual way: chosen on 10 armies, scored on the other 5, beside a placebo. It does not hold out. At every strength the pooled exam falls (−0.6 to −2.0 points). At small strengths it also falls below its own placebo, which deals the same credits to the wrong leaders. The penalty half alone is the least bad, and it lifts Votann +6.7 and World Eaters +3.6. It still loses pooled and held out.
The reason is in the inputs, not in the competition. Our engine measures what each leader adds to each squad, and from those numbers it usually crowns a different winner than the reviewers do:
- Votann: the engine's best leader for Hearthguard is Ûthar (2,730 a hundred points), not the Kâhl (1,594).
- Votann: the Iron-master wins both Thunderkyn and Warriors.
- Space Marines: the Apothecary wins ten squads. Its healing is measured large: up to 6,113 a hundred points on Hellblasters.
Across 129 squad-and-list cases where the reviewers' letters separate the candidate leaders, the engine's winner holds the best letter 30% of the time. Picking at random would do it 37% of the time. A competition judged on those numbers moves the wrong leaders up.
What came out right: the Grimnyr, the Champion, the Judiciar and the Lord on Juggernaut all fall, toward the reviewers. What came out wrong: the Captain in Phobos Armour rises from 75th to 8th. Only Phobos characters can lead Phobos squads, so it wins its own small contest. The Apothecary rises from 78th to 1st.
How it works
Read on compute()'s run for each army:
- The pairings. Each joined form is a squad led by one leader. No form in the pools has two leaders. Its synergy is the form's total minus the squad alone minus the leader alone, per 100 of the leader's points. "Alone" is ledger.js partnerAlone(): the same model count in the form's detachment, or else the member's best solo form. A pairing takes its best form.
- Squads worth leading. A squad counts when its best pairing adds something. Its weight is its standing in the army: the percentile of its own ranked Value, from 1 for the army's best to 0 for its worst.
- Credit. A leader gets Σ over the squads of weight × (its synergy − the best other leader's synergy, floored at 0). When no other leader adds anything, the squad would go unled.
- Penalty. A leader that is best on no squad worth leading takes weight × its gap to the best at its nearest miss. The sub-dial
leaderCompPenfollowsleaderCompunless it is set on its own. - The dial. Value′ = Value + leaderComp × credit − leaderCompPen × penalty. At 0 and 0 the run comes back untouched. Synergy per 100 of the leader's points is on Value's own scale: Q points of worth at the median rate P comes to Q × P ÷ points, the price the character credits use.
- Placebo. Each leader's net adjustment is dealt to the army's leaders in a seeded random order (seeds 11 to 13).
Results
All figures are against Track A: exam pooled 59.37 (15 armies, 11th-edition lists), consensus pooled 60.08, Tyranids' three 11th-edition lists 82.45. Held out means 3 splits, choosing on 10 armies and scoring on 5. "Consensus" scores armies with two or more 11th-edition lists on their consensus pairs and the rest as the exam does. Each placebo sits in brackets beside its real value.
| Variant | Strength | Held out ± spread (placebo) | Exam pooled (placebo) | Consensus pooled (placebo) | Tyranids (placebo) | Armies better / worse |
|---|---|---|---|---|---|---|
| whole idea | 0 | +0.00 ± 0.00 (−1.17); every fold picks 0 | 0 | 0 | 0 | – |
| whole idea | 0.25 | −0.62 (−0.43) | −0.69 (−0.41) | −0.40 (−0.27) | 6 / 9 | |
| whole idea | 0.5 | −0.98 (−1.28) | −0.98 (−1.25) | −0.74 (−0.50) | 5 / 10 | |
| whole idea | 1 | −2.03 (−1.93) | −2.11 (−1.96) | −1.21 (−0.98) | 6 / 9 | |
| credit alone | 0.25 / 0.5 / 1 | −0.98 ± 1.70 (−0.42) | −0.66 / −0.78 / −1.48 (−0.24 / −0.69 / −0.79) | −0.73 / −0.87 / −1.61 | −0.40 / −0.74 / −1.21 | 4/10, 4/9, 3/12 |
| penalty alone | 0.25 / 0.5 / 1 | −0.36 ± 0.63 (−1.04) | −0.23 / −0.41 / −0.87 (−0.16 / −0.56 / −1.20) | −0.15 / −0.29 / −0.81 | 0 / 0 / 0 | 5/8, 7/7, 7/8 |
These variants were not asked for; they are reported because the asked-for ones failed:
- A finer grid (0.02, 0.05, 0.1). Whole idea −0.37 / −0.57 / −0.60 against a placebo of −0.10 / +0.04 / −0.04. Held out it reads 0 for the whole idea, −0.92 for the credit alone and −0.31 for the penalty alone.
- Standing read from the squad's best led form (
--stand led). Held out: whole idea 0, credit −1.13, penalty −0.44. - Winners only (
--shape rank). The sizes of the gains are set aside: each squad a leader clearly wins is worth its weight × the army's median Value. Held out: whole idea 0, credit 0 (−0.36 on the fine grid), penalty −0.59. Pooled at 0.5 it reads −1.41 against a placebo of −0.27.
tools/exam.js gives the same through the hook. At 0 the result is Track A to the bit. At 0.5 the pooled exam is −1.0, with 95% intervals of [−4.0, +1.9] when units are resampled and [−2.5, +0.4] when armies are (200 draws). The Tyranids are at −0.7.
Per army at 0.5 for the whole idea:
- Up: World Eaters +3.9, Thousand Sons +3.6, Space Wolves +2.1, Astra Militarum +2.1, Dark Angels +1.9.
- Down: Drukhari −7.7, Space Marines −5.2, Agents −4.9, Blood Angels −3.2, Adeptus Mechanicus −2.8.
The wall units, before and after (the whole idea at 0.5)
Each unit's rank in its army before → after, with its 11th-edition letters.
| Army | Unit (letters) | Rank before → after | Credit / penalty | Toward the reviewers? |
|---|---|---|---|---|
| Votann (22) | Grimnyr (C) | 17 → 22 | penalty 1,514 (nearest miss: Warriors) | yes |
| Brôkhyr Iron-master (C) | 11 → 2 | credit 3,516 (best on Thunderkyn, Warriors) | no | |
| Einhyr Champion (C) | 8 → 17 | penalty 663 (Hearthguard) | yes | |
| Ûthar the Destined (B) | 16 → 9 | credit 379 (best on Hearthguard) | yes | |
| Kâhl (not graded) | 20 → 20 | penalty 379 (Hearthguard) | – (the engine has Ûthar, not the Kâhl, as Hearthguard's best) | |
| Einhyr Hearthguard (A) | 15 → 15 | a squad: the dial moves leaders only | – | |
| Space Marines (84) | Captain in Phobos Armour (BD) | 75 → 8 | credit 1,748 (best on Incursors, Reivers, Eliminators, Scouts) | no |
| Lieutenant in Phobos Armour (CC) | 76 → 65 | credit 420 (best on Infiltrators) | no | |
| Lieutenant in Reiver Armour (CD) | 79 → 74 (Value 202 → 27) | penalty 350 (Reivers) | Value yes, rank no (others fell further) | |
| Apothecary (BD) | 78 → 1 | credit 24,068 (best on ten squads) | no | |
| Judiciar (BD) | 43 → 80 | penalty 1,993 (Tactical Squad) | yes | |
| Tyranids (52) | Old One Eye (BBA) | 13 → 10 | credit 1,148 (Carnifexes) | yes, slightly |
| Neurotyrant (A, –, –) | 29 → 19 | credit 691 (Zoanthropes, Neurogaunts) | yes | |
| Winged Hive Tyrant (ABB) | 3 → 3 | not a leader in our data | – | |
| Broodlord (AAA) | 10 → 8 | credit 1,152 (Genestealers) | yes | |
| World Eaters (30) | Lord on Juggernaut (CBC) | 10 → 25 | penalty 1,049 (Eightbound) | yes |
The prediction, written first, was that Votann's non-Kâhl leaders and Space Marines' Phobos leaders would fall, toward the reviewers. It was half right:
- Grimnyr and the Champion fall.
- The Iron-master rises, and so do both Phobos leaders.
- The Kâhl falls, because the engine prefers Ûthar on Hearthguard.
What it means
- The competition is the right shape for Jordan's notes, but it can only be as good as the synergy numbers it judges on. Those numbers come from leaderfx.json through the engine. On them, the engine picks the reviewers' preferred leader less often than chance.
- The overrated winners show what is wrong:
- The Apothecary's healing is valued far beyond any other leader buff.
- The Kâhl's Hearthguard buff is measured below Ûthar's.
- Restricted keywords (Phobos, Gravis, Terminator) hand a weak leader a contest it can't lose.
- A next try would need two things before it measures anything new:
- an audit of those leader effects (the Kâhl, the Apothecary, the Iron-master);
- a competition among leaders of the same restricted keyword, so that "best on its own squads" counts against the army's best squads, not only its own.
Reproduce
node --max-old-space-size=2500 tools/leadercomp.js --measure --json out.json # the table node tools/leadercomp.js --army leagues-of-votann --w 0.5 # one army's pairings and leaders node --max-old-space-size=2500 tools/leadercomp.js --measure --grid 0.02,0.05,0.1 # the finer grid node --max-old-space-size=2500 tools/leadercomp.js --measure --stand led | --shape rank # the two variants
As an exam setting: post: { module: "tools/leadercomp.js", args: { leaderComp, leaderCompPen, stand, shape, seed } }, or node tools/exam.js <setting> --post tools/leadercomp.js --post-args leaderComp=0.5.