Jordan's question: "score units based on roles where certain stats are ignored. Is there a statistical method for this? And can we just look at the aggregate data and see what roles fall out naturally? Or are we back to the 4 we had tried ages ago?" Scope (Jordan, later on 3 Oct): Tyranids only for now; an all-armies run made before that call is a side note at the end. Command: node tools/roles.js (16 s; --all for the 16 armies, 5 min). Numbers in reports/tiers/2026-10-03-roles.json (each Tyranid unit's memberships) and 2026-10-03-roles-all-armies.json. A report: nothing is adopted and no setting, engine or data file changed. Our own numbers; unit names only.
The short answer
- Yes, there are methods for both. Two standard tools find the roles in the data: factor analysis (how many separate ways units really differ) and archetypal analysis (a few extreme types, with every unit a mix of them). "Score each role, ignoring some stats" is a mixture of regressions. Each role gets its own weights. A weight may fall to 0, which is the role ignoring that stat. A unit's Value is its roles' scores, mixed by how much of each role it is.
- Three roles fall out of the 52 Tyranid units. No reviewer letter was used to find them.
- Big and tough: monsters, many wounds, Kill into vehicles, monsters and heavy infantry, few bodies per point. The Swarmlord, Hive Tyrant, Norn Assimilator, Hierophant, Exocrine, Tyrannofex.
- Bodies on objectives: many models, OC and Hold, Kill into chaff and elite infantry. Termagants, Hormagaunts, Von Ryan's Leapers, Ranged Warriors, Neurogaunts, Genestealers, Gargoyles.
- Cheap action pieces: very cheap, many bodies per point for actions, almost no Kill or Soak. Mucolid Spores, Spore Mines, half of Ripper Swarms.
- Partly the old four, not all of them. Bodies on objectives is part Anvil (holding) and part Runner. Cheap action pieces are Runners. Big and tough is part Anvil and part anti-tank Hammer. Damage (Hammer) is not a role of its own. Kill per point varies across every role, so it measures how good a unit is, not what kind of unit it is. Support (Banner) can't be seen at all, because no solo line measures what a unit gives its neighbours. The Tyranid characters scatter across the first two roles (the Broodlord 55% big, the Winged Prime 51% bodies).
- Scoring each role by its own weights: a small held-out gain that noise matches. Under leave one list out over the five Tyranid lists, role weights gain about +0.02 in ρ at the weaker ridges (4 of 5 lists better). On shuffled letters the same fit "gains" +0.01 to +0.03 (SD 0.03 to 0.04). Neither fitted ledger beats the step-6 baseline (0.565), though that baseline was tuned on these same lists. Not a finding yet; don't adopt. The roles are worth keeping as a lens on the page (which kind of unit is this?), not as a scoring change.
- What each role ignores (fitted on all five lists; floors at 0): big and tough ignores Presence (0, against 0.7 for the single ledger), so bodies on the table don't count for a monster. Cheap action pieces ignore Kill (0, against 76). Bodies on objectives keep Kill but take less for OC (Score 5.2, against 7.2). At tools/ledgerfit.js's default ridge (λ 0.01) nothing reaches 0; the roles' Kill weights only spread (big 365, bodies 312, spores 226, against 368).
The roles, unit by unit (Tyranids)
Memberships sum to 1; the main role and its share in %. Archetypal analysis, k = 3 (the knee of the fit curve, also the most stable k; Jordan's budget allows 2 to 3).
| Role | Units (main role, share) |
|---|---|
| Big and tough (27) | The Swarmlord 100, Hive Tyrant 100, Norn Assimilator 100, Hierophant 100, Harridan 100, Old One Eye 95, Winged Hive Tyrant 94, Maleceptor 93, Haruspex 92, Hive Crone 91, Exocrine 90, Tyrannofex 88, Tervigon 85, Norn Emissary 82, Trygon 82, The Red Terror 82, Tyrannocyte 81, Harpy 77, Screamer-killer 76, Toxicrene 67, Sporocyst 66, Psychophage 66, Carnifexes 63, Deathleaper 55, Mawloc 55, Broodlord 55, Neurotyrant 53 |
| Bodies on objectives (22) | Von Ryan's Leapers 100, Hormagaunts 100, Termagants 100, Ranged Warriors 100, Neurogaunts 96, Genestealers 93, Gargoyles 91, Barbgaunts 90, Venomthropes 87, Melee Warriors 87, Pyrovores 77, Tyrant Guard 71, Neurolictor 68, Raveners 63, Parasite of Mortrex 63, Biovores 62, Hive Guard 59, Hyperadapted Raveners 54, Winged Tyranid Prime 51, Zoanthropes 51, Tyranid Prime with Lash Whip 50, Lictor 48 |
| Cheap action pieces (3) | Mucolid Spores 100, Spore Mines 100, Ripper Swarms 51 |
Mixed units (no role above 60%): the Lictor, Zoanthropes, Deathleaper, Hyperadapted Raveners, both Primes, the Broodlord, Hive Guard, Ripper Swarms, the Neurotyrant and the Mawloc. These are mostly the characters and mid-sized specialists the reviewers argue about.
Each role's profile (z-scores within Tyranids):
| Role | High | Low | Old jobs' proxies (correlation with membership) |
|---|---|---|---|
| Big and tough | monster +1.03, wounds per model +0.99, Kill into light vehicle +0.75, heavy vehicle +0.63, heavy infantry +0.63 | infantry −0.92, Presence −0.83, Actions −0.81, Kill into chaff −0.72 | Hammer +0.21, Anvil +0.12, Banner +0.25, Runner −0.56 |
| Bodies on objectives | Kill into elite infantry +1.24, OC +1.22, Kill into chaff +1.21, model count +1.14, Hold +0.95 | wounds per model −1.13, monster −0.95, Kill into light vehicle −0.64 | Hammer +0.30, Anvil +0.33, Banner −0.21, Runner +0.36 |
| Cheap action pieces | mounted or beasts +3.31, Actions +2.16, Presence +2.16, fly +1.41 | Kill into character −3.65, Soak −2.36, Kill into elite infantry −2.04 | Hammer −0.86, Anvil −0.75, Banner −0.08, Runner +0.37 |
The proxies: Hammer is the log of total Kill per point, Anvil is Soak with Hold, Runner is Actions, Presence and Move, and Banner is the character flag (no solo line measures support).
How many dimensions matter: five, by parallel analysis (eigenvalues 6.1, 5.2, 3.0, 2.5, 2.2 against shuffled-data thresholds 2.9, 2.6, 2.3, 2.1, 1.9; Kaiser's rule says 8). After rotation:
- small and cheap (Actions, Presence, beasts) against big and dangerous (Kill into characters, cavalry and heavy infantry, Soak, wounds);
- objective bodies (OC, model count, Hold, Kill into chaff);
- speed (charge lands, Move, melee share, arrival);
- anti-armour Kill (into monsters and vehicles);
- transport (the Carry line).
The three roles span the first two dimensions well. Speed, anti-armour and transport cut across them: inside each role they separate the better units from the worse. The three roles explain 38.5% of the table's variance (two roles 21.5%, four 48.1%, five 55.7%). Bootstrap stability is 0.874 (mean matched-profile correlation over 20 resamples), and 89% of units keep their main role. A Gaussian mixture on the five components prefers 5 clusters by BIC and agrees only weakly with the roles (adjusted Rand 0.27; 0.37 at 3 clusters). The units sit on a continuum rather than in natural clumps.
Each role scored by its own weights: leave one list out (Tyranids)
Fit on four lists by the pairwise likelihood (tools/ledgerfit.js's, reliability-weighted). Score the list left out. The single-weight ledger is fitted the same way on the same folds, with the same ridge (toward the step-6 weights). The role ledger's ridge pulls toward that fold's single fit. Each role frees the 3 lines its profile is furthest from 0 on (from Part 1, never from letters): big and tough Kill, Actions and Presence; bodies on objectives Kill, Score and Hold; cheap action pieces Kill, Soak and Actions.
The budget: 5 lists, 3,946 cross-tier pairs over 226 graded entries of 52 units. The single ledger fits 7 weights and 5 list scales. The role ledger adds 9 role weights, 21 parameters in all, about 188 pairs per parameter. The pairs share their units, so the honest count is nearer 2.5 units per parameter.
| Ridge λ | Single ρ | Roles ρ | Δρ (SE over 5 folds) | Lists better | Roles, all 7 lines free | Single pairwise | Roles pairwise | Scrambled letters: Δρ mean, SD (3 shuffles) |
|---|---|---|---|---|---|---|---|---|
| 0.01 | 0.546 | 0.546 | −0.000 (0.006) | 3 of 5 | 0.547 | 74.3% | 74.1% | +0.025, 0.023 |
| 0.001 | 0.539 | 0.560 | +0.021 (0.008) | 4 of 5 | 0.563 | 74.2% | 75.0% | +0.011, 0.032 |
| 0.0001 | 0.540 | 0.563 | +0.023 (0.009) | 4 of 5 | 0.557 | 74.2% | 75.2% | +0.011, 0.042 |
References: the step-6 weights, unfitted, give ρ 0.565 (pairwise 75.1%), but they were tuned on these five lists. Every weight at 1 gives ρ 0.347.
| List held out | Graded | λ 0.01: single → roles | λ 0.001: single → roles | λ 0.0001: single → roles | Step-6 |
|---|---|---|---|---|---|
| Auspex | 44 | 0.640 → 0.621 | 0.645 → 0.669 | 0.652 → 0.674 | 0.662 |
| Hivemind | 45 | 0.604 → 0.595 | 0.604 → 0.597 | 0.605 → 0.593 | 0.622 |
| Second | 41 | 0.556 → 0.560 | 0.529 → 0.551 | 0.525 → 0.557 | 0.587 |
| Maelstrom | 48 | 0.552 → 0.569 | 0.544 → 0.585 | 0.541 → 0.585 | 0.571 |
| Astrategas | 48 | 0.378 → 0.385 | 0.374 → 0.400 | 0.378 → 0.404 | 0.381 |
Reading it.
- At the default ridge the roles change nothing.
- At the weaker ridges they gain about 0.02. Part of that only recovers what the single ledger loses when it is let off its prior (0.546 → 0.539): the role ledger rebuilds most of the step-6 shape from three role weight sets.
- The gain is about the size of what the same fit finds in shuffled letters, so it does not clear noise.
- The keep rule (beat the single weights held out, with no big loss) is met on its face at λ 0.001 and 0.0001 (Hivemind −0.007 and −0.012), but not beyond the scrambled spread. It is also not against step 6, the ledger in use.
The weights fitted on all five lists (bold: freed for that role; the rest held at the single fit):
| λ | Ledger | Kill | Soak | Score | Actions | Hold | Presence | Spawn |
|---|---|---|---|---|---|---|---|---|
| 0.01 | Single | 368 | 16.2 | 10.1 | 1.30 | 19.2 | 2.74 | 14.8 |
| Big and tough | 365 | 16.2 | 10.1 | 1.30 | 19.2 | 1.95 | 14.8 | |
| Bodies on objectives | 312 | 16.2 | 10.4 | 1.30 | 18.4 | 2.74 | 14.8 | |
| Cheap action pieces | 226 | 16.2 | 10.1 | 1.29 | 19.2 | 2.74 | 14.8 | |
| 0.001 | Single | 75.9 | 22.5 | 7.23 | 0.78 | 0 | 0.70 | 0 |
| Big and tough | 61.8 | 22.5 | 7.23 | 0.87 | 0 | 0 | 0 | |
| Bodies on objectives | 57.5 | 22.5 | 5.18 | 0.78 | 0 | 0.70 | 0 | |
| Cheap action pieces | 0 | 16.5 | 7.23 | 0.43 | 0 | 0.70 | 0 |
With a weak ridge the single ledger itself drops Hold and Spawn and halves Kill against Soak (the weights' overall size is set only by the ridge; read them against each other). With every line free, the roles' weights wander further (in the JSON) for no held-out gain.
Side note: all 16 armies (done before the scope change)
Before Jordan narrowed the scope, the full brief had been run on all 16 armies (733 units, 26 lists). It had a smoke test and its own results; node tools/roles.js --all reproduces it (312 s). In short:
- Eight dimensions by parallel analysis (Kaiser 9): infantry against vehicles, objective bodies, anti-armour Kill, speed, actions and presence, anti-elite Kill, monster melee, transport.
- Four roles, chosen as the knee (45% of variance; bootstrap stability 0.84, the highest of k = 2 to 7):
- big and tough (265 units: Shadowsword, Land Raider, Death Company Dreadnought, Defiler);
- characters (224: Commissar, Painboy, Chaplain, Master of Executions);
- bodies on objectives (219: Boyz, Wyches, Catachan Jungle Fighters, Scout Squad);
- not a fighter (25: Drop Pods, the Aegis Defence Line, aircraft, spores).
Characters are a role across armies but not within Tyranids, whose characters are mostly monsters. Damage is again not a role.
- Leave one army out, mean of 16: single ρ 0.174 against roles 0.170 (Δ −0.004, SE 0.011, 7 of 16 better) at λ 0.01. At λ 0.001 it was −0.004 and at λ 0.0001 −0.003. Scrambled letters gave Δ −0.001 to +0.003 (SD up to 0.035).
- Notable armies at λ 0.01: Tyranids −0.019, Space Marines +0.028, Dark Angels +0.053, Emperor's Children +0.064, Death Guard −0.098, World Eaters −0.053, Leagues of Votann −0.058.
- Verdict: role weights did not transfer across armies.
- Changes made for the Tyranid run: the k rule (below) and the names. The first full run's own k rule, "step down while stability is under 0.85", fell to k = 2, the least stable k, because every k sat under 0.85 on 20 bootstraps. It was replaced by "the knee unless a smaller k is clearly more stable" after seeing that run. The names were refined after seeing the profiles.
Method
- Features (tools/roles.js; every unit at its best solo form at tools/ledger-benchmarks/2026-10-03-step6-baseline.json, compute()'s ranking row):
- Kill into each of the 8 target classes at β 6 (
kill["6"].byClass) and the melee share of Kill; - Soak (raw points saved), OC, Actions, Hold, Presence (own bodies plus spawned units, without the placed-anywhere dial), Spawn and Carry;
- Move (the datasheet's M), the charge model's lands, arrival (0 or 1), wounds per model and model count;
- flags from the datasheet keywords: character, vehicle, monster, infantry, mounted or beasts, transport, fly.
- Kill into each of the 8 target classes at β 6 (
Rates and counts take log(1 + x). Every feature is then standardised within the army.
- PCA on the correlation matrix (Jacobi eigen-decomposition), Horn's parallel analysis (100 column shuffles, 95th percentile), and varimax on the kept components.
- Archetypal analysis (Cutler and Breiman) for k = 2 to 5 (2 to 7 for all armies): alternating projected FISTA on the simplex, the furthest-sum start, best of 4 starts. Stability comes from 20 bootstrap resamples per k, with archetypes matched by the best permutation of profile correlations. k is the knee of the fit curve unless a smaller k is more stable by 0.05, capped at 3 for Tyranids.
- Names come from a fixed rule list over each profile (tools/roles.js
nameOf). They read the profile only, never a letter. - Cross-check: a diagonal Gaussian mixture on the kept components' scores (k-means++ start, EM, best of 4), compared with the archetypes by BIC and the adjusted Rand index. On the raw table the binary flags make a diagonal mixture degenerate.
- Role ledger:
- Value_u = max over the unit's solo forms f of (Σ_r m_ur w_r) · b_f, where b_f is the form's seven scaled lines from valueForm() with every weight at 1 (the threat weighting at 0, every rule multiplier as step 6).
- With one role this is the single ledger. Its Values match tools/ledgerfit.js's valuesAt() at the step-6 weights to 1e-9, checked on every run.
- Fitted with ledgerfit's nll() and bfgs(): weights w = size · x², floors at 0, a log scale per list, lists weighted by reliability. The ridge λ · ((w − prior) ÷ size)², where size is each line's step-6 weight, pulls toward the step-6 weights (single) or that fold's single fit (roles).
- λ in {0.01, 0.001, 0.0001}, reported in full rather than chosen on the held-out lists.
- The memberships m are fixed from Part 1.
- Judging: leave one list out over the five Tyranid lists (one army out over 16 for the side run). Scrambled letters run the same folds on each list's letters shuffled among the units it grades (3 shuffles).
- Caveats:
- The five lists grade the same 52 units, so leave one list out tests agreement with a new reviewer, not with new units.
- The step-6 weights were set on these lists, so 0.565 is in-sample.
- Two lists are 10th edition (Maelstrom, Astrategas).
- Banner can't be measured on solo forms.
- Run time: Tyranids 16 s (archetypes and bootstraps 10.6 s, Part 2's 3 ridges × (fits, 5 folds, 3 × 5 scrambled folds) 5.2 s). All 16 armies 312 s (archetypes 144 s on 6 threads, Part 2 162 s).
- Tests: tests/unit/ledgerroles.test.js (archetypes find a triangle's corners; two alike roles are one; the fit recovers known weights on a toy army).