Jordan's challenge: "You have every variable available to you based on the datasheets. If you had NOTHING ELSE there should be a mathematical approach to solve this mystery. Either we are not attacking this with the correct statistical approach, or we have the right approach but not enough variables isolated, or we have not tried enough rigor." Command: node tools/diagnose.js (5 minutes; --quick for a smoke run). Every number below is in reports/tiers/2026-10-03-diagnosis-tables.md and -diagnosis.json; the table of variables is 2026-10-03-tyranid-features.csv with its dictionary in -tyranid-features.md. A report: no ledger, page, data file, setting or input lock changed. Our own numbers; unit names only.
The short answer
- Mostly the third: not enough data for the rigour the question needs. It is partly the second as well: some of what the reviewers see isn't on a datasheet. It is not the first. A better statistical method doesn't beat the ledger on units it hasn't seen.
- On the 52 units they were fitted on, the flexible models reach the reviewers easily. Against the five lists they score 88 to 92%, the same as the reviewers' 88.5% agreement with each other. The ledger's own form can't get there however its seven weights are set: 74% at most. That looks like "the ledger's form is the problem", but it isn't the test that matters. With 144 variables and 52 units, a flexible model can fit almost anything, including noise.
- On units held out, every model lands at 64 to 74% against the lists. The ceiling is 88.5%. The best flexible model, linear with every pairwise interaction, gets 73.3% against the consensus. The ledger as it is gets 72.9%. The gap between them is well inside noise. Boosted trees do worse (68%). So do plain linear models (66 to 70%).
- In plain words: with only 52 Tyranid units, nothing learns the reviewers' taste from the datasheets well enough to carry it to a unit it hasn't seen. The flexible models memorise. On new units they do no better than the ledger. The learning curve rises slowly with more training units, so more graded units, which means pooling armies, is the route. A cleverer fit on Tyranids alone is not.
- The units every model misses are the same handful, which says something is missing from the variables, not just from the fit:
- Rated much higher than any model gives them: Biovores, Tyrannofex, Exocrine, Lictor, Neurolictor, Hormagaunts.
- Rated much lower: Toxicrene, Hierophant, Tervigon, Tyrannocyte, Ranged Warriors, Deathleaper.
The numbers
The target: each unit's consensus, its mean letter over the five headline Tyranid lists that grade it (S 5 … F 0; every unit is graded by 2 to 5 lists).
How it is measured:
- Pairwise: of the pairs the consensus orders, the share the model orders the same way (tools/ledgerjoint.js shape()'s count; a tie in the model's scores counts as wrong). Against each list, it is shape() exactly: tiers S to D. The tool checks this against the page's own numbers.
- The ceiling: each list against the mean letter of the other four, averaged: 88.5% (91.7 / 87.7 / 81.9 / 94.9 / 86.5).
- "Held out": 10 repeats × 5-fold over units. Every setting (λ, ridge or lasso, depth and trees, k) is picked by a 4-fold cross-validation inside the training units. Only pairs of units that were both left out of the same fit are scored, so no fit ever saw either unit. A pooled variant, where each fold's scores are put on the consensus scale, sits beside it and agrees.
- Scrambled: the consensus and letters shuffled among the units, 5 shuffles × 2 repeats: the noise floor.
| Model | In-sample vs consensus (tuned / most flexible) | In-sample vs lists | Held out vs consensus (± SE over repeats) | Held out vs lists (± SE) | Scrambled vs consensus (SD) |
|---|---|---|---|---|---|
| Ceiling: each list against the other four | – | 88.5% | – | 88.5% | – |
| (a) the ledger as it is (step 6) | 73.3% / 73.3% | 74.5% | 72.9% ± 0.4 | 74.1% ± 0.6 | 50.8% (4.5) |
| (a1) the ledger's form, its 7 weights re-fitted in the folds | 73.3% / 73.8% | 74.5% | 69.7% ± 1.0 | 70.1% ± 1.0 | 49.1% (3.8) |
| (b) linear on all 144 variables, ridge or lasso | 89.1% / 98.8% | 87.7% | 69.7% ± 1.3 | 68.6% ± 0.9 | 45.9% (6.9) |
| (b2) linear on a short core list (25), ridge | 72.2% / 81.7% | 71.0% | 66.0% ± 0.9 | 64.2% ± 0.8 | 46.5% (8.4) |
| (c) linear + every pairwise interaction (kernel ridge) | 97.2% / 97.2% | 92.2% | 73.3% ± 0.9 | 71.4% ± 1.0 | 44.8% (6.6) |
| (d) boosted trees, depth 1 to 3, monotone constraints | 95.3% / 99.7% | 91.5% | 68.3% ± 0.9 | 67.6% ± 0.8 | 49.3% (3.8) |
| (e) nearest neighbours (similar units) | 100% / 100% | 92.3% | 71.9% ± 0.8 | 70.1% ± 1.1 | 45.2% (7.5) |
Caveat on (a): the step-6 ledger was tuned on these same 52 units and lists, and carries Jordan's hand dial on the Biovores. Its "held out" number is really in-sample, so it is optimistic. (a1) is the fair version of the ledger: its weights are fitted only on the training units. The SEs over repeats measure only how the folds were split. The honest noise is the bootstrap over units and the scrambled spread below.
Held out, per list:
| Auspex | Hivemind | Second | Maelstrom | Astrategas | |
|---|---|---|---|---|---|
| ceiling | 91.7% | 87.7% | 81.9% | 94.9% | 86.5% |
| (a) ledger | 79.9% | 78.4% | 76.2% | 70.8% | 65.1% |
| (a1) ledger re-fitted | 76.7% | 74.3% | 71.9% | 66.4% | 61.3% |
| (c) interactions | 77.9% | 71.7% | 71.2% | 69.9% | 66.5% |
| (d) trees | 76.0% | 69.0% | 72.5% | 61.4% | 59.1% |
| (e) neighbours | 76.6% | 71.2% | 69.1% | 69.4% | 64.2% |
The learning curve: trained on m random units with tuning inside, scored on the pairs among the rest, 12 splits each.
| Model | 16 | 21 | 26 | 31 | 36 | 41 |
|---|---|---|---|---|---|---|
| (c) interactions | 65.5% | 67.4% | 70.3% | 69.7% | 75.0% | 71.4% |
| (a1) ledger re-fitted | 71.8% | 69.7% | 69.8% | 69.7% | 70.7% | 70.7% |
| (a) ledger as is | 73.4% | 73.8% | 71.6% | 74.2% | 74.1% | 73.3% |
| (b2) core linear | 60.7% | 61.6% | 63.1% | 62.2% | 67.1% | 65.8% |
- The flexible model climbs about 6 to 8 points from 16 to 36–41 units. The ledger's form is flat: it has nothing more to learn.
- Projected forward, 52 Tyranid units would never take the flexible model near 88%. Several hundred graded units might get it somewhere. That number of units exists only across armies.
The three tests, answered
(1) In-sample: can any model reach the consensus? Yes, every flexible one does: 88 to 92% against the lists, at or above the reviewers' 88.5%.
- So the datasheet variables can describe the consensus, and no hidden variable is strictly needed to reproduce it.
- But with 144 variables for 52 units that is close to guaranteed. Even the 25-variable linear model reaches 82% at its most flexible.
- The one firm finding here: the ledger's form cannot describe it. Its seven lines, re-weighted freely, top out at 73.8%. Its re-fitted weights on all 52 units come back at step 6's exactly: the inner cross-validation picks the strongest pull toward step 6. Step 6 is already the best the form can do.
(2) Units held out: 64 to 74% against the lists for every model, 14 to 24 points under the ceiling. On scrambled letters the same procedure gives 45 to 51%, so the models do learn something real. They learn about as much as the ledger already knows, no more.
(3) Flexible against the ledger, held out (paired, on the same folds):
| Pair | Gain vs consensus (± SE over repeats) | Gain vs lists | Bootstrap over units, SD | Share of resamples ≤ 0 | Scrambled spread (SD) |
|---|---|---|---|---|---|
| (c) − (a1) | +3.6 ± 0.8 | +1.3 ± 0.9 | 3.6 | 0.17 | 8.7 |
| (e) − (a1) | +2.2 ± 1.3 | 0.0 ± 1.4 | 3.9 | 0.36 | 10.4 |
| (b) − (a1) | 0.0 ± 1.1 | −1.5 ± 1.1 | 3.4 | 0.54 | 8.6 |
| (d) − (a1) | −1.4 ± 1.2 | −2.5 ± 1.3 | 3.0 | 0.70 | 5.9 |
| (c) − (a) | +0.3 ± 0.7 | −2.6 ± 0.7 | 4.7 | 0.47 | 9.5 |
- None of these beats the ledger beyond noise. The best case, interactions over the re-fitted ledger, gains 3.6 points against the consensus but only 1.3 against the lists. One unit resample in six reverses it.
Which of Jordan's three:
| Explanation | Verdict | Why |
|---|---|---|
| Wrong statistical approach | No, not on held-out units | Pairwise losses, ridge, lasso, interactions, monotone trees and neighbours all land within noise of the ledger. The ledger's additive form does cap the in-sample fit at 74%, so the form is a limit. But nothing more flexible carries better to new units at this sample size |
| Not enough variables isolated | Partly | The same units are missed by every model in the same direction. What they share is play the datasheet doesn't hold (below) |
| Not enough rigour | Yes, chiefly as too little data | In-sample 90%+ falls to about 70% held out: the models memorise 52 units. The learning curve still rises at 41. Every earlier test held out lists, not units, so the ledger's 74% was never checked on unseen units before. Re-fitted, it holds 70% |
What "the same consensus %" would take. The reviewers' 88.5% is one expert against four others, each with play experience and the meta in mind. A model built from the datasheet alone, with 52 examples, isn't on the same footing. The math doesn't fail here; the information is thin. Five noisy letters on 52 units carry roughly 100 to 150 bits, and 144 variables can't be priced from that. The sports and pricing literature in design/ranking-methods.md says the same: about one free weight per 15 to 20 ranked items.
The why: what the models key on
Interactions (c), the best held-out model. Permutation importance on held-out pairs, in points of pairwise lost when one variable is shuffled. By family:
| Family | Points lost when shuffled |
|---|---|
| Rule flags | 6.3 |
| Weapons | 3.6 |
| Ledger lines | 3.4 |
| Kill by class | 1.5 |
| Datasheet | 1.0 |
| Interactions | 0.6 |
The top variables:
| # | Variable | Points lost | Direction | Who carries it |
|---|---|---|---|---|
| 1 | Battleline keyword | 1.6 | + | Hormagaunts, Termagants, Gargoyles |
| 2 | Transport keyword | 1.1 | − | Hierophant, Harridan, Tyrannocyte |
| 3 | Precision weapons | 1.1 | + | Lictor, Neurolictor, Red Terror, Norn Emissary, Deathleaper, Haruspex |
| 4 | Burrowers keyword | 0.9 | + | both Ravener units, Red Terror |
| 5 | Kill into light vehicles | 0.8 | + | |
| 6 | Invulnerable save | 0.8 | + | a better save is rated higher |
| 7 | Can arrive away from its zone | 0.8 | + | 38 of 52 units: 3.2 against 2.3 consensus |
| 8 | Psychic weapons | 0.8 | + | Swarmlord, Zoanthropes, Maleceptor, Norn Emissary, Neurotyrant: 4.0 against 2.8 |
Then Heavy guns, which the reviewers rate 3.8 against 2.8: Biovores, Exocrine, Tyrannofex, Termagants, Barbgaunts. After those come Soak in melee (−), Value at the cheapest size (+) and the charge's lands (+).
- Each of these flags marks 3 to 6 units. A model leaning on them is learning which small groups the reviewers like, not a law.
- That is why it doesn't travel to new units. It is also why each flag points at a mechanic worth measuring properly.
The trees (d). By family, the ledger lines dominate (12.4 points). The thresholds, from partial dependence on all 52 units:
- Score line (OC per 100 points, weighted): below about 9.5 the trees take −0.5 letter, and above about 13.6 they give another +0.3. Being cheap bodies with OC lifts a unit, but only up to a point.
- Arrives × melee Kill: any unit that can be placed forward and has melee Kill gets +0.55 letter over one that can't. The amount barely matters past the first step. This is "a threat that arrives" as a yes/no.
- Value: flat until the very top, where the top handful (Value above about 3,100) jump +0.4.
- Kill into monsters: a step at about 0.5 points per 100 (+0.2). Having any answer to monsters counts, not how much.
- Heavy guns +0.6, psychic +0.4, Precision +0.3, Move × melee share above about 7 (fast melee) +0.2, Kill into light vehicles above about 12 per 100 (+0.1).
The interactions the trees use: the depth-3 fit on all 52, by gain.
- Score × Value: rank by Value among units that hold objectives, separately among those that don't.
- Kill into monsters × Value.
- Score × lands.
- Arrives × Value.
- Invulnerable save × Kill into monsters: a durable monster-killer.
- Move × melee share × Heavy guns.
The units the interactions model gets right and the ledger wrong (held-out share of each unit's pairs ordered right):
| Unit | Model / ledger | Consensus | Ledger | What the model sees |
|---|---|---|---|---|
| Toxicrene | 89% / 51% | D, 48th | 23rd | expensive monster, none of Battleline, Precision, Heavy or psychic |
| Harpy | 94% / 61% | about F | 33rd | flyer, nothing the reviewers reward |
| Gargoyles | 66% / 35% | 15th | 47th | Battleline, fly, cheap OC bodies |
| Tyranid Prime | 71% / 51% | 35th | 14th | |
| Tyrant Guard | 78% / 60% | 29th | 49th | |
| Genestealers | 71% / 54% | 15th | 34th | arrive and fight |
| Neurotyrant | 64% / 47% | 17th | 39th | psychic |
| Norn Emissary | 76% / 61% | 11th | 19th | Precision, psychic, 4+ invulnerable save |
Several of these are the units Jordan flagged in his review (2026-10-03-jordan-tyranid-review.md).
The reverse, where the ledger is right and the model wrong:
- Biovores are the starkest. The ledger has them right only through Jordan's hand dial. The datasheet has nothing a model could learn their S from.
- Psychophage, Lictor, Hive Tyrant, Tyrannofex and the Red Terror also go this way.
Missed by every model in the same direction, in letters (mean held-out prediction minus consensus):
| Rated higher than the models give | Gap | Rated lower than the models give | Gap |
|---|---|---|---|
| Biovores | −3.4 | Toxicrene | +1.8 |
| Tyrannofex | −1.5 | Hierophant | +1.5 |
| Exocrine | −1.4 | Tervigon | +1.4 |
| Neurolictor | −1.4 | Ranged Warriors | +1.3 |
| Lictor | −1.4 | Deathleaper | +1.3 |
| Hormagaunts | −1.2 | Tyrannocyte | +1.2 |
- The first group is cheap reach and actions (Biovores, Lictor, Neurolictor), the army's few anti-elite and anti-tank guns (Exocrine, Tyrannofex), and fast objective bodies (Hormagaunts).
- The second group is expensive monsters whose worth is auras, spawning or carrying (Toxicrene, Hierophant, Tervigon, Tyrannocyte), plus units out of fashion.
- None of these is a datasheet number. They are how a unit fits a list and a mission.
Candidate ledger mechanics, ranked
Each is a measurable game quantity that would carry the same signal as a flag above. Each is marked "a known rule, measurable" or "unclear". Each would be one step, with a prediction written first, and judged held out on units (this tool's procedure) as well as on lists.
| # | Mechanic | Status | What it replaces | Units it should move |
|---|---|---|---|---|
| 1 | Threat on arrival: Kill in the turn a unit arrives forward (deep strike, tunnels, infiltrate, scout) × the chance it survives to strike. In place of arrival as Presence bodies | known rule, measurable (the charge model already has arrival; the turn-of-arrival Kill is new) | arrives × melee Kill, the trees' second-strongest split | Genestealers, Raveners, Red Terror, Lictor up |
| 2 | Scarce answers: each unit's share of its army's Kill into a class few units can reach (monsters and vehicles for Tyranids; elite infantry), as the army's marginal loss if the unit is removed. In place of role scarcity's factor | known rule, measurable (the matrix has kill by class; the "removed from a typical list" pass is new) | Heavy guns, Kill into monsters as a step | Exocrine, Tyrannofex, Hive Guard up |
| 3 | Objective bodies with a cap: Score as OC per 100 points, saturating at about the trees' 13.6 rather than linear | known rule, measurable | the Score threshold | Hormagaunts, Gargoyles, Termagants up; big OC-heavy monsters down |
| 4 | Character sniping: Kill into characters through Precision (the target picked, not allocated), against the field's typical leaders | known rule, measurable (the allocation code exists; Precision isn't modelled in the ledger's Kill) | Precision weapons | Lictor, Neurolictor, Red Terror, Norn Emissary up |
| 5 | Transports at their dedicated value only, with walkers and monsters that carry not counted as transports (T-629's 58.1% finding) | known rule, measurable | the Transport flag (negative) | Hierophant, Harridan, Tyrannocyte down |
| 6 | Psychic attacks (they skip cover and some hit modifiers) and the Synapse casters' army reach | partly known; the cover part measurable, the reach part unclear | Psychic weapons | Zoanthropes, Neurotyrant, Maleceptor up |
| 7 | Cheap reach for actions and indirect fire: points per action-capable body, and Kill without line of sight | measurable but the worth is unclear (the Biovores' S rests on how missions are played) | the Biovores' hand dial | Biovores, Lictor, Spore pieces |
| 8 | Value at the cheapest legal size beside the best size: a unit that is good small is easy to fit | measurable | the cheap-Value variable (+) | min-size specialists up |
| 9 | Melee Soak counted down (the models give a negative sign to melee Soak: monsters built to absorb melee are over-credited) | unclear: may be the same "designated anti-tank" effect as step 12's failure | Soak in melee (−) | Toxicrene, Haruspex, Tervigon down |
Next steps
- Pool the armies before fitting anything flexible. The learning curve says data, not method, binds. The 16 armies' graded units (about 800) through the same variables would show whether a model generalises across armies. That is the real held-out test: leave one army out. This tool's table extends to it with the per-army scales the roles work used.
- Add held-out units to the keep rule. Every step so far was judged on lists. A step that also holds on 10 × 5-fold over units is far less likely to be fitting the 52 units.
- Try mechanics 1 to 4 as single steps, in that order. Each replaces a flag the flexible models found with a game quantity, so it can travel to other armies.
- Leave Biovores, Tyrannocyte and Toxicrene to Jordan's calls. No variable we have explains them, and a model that reproduced them would be fitting reputation.
Method notes
- The variables are 144 per unit:
- the datasheet;
- weapons summarised against three reference targets of our own;
- keyword counts, core-rule flags and the effects files' kinds;
- the ledger's lines and Value at step 6;
- Kill by class at β 6;
- 8 interactions by formula. The best form is the ledger's ranking row. The cheapest size enters as its Value, points and models. Datasheets and effects files are read as frozen (BSData 374f505, tools/ledger-inputs.lock.json).
- The loss is pairwise logistic (Bradley–Terry), each pair weighted by its consensus gap. (a1) is the ledger's Value with each weight free as w6·e^u (ridge on u) and a free scale. (c) is a kernel ranker with (1 + x·z/p)², which is linear plus every pairwise product and square. (d) is Newton-leaf boosting on the same loss, at least 3 units a leaf, colsample 1 or 0.3, with the dictionary's monotone signs: more wounds, Toughness, Kill, OC or Value is never worse, and more points is never better. It is written in Node, since no Python ML library is installed here. (e) is the distance-weighted mean consensus of the k nearest units on the standardised variables.
- Scrambled letters score 45 to 51% held out. Below 50% is the usual small-sample bias of cross-validation on shuffled targets. The spread, 4 to 9 points, is the yardstick for a "gain".
- Permutation importance shuffles a variable's column across all 52 units. Each unit is then scored by the fold fit that never saw it, on 10 repeats × 3 shuffles. Correlated variables share their importance, so families are reported too.
- Tests: tests/unit/diagnose.test.js, a smoke test on toy data (pairwise counts, rankers recover a known order, trees stay monotone).