Tactical Reroll
⋯

Why the ledger can't reach the reviewers on Tyranids: the three-way diagnosis (T-632, 3 Oct)

From reports/tiers/2026-10-03-diagnosis.md , rendered when the site is built.

Jordan's challenge: "You have every variable available to you based on the datasheets. If you had NOTHING ELSE there should be a mathematical approach to solve this mystery. Either we are not attacking this with the correct statistical approach, or we have the right approach but not enough variables isolated, or we have not tried enough rigor." Command: node tools/diagnose.js (5 minutes; --quick for a smoke run). Every number below is in reports/tiers/2026-10-03-diagnosis-tables.md and -diagnosis.json; the table of variables is 2026-10-03-tyranid-features.csv with its dictionary in -tyranid-features.md. A report: no ledger, page, data file, setting or input lock changed. Our own numbers; unit names only.

The short answer

The numbers

The target: each unit's consensus, its mean letter over the five headline Tyranid lists that grade it (S 5 … F 0; every unit is graded by 2 to 5 lists).

How it is measured:

ModelIn-sample vs consensus (tuned / most flexible)In-sample vs listsHeld out vs consensus (± SE over repeats)Held out vs lists (± SE)Scrambled vs consensus (SD)
Ceiling: each list against the other four–88.5%–88.5%–
(a) the ledger as it is (step 6)73.3% / 73.3%74.5%72.9% ± 0.474.1% ± 0.650.8% (4.5)
(a1) the ledger's form, its 7 weights re-fitted in the folds73.3% / 73.8%74.5%69.7% ± 1.070.1% ± 1.049.1% (3.8)
(b) linear on all 144 variables, ridge or lasso89.1% / 98.8%87.7%69.7% ± 1.368.6% ± 0.945.9% (6.9)
(b2) linear on a short core list (25), ridge72.2% / 81.7%71.0%66.0% ± 0.964.2% ± 0.846.5% (8.4)
(c) linear + every pairwise interaction (kernel ridge)97.2% / 97.2%92.2%73.3% ± 0.971.4% ± 1.044.8% (6.6)
(d) boosted trees, depth 1 to 3, monotone constraints95.3% / 99.7%91.5%68.3% ± 0.967.6% ± 0.849.3% (3.8)
(e) nearest neighbours (similar units)100% / 100%92.3%71.9% ± 0.870.1% ± 1.145.2% (7.5)

Caveat on (a): the step-6 ledger was tuned on these same 52 units and lists, and carries Jordan's hand dial on the Biovores. Its "held out" number is really in-sample, so it is optimistic. (a1) is the fair version of the ledger: its weights are fitted only on the training units. The SEs over repeats measure only how the folds were split. The honest noise is the bootstrap over units and the scrambled spread below.

Held out, per list:

AuspexHivemindSecondMaelstromAstrategas
ceiling91.7%87.7%81.9%94.9%86.5%
(a) ledger79.9%78.4%76.2%70.8%65.1%
(a1) ledger re-fitted76.7%74.3%71.9%66.4%61.3%
(c) interactions77.9%71.7%71.2%69.9%66.5%
(d) trees76.0%69.0%72.5%61.4%59.1%
(e) neighbours76.6%71.2%69.1%69.4%64.2%

The learning curve: trained on m random units with tuning inside, scored on the pairs among the rest, 12 splits each.

Model162126313641
(c) interactions65.5%67.4%70.3%69.7%75.0%71.4%
(a1) ledger re-fitted71.8%69.7%69.8%69.7%70.7%70.7%
(a) ledger as is73.4%73.8%71.6%74.2%74.1%73.3%
(b2) core linear60.7%61.6%63.1%62.2%67.1%65.8%

The three tests, answered

(1) In-sample: can any model reach the consensus? Yes, every flexible one does: 88 to 92% against the lists, at or above the reviewers' 88.5%.

(2) Units held out: 64 to 74% against the lists for every model, 14 to 24 points under the ceiling. On scrambled letters the same procedure gives 45 to 51%, so the models do learn something real. They learn about as much as the ledger already knows, no more.

(3) Flexible against the ledger, held out (paired, on the same folds):

PairGain vs consensus (± SE over repeats)Gain vs listsBootstrap over units, SDShare of resamples ≤ 0Scrambled spread (SD)
(c) − (a1)+3.6 ± 0.8+1.3 ± 0.93.60.178.7
(e) − (a1)+2.2 ± 1.30.0 ± 1.43.90.3610.4
(b) − (a1)0.0 ± 1.1−1.5 ± 1.13.40.548.6
(d) − (a1)−1.4 ± 1.2−2.5 ± 1.33.00.705.9
(c) − (a)+0.3 ± 0.7−2.6 ± 0.74.70.479.5

Which of Jordan's three:

ExplanationVerdictWhy
Wrong statistical approachNo, not on held-out unitsPairwise losses, ridge, lasso, interactions, monotone trees and neighbours all land within noise of the ledger. The ledger's additive form does cap the in-sample fit at 74%, so the form is a limit. But nothing more flexible carries better to new units at this sample size
Not enough variables isolatedPartlyThe same units are missed by every model in the same direction. What they share is play the datasheet doesn't hold (below)
Not enough rigourYes, chiefly as too little dataIn-sample 90%+ falls to about 70% held out: the models memorise 52 units. The learning curve still rises at 41. Every earlier test held out lists, not units, so the ledger's 74% was never checked on unseen units before. Re-fitted, it holds 70%

What "the same consensus %" would take. The reviewers' 88.5% is one expert against four others, each with play experience and the meta in mind. A model built from the datasheet alone, with 52 examples, isn't on the same footing. The math doesn't fail here; the information is thin. Five noisy letters on 52 units carry roughly 100 to 150 bits, and 144 variables can't be priced from that. The sports and pricing literature in design/ranking-methods.md says the same: about one free weight per 15 to 20 ranked items.

The why: what the models key on

Interactions (c), the best held-out model. Permutation importance on held-out pairs, in points of pairwise lost when one variable is shuffled. By family:

FamilyPoints lost when shuffled
Rule flags6.3
Weapons3.6
Ledger lines3.4
Kill by class1.5
Datasheet1.0
Interactions0.6

The top variables:

#VariablePoints lostDirectionWho carries it
1Battleline keyword1.6+Hormagaunts, Termagants, Gargoyles
2Transport keyword1.1−Hierophant, Harridan, Tyrannocyte
3Precision weapons1.1+Lictor, Neurolictor, Red Terror, Norn Emissary, Deathleaper, Haruspex
4Burrowers keyword0.9+both Ravener units, Red Terror
5Kill into light vehicles0.8+
6Invulnerable save0.8+a better save is rated higher
7Can arrive away from its zone0.8+38 of 52 units: 3.2 against 2.3 consensus
8Psychic weapons0.8+Swarmlord, Zoanthropes, Maleceptor, Norn Emissary, Neurotyrant: 4.0 against 2.8

Then Heavy guns, which the reviewers rate 3.8 against 2.8: Biovores, Exocrine, Tyrannofex, Termagants, Barbgaunts. After those come Soak in melee (−), Value at the cheapest size (+) and the charge's lands (+).

The trees (d). By family, the ledger lines dominate (12.4 points). The thresholds, from partial dependence on all 52 units:

The interactions the trees use: the depth-3 fit on all 52, by gain.

The units the interactions model gets right and the ledger wrong (held-out share of each unit's pairs ordered right):

UnitModel / ledgerConsensusLedgerWhat the model sees
Toxicrene89% / 51%D, 48th23rdexpensive monster, none of Battleline, Precision, Heavy or psychic
Harpy94% / 61%about F33rdflyer, nothing the reviewers reward
Gargoyles66% / 35%15th47thBattleline, fly, cheap OC bodies
Tyranid Prime71% / 51%35th14th
Tyrant Guard78% / 60%29th49th
Genestealers71% / 54%15th34tharrive and fight
Neurotyrant64% / 47%17th39thpsychic
Norn Emissary76% / 61%11th19thPrecision, psychic, 4+ invulnerable save

Several of these are the units Jordan flagged in his review (2026-10-03-jordan-tyranid-review.md).

The reverse, where the ledger is right and the model wrong:

Missed by every model in the same direction, in letters (mean held-out prediction minus consensus):

Rated higher than the models giveGapRated lower than the models giveGap
Biovores−3.4Toxicrene+1.8
Tyrannofex−1.5Hierophant+1.5
Exocrine−1.4Tervigon+1.4
Neurolictor−1.4Ranged Warriors+1.3
Lictor−1.4Deathleaper+1.3
Hormagaunts−1.2Tyrannocyte+1.2

Candidate ledger mechanics, ranked

Each is a measurable game quantity that would carry the same signal as a flag above. Each is marked "a known rule, measurable" or "unclear". Each would be one step, with a prediction written first, and judged held out on units (this tool's procedure) as well as on lists.

#MechanicStatusWhat it replacesUnits it should move
1Threat on arrival: Kill in the turn a unit arrives forward (deep strike, tunnels, infiltrate, scout) × the chance it survives to strike. In place of arrival as Presence bodiesknown rule, measurable (the charge model already has arrival; the turn-of-arrival Kill is new)arrives × melee Kill, the trees' second-strongest splitGenestealers, Raveners, Red Terror, Lictor up
2Scarce answers: each unit's share of its army's Kill into a class few units can reach (monsters and vehicles for Tyranids; elite infantry), as the army's marginal loss if the unit is removed. In place of role scarcity's factorknown rule, measurable (the matrix has kill by class; the "removed from a typical list" pass is new)Heavy guns, Kill into monsters as a stepExocrine, Tyrannofex, Hive Guard up
3Objective bodies with a cap: Score as OC per 100 points, saturating at about the trees' 13.6 rather than linearknown rule, measurablethe Score thresholdHormagaunts, Gargoyles, Termagants up; big OC-heavy monsters down
4Character sniping: Kill into characters through Precision (the target picked, not allocated), against the field's typical leadersknown rule, measurable (the allocation code exists; Precision isn't modelled in the ledger's Kill)Precision weaponsLictor, Neurolictor, Red Terror, Norn Emissary up
5Transports at their dedicated value only, with walkers and monsters that carry not counted as transports (T-629's 58.1% finding)known rule, measurablethe Transport flag (negative)Hierophant, Harridan, Tyrannocyte down
6Psychic attacks (they skip cover and some hit modifiers) and the Synapse casters' army reachpartly known; the cover part measurable, the reach part unclearPsychic weaponsZoanthropes, Neurotyrant, Maleceptor up
7Cheap reach for actions and indirect fire: points per action-capable body, and Kill without line of sightmeasurable but the worth is unclear (the Biovores' S rests on how missions are played)the Biovores' hand dialBiovores, Lictor, Spore pieces
8Value at the cheapest legal size beside the best size: a unit that is good small is easy to fitmeasurablethe cheap-Value variable (+)min-size specialists up
9Melee Soak counted down (the models give a negative sign to melee Soak: monsters built to absorb melee are over-credited)unclear: may be the same "designated anti-tank" effect as step 12's failureSoak in melee (−)Toxicrene, Haruspex, Tervigon down

Next steps

  1. Pool the armies before fitting anything flexible. The learning curve says data, not method, binds. The 16 armies' graded units (about 800) through the same variables would show whether a model generalises across armies. That is the real held-out test: leave one army out. This tool's table extends to it with the per-army scales the roles work used.
  2. Add held-out units to the keep rule. Every step so far was judged on lists. A step that also holds on 10 × 5-fold over units is far less likely to be fitting the 52 units.
  3. Try mechanics 1 to 4 as single steps, in that order. Each replaces a flag the flexible models found with a game quantity, so it can travel to other armies.
  4. Leave Biovores, Tyrannocyte and Toxicrene to Jordan's calls. No variable we have explains them, and a model that reproduced them would be fitting reputation.

Method notes