Tactical Reroll
⋯

Critique of the unit tier brief (2026-10-02)

From reports/tiers/2026-10-02-brief-critique.md , rendered when the site is built.

Verdict. The pipeline is careful and well documented, but its only real evidence of validity is a ρ of 0.50 against one YouTube list for the army it has been tuned on for eight rounds, while both untouched panels sit at or below zero. As it stands, the maths shows that the system can be fitted to Tyranids. It does not yet show that the system measures worth.

Critiques, ranked by how much fixing each would move the letters

1. Scoring units alone, at minimum size, with leaders credited apart, removes the main reason many units are taken. In real lists a unit's value often comes from its partner: a squad that a character makes lethal, or a horde that only works at full size. Scoring the bare squad at its smallest size, and handing the gain to the leader, demotes the "chassis" units and promotes characters, and it misprices units whose per-point profile changes with size. If this is right you'd see: the biggest gaps against experts and against list share fall on units usually fielded at full size or with an attached leader, with the system rating them low. Cheapest test: re-score the 44 Tyranid units and the 30 Marine units at their most-listed size and with their most-listed leader attached, then recompute ρ.

2. Taking the maximum over detachments, loadouts, jobs and transport bundles builds in an optimistic bias that grows with the number of options. The maximum of several noisy scores rises with the number of draws. Armies with many detachments (the Marine chapters especially) and units with many builds or eligible jobs get lifted for having choices, not for being good. If this is right you'd see: the letter correlates with the count of a unit's detachments, loadouts and eligible jobs, after controlling for raw Offense and Defense. Cheapest test: regress percentile on those option counts, then re-score with each army's most-listed detachment fixed and compare the Marine ρ.

3. Hammer's 0.75 weight on the peak target assumes the unit always meets the target it wants. Against an opponent who chooses, a specialist meets its ideal target only some of the time. Weighting the best case at three to one rewards narrow anti-tank or anti-elite tools over generalists, and the 38% TEQ weight compounds this. If this is right you'd see: Hammer leaders whose peak target is far above their mean, and experts rating them lower than the system does. Cheapest test: set the split to 0.25/0.75 and to 0/1, then report the letter churn and ρ on all three panels.

4. Anvil counts points twice. Defense per 100 points multiplied by OC per 100 points scales roughly as the inverse square of cost, so cheap bodies dominate Anvil percentiles almost by construction. If this is right you'd see: Anvil percentile tracks points cost more closely than it tracks Defense or OC. Cheapest test: correlate Anvil with 1/points and 1/points², then re-run with a geometric mean or a single per-point division.

5. The validation is one panel, used to steer, and it fails to transfer. Eight rounds of changes were judged by ρ against Auspex Tyranids, and Tyranids also have hand files (Synapse at 75%, Norn approximation, Spore Mines). The rise from 0.07 to 0.50 is what selecting rules on a single 44-unit sample produces. The −0.05 and −0.13 on the clean panels are the honest out-of-sample figure, and treating a gap under one SE as noise means nothing can ever be shown to have failed. If this is right you'd see: the Marine and Ork ρ stay flat or fall across rounds while the Tyranid ρ climbs. Cheapest test: score every historical version (V2b to round 7) against the Marine and Ork panels and plot all three lines together.

6. The weights are fitted on list share, then the system is checked against list share and against experts who read the same lists. Target mix, incoming fire, the S/A split, the aura partners and the Battle-shock rate all come from winning lists. Agreement with list share or with experts is partly agreement with yourself, and the 0.65 of the live letter, which includes list share, shows how much popularity alone buys. If this is right you'd see: fitted weights taken from a different period or corpus change the letters little, yet ρ against list share drops. Cheapest test: fit the weights on July only and validate on September lists, or swap in uniform weights and compare ρ.

7. With no terrain, no line of sight and no opponent choices, the model favours shooting over melee, infiltration and screening. Tournament tables are dense with ruins. Melee, Scouts, Infiltrators, Stealth and indirect fire earn their value from terrain, and screening has no job here at all. If this is right you'd see: gunline units rated above expert consensus, and screens, infiltrators and melee units below it. Cheapest test: split the panel residuals by ranged versus melee and by Scouts/Infiltrate keywords, then check the sign.

8. A flat 30-point CP, with no stratagems, hides the most army-specific value in the game. Many units are taken because of a stratagem they unlock or receive, and CP value differs by army and detachment. The brief records that the CP value moved from 0 to 15 to 30 by judgement, and that change alone reshuffles Banner. If this is right you'd see: Banner letters for CP generators swing sharply between CP = 15 and CP = 45. Cheapest test: re-run at 15 and 45 points and count how many letters change.

9. Most of the driving constants are set by hand and have never been through a sensitivity analysis. The position floor, the Runner gate (taken from an unrelated weight-class scheme), the revive ×1.15, Synapse coverage at 75%, a lone character at ×0.5, a Lone Operative at ×2 and the new +5/+2 bonus all lack a test showing that the letters are robust to them. If this is right you'd see: varying each constant by ±50% moves a large share of units by a letter. Cheapest test: a one-at-a-time sweep that reports how many letters change for each constant, then rank the constants by influence.

10. The comparison field and the absolute letters are distorted by duplicate datasheets and by internal competition. A shared datasheet counts once per army, so the many Marine chapters crowd the 1,161-unit field and the percentiles become partly marine-relative. Players also judge a unit against the alternatives in its own codex, and fixed cuts across armies ignore that. If this is right you'd see: Marine-family units bunched in the middle bands, and within-army ranks agreeing with experts better than the all-army letters do. Cheapest test: dedupe datasheets before taking percentiles, then compute each panel's ρ on within-army ranks against all-army ranks.

Do not undo