Tactical Reroll
⋯

A blind appraisal of the tier method (statistics and economics)

From reports/tiers/2026-10-05-blind-appraisal.md , rendered when the site is built.

5 Oct 2026. An independent appraisal, written blind first (section 1, from the brief alone, before reading the repo's method; not revised afterwards), then compared with the points ledger and the T-662 plan (family dials, an exhaustive grid, a winning set, Tyranids first, the other armies as the exam). A report: no model, setting, benchmark, data file or frozen input changed. The cheap checks in section 2 ran read-only through docs/lab/ledger/ledger.js at Track B. Our own words and numbers; unit and keyword names only. Tournament counts from grimstat-corpus (https://github.com/N041M/grimstat-corpus, CC BY 4.0) via docs/data/inclusion.json.

In plain English

1. Phase 1: the blind plan (written before reading the repo's method; not revised)

(a) How I would solve it

Name the target first. A reviewer's letter is not "damage per point". It is, roughly, "how much does taking this unit, at today's points, raise the win chance of a well-built competitive list of this army". That is a marginal value inside a list (a role filled, against a field, under a points budget), with a dose of "how often is it an auto-include". So I treat each unit as having one latent worth θ_u (continuous), and every data source as a noisy, biased measurement of it.

Two layers: a structural layer the engine supplies, a statistical layer that is small.

  1. Structural layer (the game, not fitted). For every unit, the dice engine turns the datasheet into a handful of capability numbers per point, each against a field weighted by what is actually played:
    • Kill: expected damage / burst-kill probability against 4 to 6 canonical target classes (horde, elite infantry, monster/vehicle, character), weighted by the field's mix.
    • Durability: points of enemy attacks needed to remove it, against the field's attackers (this is where T, W, Sv, invuln, FNP, Stealth, -1 to hit, -1 damage all live, mechanically).
    • Board: OC × expected survival turns; movement/reach (M, Deep Strike, Scouts, Fly, Infiltrate) as "turn-1 / turn-2 objective reach probability" from a simple board model.
    • Utility: actions, auras/buffs (valued by playing the buffed unit with and without), transport. This gives maybe 6 to 8 interpretable capability scores, not 341 variables. Perks enter the model through the engine (counterfactual: same unit minus Stealth, recompute), so a perk's worth is a derived quantity, not a free parameter.
  2. Statistical layer (fitted, few parameters). θ_u = Σ_k w_k · g(cap_k,u) + a_army + e_u, with monotone constraints (w ≥ 0), a concave g (diminishing returns), and a points term that makes it value-for-cost (log value minus log points, hedonic style). Letters enter through a many-facet ordinal model (cumulative probit / many-facet Rasch): each reviewer r has their own cut points τ_r (one reviewer's "B" is another's "C"), their own noise σ_r, and the ratings are correlated through a shared "consensus" factor (they read the same tournament results). Fit in a Bayesian way (Stan/PyMC or a hand-written Laplace approximation), with priors centred on "trust the engine": w equal across lines after scaling, residual e_u shrunk to 0.
  3. A second, independent target: revealed preference. The ~190 tournament lists are choices under a 2,000-point budget. A conditional-logit / discrete-choice model of "which units appear, how many copies" with price (points) as a regressor, sharing θ_u with the letters but with its own biases (army popularity, model availability, rule-of-three limits, synergy with the chosen detachment). Joint fitting gives a check: units the reviewers love but nobody takes (or vice versa) are diagnostic, not noise to be averaged away.
  4. A third, free and large data source: the points themselves. Game designers price ~1,000+ units across all armies with a hedonic model in their heads. Regressing log(points) on engine capabilities and perk indicators across every army gives the designer's implied price of each perk with n ≈ 1,000, not 50. A unit's tier should then largely be its residual: worth (our model) minus price. This is the cleanest way to get "true worth per perk" with real statistical power, and it's GW's numbers (points), not GW's text.

Perk values are reported as: (i) the engine's counterfactual change in capability (exact, no fitting), times (ii) the fitted line weight, with a posterior interval; for perks the engine can't play (Lone Operative, Scouts, Deep Strike, actions) one explicit parameter each, with a prior from the hedonic points regression. Interactions are attributed with Shapley values over the perks.

Validation, decided in advance. Fit on all armies with reviewers (not one), with leave-one-army-out cross-validation and leave-one-reviewer-out. Report held-out pairwise ordering accuracy against (a) the inter-rater ceiling computed the same way, (b) dumb baselines: points alone, damage-per-point alone, tournament inclusion alone, "last edition's letter". A model that can't beat "inclusion rate" on held-out armies doesn't explain anything. A permutation control (shuffle letters within army) gives the chance distribution of the whole search, not of one fit.

(b) Main pitfalls for any approach

  1. Tiny effective sample. ~50 Tyranid units, 5 reviewers who are not independent: the effective number of independent observations is close to 50 units, maybe fewer (S and D are a handful each). Any model with more than ~5 free parameters fit on one army is mostly fitting noise.
  2. Searching a huge space on a tiny target (garden of forking paths). The best of millions of combinations is optimistically biased; the optimism must be estimated by redoing the whole search inside resampling (nested CV / bootstrap of the search), not by bootstrapping the winner's score.
  3. Construct validity. Letters mix per-point efficiency, ceiling power, role scarcity (the army's only anti-tank), synergy with detachments and leaders, and fashion. A pure per-point efficiency model can't reach the ceiling and shouldn't try to by bending perk values.
  4. Reviewers aren't independent ground truth. They read the same tournament results and each other; their 84 to 89% agreement is partly shared information, so the "ceiling" overstates how much truth there is, and fitting them partly fits the meta's popularity.
  5. Non-identifiability. Many dials multiply the same line; correlated perks (big monsters have both high T and W and high points) make individual values unidentifiable from 50 units. Multiple different perk values fit equally well; the perk "values" then mean nothing, even if the ranking is right.
  6. Context: detachments, leaders, transports, army rules. A unit is worth more where its army lacks that role (diminishing returns within a list); single-unit valuation ignores this.
  7. Time: points and reviewers' dates differ; a letter given at old points is a different item.
  8. Selection of graded units: reviewers skip Legends, irrelevant or brand-new units; the missing are not random (mostly bad), which squashes the bottom of the scale.
  9. Coarse ordinal scale. Five letters, uneven use (few S, many B/C), ties: Spearman is unstable, distances between letters are not equal, and reviewers' cut points differ.
  10. Engine misspecification. Expected damage ignores variance, positioning, terrain, the missions' timing; a perk the engine can't see (Deep Strike) will be compensated by inflating something it can see.
  11. Goodhart / transfer. Tuning to one army's reviewers makes the system learn that army's quirks (Synapse, Shadow in the Warp) and attribute them to generic perks.

(c) Analogous problems and how they're solved


2. Our method against the blind plan: where they agree, where they differ, who is right

Read for this section: design/tier-engine.md, design/tier-briefing.md, reports/tiers/ledger-steps.md (to step 39 and the Two tracks and Tyranids-first rules), reports/tiers/2026-10-03-ledger-crack.md, 2026-10-05-hedonic.md, 2026-10-05-dials.md, design/decisions.md (3 to 5 Oct), design/mission-notes.md (structure), docs/lab/ledger/ledger.js (header, SWITCHES, BASE), and, after writing section 1, the headers of the two earlier reviews (2026-10-04-review.md, 2026-10-05-critique.md) so as not to repeat them.

Where they agree

PointBlind planOur methodVerdict
The engine supplies the structure; perks act through itcapability numbers per point from the engine; perks valued by counterfactualthe ledger's lines (Kill as Burst, Soak, Score, Hold, Actions, Spawn, Presence, Carry) per 100 points; "1× = trust the dice engine"Same idea. The ledger's lines are a better-worked version of my capability layer (Burst, preferred targets, the mission catalogue).
Few fitted numbers~5 to 8 weights on one armya budget of about 8 knobs, the rest hibernated (3 Oct)Agree. The T-662 grid quietly breaks this budget (below).
A chance controlpermutation of letters, run through the whole searchshuffled-letters controlAgree in spirit; ours must shuffle through the whole pipeline (it does in ledgercrack.js; T-662 must too).
Report the set of good models, not a winnerRashomon setthe "winning set" within the bootstrap band, weighted by fitAgree. Right instinct; the implementation needs a calibration (section 3, pitfalls 1 and 2).
Score held-out armies and baselinesleave-one-army-out; baselines points, inclusionthe 14 never-fitted armies as the exam; list share and points as baselinesAgree.
Timeletters at old points are different itemsunits whose points changed flagged unfittableAgree, but it doesn't go far enough (pitfall 6).
Ceilinginter-rater ceiling the same way as the model's metric~89%, each reviewer against the others' consensusAgree; I reproduce it (below).

Where they differ, and who is right

  1. Fitting on one army (52 units) vs pooling. I'd fit game-wide numbers on every army with a panel, weighted by panel reliability, holding armies out. We fit on Tyranids and use the rest as an exam. Who is right: both, for different jobs. Tyranids-first is right for debugging mechanisms (five reviewers, clean residuals; the step log shows why). It is wrong for estimating game-wide perk values: a "game-wide" dial fit on Tyranids is estimated from the Tyranid carriers only, and many families have three to seven of them (in my rough family split: Fights First 3, Lone Operative 4, Scouts/Infiltrators 6, Precision/Indirect 7). That isn't a perk value; it's a per-unit adjustment with a game-wide label. And one "family", Synapse, exists only in this army. The pooled diagnosis (4 Oct) found flexible models don't travel; that is not evidence against pooling a dozen constrained dials, which is exactly the case pooling helps.
  2. An exhaustive grid with a fit-weighted winning set vs a likelihood with a prior. I'd fit an ordinal (many-facet) likelihood with a prior centred on the engine and read intervals off a posterior. T-662 enumerates tens of millions of combinations on a coarse grid and keeps those within noise of the best. Who is right: the grid is more transparent and easier to show Jordan, and that matters. But a fit-weighted grid is a pseudo-posterior whose prior is "uniform on {0, 0.5, 1, 2, 4} for every family", and for any dial the data can't see, the winning set's 10 to 90% range is just that prior coming back (my toy below: a family every unit carries shows exactly 20/20/20/20/20%). The "gentle pull toward 1×" reported with and without is the prior; it should be the default, not an option. Keep the grid as the engine, but turn its weight into a proper likelihood × prior, and calibrate it (pitfalls 1 and 2).
  3. What a dial multiplies. I'd value a perk as the engine's increment (line with the perk minus line without), times a weight. T-662 says a dial "multiplies the line it affects" (defensive perks scale survival). If that means the carrier's whole line, "Stealth ×2" means "every unit with Stealth has its whole Soak doubled", which mixes the perk with everything else the carrier has. That is the same flaw the 5 Oct critique found in the support multipliers (Winged Hive Tyrant ×4.1 for Synapse on its whole worth). I am right here, if the summary reflects the design.
  4. A second target. I'd fit reviewer letters and tournament inclusion jointly, with inclusion's biases modelled. Ours keeps popularity out of the score (Jordan's principle) and tried fitting the dials to inclusion (the "hedonic" report): no transfer. Who is right: the product principle is Jordan's call and sound for the published letter; but for estimating what a perk is worth, inclusion is a measurement, not a display, and the report shows it carries most of what the reviewers see (list share alone orders the Tyranid consensus right on 86%, the ledger on 78%). Also the repo's "hedonic" isn't hedonic pricing in the economist's sense (regressing price on attributes); it's fitting dials to popularity. The real hedonic regression (GW's points on attributes) hasn't been done (section 4, idea 1).
  5. Process. Our method pre-registers a prediction per step, names the aimed units, and keeps a decision log. That is better practice than my blind plan, which had none of it. We are right here. Keep it.
  6. Explanation. My plan's statistical layer explains less than the ledger's named lines and residual loop. For a site whose product is "show your working", ours is right on the shape of the explanation; the statistics must sit under it, not replace it.

Cheap checks I ran (Node, read-only, a few seconds each; no repo file changed)

All on docs/lab/ledger/tyranids.json at Track B (tools/ledger-benchmarks/2026-10-05-trackB.json) through Ledger.compute(), the five Tyranid lists, pairwise with ties in the model counted half, mean over the five lists. Tournament counts from grimstat-corpus (CC BY 4.0) via docs/data/inclusion.json.


3. Pitfalls we are missing, ranked by how much each could mislead us

  1. The grid can manufacture the gain it reports (selection on a tiny target). Effective sample: ~52 units, five correlated lists. With tens of millions of combinations the best is biased upward by about the size of a real effect (the toy null: 0 to 4 points from eight families; more families, more bias). The shuffled-letters control doesn't measure this well: shuffled letters start at ~50%, the real ones at ~78%, and gains near a ceiling aren't comparable to gains from chance level. The 3 Oct crack already showed the problem with 11 knobs (ρ 0.63 on shuffled letters, 0.70 on real), and the 5 Oct dials report admits one more layer (the four dials were chosen from one-at-a-time sweeps on the same 52 units). Check: a calibrated null world (simulation-based calibration): simulate letters from the 1× ledger with reviewer-like noise matched to the real agreement, run the entire T-662 pipeline (grid, band, winning set, family split, Tyranid shortlist) on 50 such worlds, and report the real gain's percentile. Publish nothing unless it clears the null's 95th percentile. Also run the pipeline inside unit folds (nested), and on the Track B history (the choice of Track B's members is itself a forking path).
  2. The winning set's ranges are mostly the grid's prior (identification). A family carried by almost every unit is a copy of its line's weight; rank metrics ignore overall scale; a family with 3 to 7 Tyranid carriers is fitted by those few units' letters. The toy shows the signature: a universal family comes back 20/20/20/20/20 (a 10 to 90% range of 0 to 4×, which reads like a finding), and few-carrier families come back wide and skewed in null worlds too. Fix: for each family print carriers (per army), the winning set's concentration against (a) the grid's uniform prior and (b) the null worlds' concentration, and the units that move it (drop-one-unit influence). Add knockoff families (the same number of carriers, assigned at random within unit type) to the grid; a real family whose range is no tighter than its knockoffs' is "not identified", and the site says so instead of showing a range.
  3. Dial semantics: a dial on the carrier's line is not a perk's value. If "defensive perks × 2" doubles a carrier's whole Soak, the dial confounds the perk with the carrier's other traits and stacks (a Stealth + Feel No Pain + invulnerable unit gets three multipliers on one line). Fix: a dial scales the perk's engine increment, Δ = line with the perk − line with it removed in the engine; perks the engine can't remove (Deep Strike, Scouts, Lone Operative, actions) enter as an additive credit in points per 100, with the carrier count printed beside it.
  4. Multiplicative stacking and the grid's corners. {0, 4} on several families multiplies to 0 or 64× on units carrying several. 0 isn't "a smaller effect"; it is a different hypothesis (the perk is worthless, or the line is switched off), and a grid that mixes them reports a range spanning "worthless to quadruple". Fix: dials in log space with a prior (log-dial ~ Normal(0, 0.35), so ×0.5 and ×2 are each about 2 SD away); drop 0 from the grid or treat it as a separate model; report the largest combined multiplier any unit gets in the winning set.
  5. What the letters measure (construct). On Tyranids, list share alone is within about 3 points of the reviewers' own ceiling. Reviewers are, to a large degree, reporting what the meta takes. A dial fitted to them learns "what the meta rewards", including army context (Synapse, the best detachment, the role a unit fills in a list), and labels it as the worth of a generic perk. That's the opposite of "true worth per perk". Check: fit the same grid with list share as the target (one army out at a time) and with the consensus residual on list share as the target; a perk value that survives only on raw letters is popularity, not mechanism. Report the partial agreement with share held fixed (the hedonic report's 0.33 on Tyranids) as the headline of "what the ledger adds".
  6. Reviewers are not five independent truths, and two of them graded a different game. Maelstrom and Astrategas are 10th edition and cluster together. Flagging units whose points changed misses rule changes, and the pairwise mean double-counts the 10th-edition view. Fix: a Thurstonian rater model with a shared factor and a date/edition facet; report every fit on the three 11th-edition lists alone beside the five; compute the ceiling within the 11th-edition cluster. Weight lists by estimated reliability, not one each.
  7. Too little power for the exam. The 14 armies carry 1 to 3 lists each, some near or below chance for any model (Thousand Sons 38.5%, Space Wolves 51%). A third of them held out is about five panels, each with a unit-bootstrap SE of roughly 4 to 6 points, pooled perhaps 2 to 2.5. Choosing among a Tyranid shortlist whose members differ by under 2 points on those armies is choosing by noise. Fix: before running it, compute the minimum detectable difference on the held-out third (paired bootstrap over units and lists) and say in advance that the shortlist is chosen by prior (closest to 1×) unless a member beats it by more than that.
  8. Points-to-tier non-linearity and replacement level. Value per 100 points assumes tier ∝ worth per point. The check above says the reviewers lean even further toward cheap units (points^−0.25 helps 1.6 points), which is what a replacement-level model predicts (a cheap unit that fills a slot frees points). One exponent or a replacement level is one interpretable parameter; fit it before family dials, or the families will soak it up (cheap units carry certain perks more often).
  9. Best-of-many selection inside the ledger. Each unit is scored at its best solo form over sizes, loadouts and detachments. Units with more options get a higher maximum by chance of measurement, not worth. Check: regress the consensus residual on the number of forms per unit; if it's negative, score the form players field (the mode) or shrink the max.
  10. A metric with a surrogate. The search steers by Spearman or an ordinal likelihood and reports pairwise. Fine, but pick one surrogate in advance; trying both and keeping the better is one more fork. Ties: the model never ties, the reviewers often do; keep the ceiling computed on exactly the model's rule.

4. Directions not yet tried

#Idea (field it comes from)What it needsDo we have itEffortExpected payoff
1The designers' hedonic price (economics, Rosen). Regress log(points) for every fieldable unit in all 36 armies (~1,100) on the engine's capabilities (Kill, Soak, OC, Move) and perk flags (Stealth, Feel No Pain, Deep Strike, Scouts, Lone Operative, Fights First, Lethal/Sustained/Devastating, Anti, Precision, Indirect), monotone, with army fixed effects. The coefficients are what the game's designers charge for each perk, with standard errors from n ≈ 1,100, not 52. A unit's residual is its mispricing, which is what a tier list is about.points and datasheetsyes (build/units.json, BSData; points are numbers the ledger already reads; our coefficients are our own)1 to 2 daysHigh. The first perk values with real statistical power, independent of reviewers; a natural prior for every dial; a headline the site can show ("the game charges about X points for Stealth on infantry; our engine says it's worth Y").
2Calibrated null worlds and knockoff families (simulation-based calibration, model-X knockoffs). As in my toy: letters simulated from the 1× ledger with reviewer-like noise; fake families with matched carrier counts. Run the whole T-662 pipeline on them.the ledger, the lettersyeshalf a day (the toy is most of it)High as a guard: it turns "within the band" into "beats chance by this much", per perk.
3A many-facet Thurstonian model of the reviewers (psychometrics; Dawid–Skene; sensory panels). Each unit has a latent worth = engine lines × dials + residual; each list has its own cut points, noise and an edition facet; a shared factor for what all reviewers see. Fit by Laplace approximation or a small sampler in Node, pooled over the 16 armies with panels. Gives honest posterior intervals for each dial, reviewer reliabilities (who to trust), and replaces the fit-weighted grid.the panelsyes2 to 3 daysHigh methodologically; medium in agreement (it fixes the arithmetic, not the missing lines).
4A discrete-choice model of list building (conjoint, industrial organisation). Each tournament list is a basket chosen under a 2,000-point budget: a conditional logit on which datasheets and how many copies, with points as the price, slot rules and army fixed effects. Its unit utilities are revealed preference with the "tax unit" bias modelled instead of named; units with similar utility but different reviewer letters are where the reviewers know something.per-list contents (not only counts)the corpus has them (CC BY 4.0, outside listhammer/); the repo keeps counts only, so it's a build-time read3 daysMedium-high: a second target with ~190 lists × ~10 choices, for all 16 armies, not only those with reviewers.
5Co-occurrence and a within-list Shapley value (recommenders; sabermetrics' replacement level). Lift of unit pairs in real lists (tools/colist.js counts already) finds roles and synergy (who goes with the Neurotyrant). Then value real winning lists with the ledger and split each list's worth among its units by Shapley value against a replacement-level unit. The support units the ledger misses (Neurolictor, Neurotyrant) are exactly those whose worth shows up only at list level.lists, the ledgeryes2 daysMedium: a principled route to "support" and "role" without a multiplier on a unit's own damage.
6Points changes as natural experiments (difference-in-differences). BSData's history gives each unit's points at each dataslate; the inclusion months (July to September) and reviewers' dates give before/after. A price cut that raises inclusion a lot says the unit was near the line. Calibrates the points exponent (pitfall 8).BSData history, monthly countsyes, thin (three months)1 to 2 daysMedium, growing with each dataslate.
7Army-level outcomes as an independent target (sports ratings). Sum the ledger over the best legal 2,000-point list per army and compare with public army win rates (36 armies). Thin, but it owes nothing to reviewers or popularity.army win ratespublic; the repo mentions them1 dayMedium as a sanity test of the line weights, not of perks.
8Targeted elicitation (active learning; RLHF-style pairwise labels). Show Jordan 60 to 100 pairwise questions chosen where the winning set disagrees most (e.g. the Fights First and Lone Operative carriers), with a Bradley–Terry fit. Thirty minutes of his time gives information exactly where the dials are unidentified.Jordan's timeyeshalf a day to buildMedium; label it as his view, not independent truth.

5. Top three for the next week, in plain English

  1. Prove the search beats luck before trusting any number it prints. Build the "fake world" test: invent reviewer letters from our own model with realistic disagreement, where every perk is worth exactly what the engine says, and run the whole T-662 search on 50 of them, plus fake perks given to random units. Only perk values that come out tighter and further from 1× than they do in the fake worlds get published; the others say "the reviewers can't tell us". Make the pull toward 1× the default, and show every result with a paired margin of error (on Tyranids a single setting's error is about 3.5 points).
  2. Make each dial mean what its label says, and fit game-wide dials on the whole game. A perk's dial should scale what the engine says that perk adds (with it minus without it), not the carrier's whole line, and dials should combine gently (in log space, no zero). Fit game-wide perks pooled over every army with a panel, with Tyranids weighted for their five lists, because on Tyranids alone several perks have three to seven carriers. Use Tyranids to find mechanisms; use the whole game to price them.
  3. Use the biggest dataset we have: the points themselves. Fit what the game charges for each perk across all ~1,100 units (a hedonic price list, our own numbers), add one "points exponent" to the ledger, and use both as the starting point (prior) for the dials. Then ask Jordan one product question: may tournament inclusion count as evidence of worth when estimating perks (not as a popularity bonus on the letter)? On Tyranids it already explains most of what the reviewers see, and a 70/30 blend with the ledger comes within about 1.5 points of the reviewers' own ceiling.