5 Oct 2026. An independent appraisal, written blind first (section 1, from the brief alone, before reading the repo's method; not revised afterwards), then compared with the points ledger and the T-662 plan (family dials, an exhaustive grid, a winning set, Tyranids first, the other armies as the exam). A report: no model, setting, benchmark, data file or frozen input changed. The cheap checks in section 2 ran read-only through docs/lab/ledger/ledger.js at Track B. Our own words and numbers; unit and keyword names only. Tournament counts from grimstat-corpus (https://github.com/N041M/grimstat-corpus, CC BY 4.0) via docs/data/inclusion.json.
In plain English
- The plan's building blocks are right. Valuing a unit through the dice engine, keeping the numbers we fit few, checking against shuffled letters, and keeping every setting that fits about as well as the best (not just one winner) are what a statistician would do.
- The weak point is how much the search can find by luck. I built a small copy of the planned search (8 perk families, 390,625 settings) and ran it on invented reviewer letters where every perk is worth exactly what the engine says. It still "found" improvements of up to 4 points, and the real letters give 3.4. On this toy, the plan can't yet tell a real perk value from noise. On Tyranids one setting's score also carries a margin of error of about 3.5 points just from which 52 units happen to exist.
- The "range" a perk gets can be a mirage. If almost every unit has a perk, or only three or four do, the search can't tell its worth, and the range it prints is just the grid it searched. That needs a check (fake perks given to random units) before anything goes on the site.
- The reviewers largely report what tournament players take. How often a unit appears in winning lists orders the Tyranid reviewers' consensus nearly as well as the reviewers agree with each other (86% against 89%; our ledger 78%). So fitting perk values to the reviewers partly teaches them popularity and army context, not what a perk is worth.
- Two things to add: (1) the game's own prices: regressing every unit's points on its perks across ~1,100 units gives what the designers charge for Stealth, Deep Strike and the rest, with far more data than 52 units; (2) a proper reviewer model that knows two of the five Tyranid lists were made for the previous edition and agree mostly with each other.
- Top three for the week: prove the search beats luck (fake-world test); make each dial scale what its perk adds rather than the carrier's whole line, and price perks on the whole game, not one army; use the points themselves as the big dataset and starting point.
1. Phase 1: the blind plan (written before reading the repo's method; not revised)
(a) How I would solve it
Name the target first. A reviewer's letter is not "damage per point". It is, roughly, "how much does taking this unit, at today's points, raise the win chance of a well-built competitive list of this army". That is a marginal value inside a list (a role filled, against a field, under a points budget), with a dose of "how often is it an auto-include". So I treat each unit as having one latent worth θ_u (continuous), and every data source as a noisy, biased measurement of it.
Two layers: a structural layer the engine supplies, a statistical layer that is small.
- Structural layer (the game, not fitted). For every unit, the dice engine turns the datasheet into a handful of capability numbers per point, each against a field weighted by what is actually played:
- Kill: expected damage / burst-kill probability against 4 to 6 canonical target classes (horde, elite infantry, monster/vehicle, character), weighted by the field's mix.
- Durability: points of enemy attacks needed to remove it, against the field's attackers (this is where T, W, Sv, invuln, FNP, Stealth, -1 to hit, -1 damage all live, mechanically).
- Board: OC × expected survival turns; movement/reach (M, Deep Strike, Scouts, Fly, Infiltrate) as "turn-1 / turn-2 objective reach probability" from a simple board model.
- Utility: actions, auras/buffs (valued by playing the buffed unit with and without), transport. This gives maybe 6 to 8 interpretable capability scores, not 341 variables. Perks enter the model through the engine (counterfactual: same unit minus Stealth, recompute), so a perk's worth is a derived quantity, not a free parameter.
- Statistical layer (fitted, few parameters). θ_u = Σ_k w_k · g(cap_k,u) + a_army + e_u, with monotone constraints (w ≥ 0), a concave g (diminishing returns), and a points term that makes it value-for-cost (log value minus log points, hedonic style). Letters enter through a many-facet ordinal model (cumulative probit / many-facet Rasch): each reviewer r has their own cut points τ_r (one reviewer's "B" is another's "C"), their own noise σ_r, and the ratings are correlated through a shared "consensus" factor (they read the same tournament results). Fit in a Bayesian way (Stan/PyMC or a hand-written Laplace approximation), with priors centred on "trust the engine": w equal across lines after scaling, residual e_u shrunk to 0.
- A second, independent target: revealed preference. The ~190 tournament lists are choices under a 2,000-point budget. A conditional-logit / discrete-choice model of "which units appear, how many copies" with price (points) as a regressor, sharing θ_u with the letters but with its own biases (army popularity, model availability, rule-of-three limits, synergy with the chosen detachment). Joint fitting gives a check: units the reviewers love but nobody takes (or vice versa) are diagnostic, not noise to be averaged away.
- A third, free and large data source: the points themselves. Game designers price ~1,000+ units across all armies with a hedonic model in their heads. Regressing log(points) on engine capabilities and perk indicators across every army gives the designer's implied price of each perk with n ≈ 1,000, not 50. A unit's tier should then largely be its residual: worth (our model) minus price. This is the cleanest way to get "true worth per perk" with real statistical power, and it's GW's numbers (points), not GW's text.
Perk values are reported as: (i) the engine's counterfactual change in capability (exact, no fitting), times (ii) the fitted line weight, with a posterior interval; for perks the engine can't play (Lone Operative, Scouts, Deep Strike, actions) one explicit parameter each, with a prior from the hedonic points regression. Interactions are attributed with Shapley values over the perks.
Validation, decided in advance. Fit on all armies with reviewers (not one), with leave-one-army-out cross-validation and leave-one-reviewer-out. Report held-out pairwise ordering accuracy against (a) the inter-rater ceiling computed the same way, (b) dumb baselines: points alone, damage-per-point alone, tournament inclusion alone, "last edition's letter". A model that can't beat "inclusion rate" on held-out armies doesn't explain anything. A permutation control (shuffle letters within army) gives the chance distribution of the whole search, not of one fit.
(b) Main pitfalls for any approach
- Tiny effective sample. ~50 Tyranid units, 5 reviewers who are not independent: the effective number of independent observations is close to 50 units, maybe fewer (S and D are a handful each). Any model with more than ~5 free parameters fit on one army is mostly fitting noise.
- Searching a huge space on a tiny target (garden of forking paths). The best of millions of combinations is optimistically biased; the optimism must be estimated by redoing the whole search inside resampling (nested CV / bootstrap of the search), not by bootstrapping the winner's score.
- Construct validity. Letters mix per-point efficiency, ceiling power, role scarcity (the army's only anti-tank), synergy with detachments and leaders, and fashion. A pure per-point efficiency model can't reach the ceiling and shouldn't try to by bending perk values.
- Reviewers aren't independent ground truth. They read the same tournament results and each other; their 84 to 89% agreement is partly shared information, so the "ceiling" overstates how much truth there is, and fitting them partly fits the meta's popularity.
- Non-identifiability. Many dials multiply the same line; correlated perks (big monsters have both high T and W and high points) make individual values unidentifiable from 50 units. Multiple different perk values fit equally well; the perk "values" then mean nothing, even if the ranking is right.
- Context: detachments, leaders, transports, army rules. A unit is worth more where its army lacks that role (diminishing returns within a list); single-unit valuation ignores this.
- Time: points and reviewers' dates differ; a letter given at old points is a different item.
- Selection of graded units: reviewers skip Legends, irrelevant or brand-new units; the missing are not random (mostly bad), which squashes the bottom of the scale.
- Coarse ordinal scale. Five letters, uneven use (few S, many B/C), ties: Spearman is unstable, distances between letters are not equal, and reviewers' cut points differ.
- Engine misspecification. Expected damage ignores variance, positioning, terrain, the missions' timing; a perk the engine can't see (Deep Strike) will be compensated by inflating something it can see.
- Goodhart / transfer. Tuning to one army's reviewers makes the system learn that army's quirks (Synapse, Shadow in the Warp) and attribute them to generic perks.
(c) Analogous problems and how they're solved
- Hedonic pricing (Rosen; house prices, computers, CPI quality adjustment): price regressed on attributes gives implicit attribute prices; here, GW's points are the price and our tier is the under/over-pricing residual. Solved with large-n regressions, log-linear forms, monotone splines.
- Conjoint analysis (marketing): part-worth utilities estimated from choices under designed variation, with hierarchical Bayes per respondent. Tournament lists are natural "choice sets" under a budget; part-worths come out of a discrete-choice model.
- Credit scorecards: an interpretable additive "points per attribute" table, binned features, monotonic constraints, fitted with logistic regression and validated out of time. This is exactly the product ("Stealth is worth +X") and shows that small, constrained additive models generalise better than tuned multiplicative stacks.
- Rater models: many-facet Rasch (Linacre), Thurstonian models, item response theory, Dawid–Skene: essay scoring, figure skating, wine judging (Hodgson showed judges are inconsistent with themselves), sensory panels (panelist severity and drift, Generalized Procrustes). They separate item quality from rater severity and noise, and estimate a latent consensus with uncertainty.
- Bradley–Terry / Plackett–Luce / Elo / TrueSkill: turn letters into pairwise comparisons; with covariates, BT becomes a feature-based ranker. Learning-to-rank (RankNet, LambdaMART) is the ML version, prone to overfitting at n = 50.
- Sabermetrics (WAR, VORP), fantasy auction values: value above a replacement-level option at the same role and price, park/era adjustments (here: army and detachment context, points dates), and decades of experience that scouts and stats disagree in predictable ways.
- Card-game analytics (17lands for Magic: GIH WR, ALSA, IWD; Hearthstone deck trackers): revealed preference (pick order) vs outcome (win rate when drawn) vs expert ratings; known biases (good players pick the good cards: selection bias; "played win rate" is confounded). Expert set reviews correlate only moderately with data, and both are informative.
- Video-game balance (MOBA pick/ban/win rates): the designers' tool is win rate conditioned on skill, plus pick rate; balance changes act as natural experiments.
- Structural estimation / simulation-based inference (economics; indirect inference, ABC): a simulator (the engine) with a few deep parameters fit to observed moments; identification is checked before estimation; priors and posterior predictive checks.
- Model multiplicity (Breiman's Rashomon effect, Rudin's Rashomon sets): many models fit equally well; report the set and what's stable across it (variable importance ranges), not a winner.
- Causal inference: points changes as natural experiments (difference-in-differences in inclusion rates before/after a points change) estimate how much a unit's worth exceeds its price.
- Recommenders / embeddings: co-occurrence of units in lists (matrix factorisation) reveals roles and synergy clusters that single-unit models miss.
2. Our method against the blind plan: where they agree, where they differ, who is right
Read for this section: design/tier-engine.md, design/tier-briefing.md, reports/tiers/ledger-steps.md (to step 39 and the Two tracks and Tyranids-first rules), reports/tiers/2026-10-03-ledger-crack.md, 2026-10-05-hedonic.md, 2026-10-05-dials.md, design/decisions.md (3 to 5 Oct), design/mission-notes.md (structure), docs/lab/ledger/ledger.js (header, SWITCHES, BASE), and, after writing section 1, the headers of the two earlier reviews (2026-10-04-review.md, 2026-10-05-critique.md) so as not to repeat them.
Where they agree
| Point | Blind plan | Our method | Verdict |
|---|---|---|---|
| The engine supplies the structure; perks act through it | capability numbers per point from the engine; perks valued by counterfactual | the ledger's lines (Kill as Burst, Soak, Score, Hold, Actions, Spawn, Presence, Carry) per 100 points; "1× = trust the dice engine" | Same idea. The ledger's lines are a better-worked version of my capability layer (Burst, preferred targets, the mission catalogue). |
| Few fitted numbers | ~5 to 8 weights on one army | a budget of about 8 knobs, the rest hibernated (3 Oct) | Agree. The T-662 grid quietly breaks this budget (below). |
| A chance control | permutation of letters, run through the whole search | shuffled-letters control | Agree in spirit; ours must shuffle through the whole pipeline (it does in ledgercrack.js; T-662 must too). |
| Report the set of good models, not a winner | Rashomon set | the "winning set" within the bootstrap band, weighted by fit | Agree. Right instinct; the implementation needs a calibration (section 3, pitfalls 1 and 2). |
| Score held-out armies and baselines | leave-one-army-out; baselines points, inclusion | the 14 never-fitted armies as the exam; list share and points as baselines | Agree. |
| Time | letters at old points are different items | units whose points changed flagged unfittable | Agree, but it doesn't go far enough (pitfall 6). |
| Ceiling | inter-rater ceiling the same way as the model's metric | ~89%, each reviewer against the others' consensus | Agree; I reproduce it (below). |
Where they differ, and who is right
- Fitting on one army (52 units) vs pooling. I'd fit game-wide numbers on every army with a panel, weighted by panel reliability, holding armies out. We fit on Tyranids and use the rest as an exam. Who is right: both, for different jobs. Tyranids-first is right for debugging mechanisms (five reviewers, clean residuals; the step log shows why). It is wrong for estimating game-wide perk values: a "game-wide" dial fit on Tyranids is estimated from the Tyranid carriers only, and many families have three to seven of them (in my rough family split: Fights First 3, Lone Operative 4, Scouts/Infiltrators 6, Precision/Indirect 7). That isn't a perk value; it's a per-unit adjustment with a game-wide label. And one "family", Synapse, exists only in this army. The pooled diagnosis (4 Oct) found flexible models don't travel; that is not evidence against pooling a dozen constrained dials, which is exactly the case pooling helps.
- An exhaustive grid with a fit-weighted winning set vs a likelihood with a prior. I'd fit an ordinal (many-facet) likelihood with a prior centred on the engine and read intervals off a posterior. T-662 enumerates tens of millions of combinations on a coarse grid and keeps those within noise of the best. Who is right: the grid is more transparent and easier to show Jordan, and that matters. But a fit-weighted grid is a pseudo-posterior whose prior is "uniform on {0, 0.5, 1, 2, 4} for every family", and for any dial the data can't see, the winning set's 10 to 90% range is just that prior coming back (my toy below: a family every unit carries shows exactly 20/20/20/20/20%). The "gentle pull toward 1×" reported with and without is the prior; it should be the default, not an option. Keep the grid as the engine, but turn its weight into a proper likelihood × prior, and calibrate it (pitfalls 1 and 2).
- What a dial multiplies. I'd value a perk as the engine's increment (line with the perk minus line without), times a weight. T-662 says a dial "multiplies the line it affects" (defensive perks scale survival). If that means the carrier's whole line, "Stealth ×2" means "every unit with Stealth has its whole Soak doubled", which mixes the perk with everything else the carrier has. That is the same flaw the 5 Oct critique found in the support multipliers (Winged Hive Tyrant ×4.1 for Synapse on its whole worth). I am right here, if the summary reflects the design.
- A second target. I'd fit reviewer letters and tournament inclusion jointly, with inclusion's biases modelled. Ours keeps popularity out of the score (Jordan's principle) and tried fitting the dials to inclusion (the "hedonic" report): no transfer. Who is right: the product principle is Jordan's call and sound for the published letter; but for estimating what a perk is worth, inclusion is a measurement, not a display, and the report shows it carries most of what the reviewers see (list share alone orders the Tyranid consensus right on 86%, the ledger on 78%). Also the repo's "hedonic" isn't hedonic pricing in the economist's sense (regressing price on attributes); it's fitting dials to popularity. The real hedonic regression (GW's points on attributes) hasn't been done (section 4, idea 1).
- Process. Our method pre-registers a prediction per step, names the aimed units, and keeps a decision log. That is better practice than my blind plan, which had none of it. We are right here. Keep it.
- Explanation. My plan's statistical layer explains less than the ledger's named lines and residual loop. For a site whose product is "show your working", ours is right on the shape of the explanation; the statistics must sit under it, not replace it.
Cheap checks I ran (Node, read-only, a few seconds each; no repo file changed)
All on docs/lab/ledger/tyranids.json at Track B (tools/ledger-benchmarks/2026-10-05-trackB.json) through Ledger.compute(), the five Tyranid lists, pairwise with ties in the model counted half, mean over the five lists. Tournament counts from grimstat-corpus (CC BY 4.0) via docs/data/inclusion.json.
- The yardstick's own noise. Track B scores 78.3% on my reimplementation of the five-list mean (the repo reports 77.7% with its own tie rule). Bootstrapping the 52 units (400 resamples) gives a standard error of 3.5 points on that mean. Most kept or reverted steps in the log moved it by 0.5 to 2 points. (Paired differences between two close settings have smaller errors, so the right number to report per step is a paired unit-bootstrap SE.)
- The ceiling. Each list against the mean of the other four: Auspex 90.6%, Hivemind 91.6%, Into the Hive Mind 86.3%, Maelstrom 95.0%, Astrategas 82.2% (mean 89.1%). Two reviewers who both put two units in different tiers order them the same way 92 to 96% of the time for most pairs; most disagreement is one-letter boundary calls.
- The reviewers come in two clusters, and they're dated. Maelstrom and Astrategas (both 10th edition, March and April 2026) agree with each other 87% (95% when both separate the pair), while Astrategas agrees with Into the Hive Mind only 70%. The five lists are closer to two or three independent views than five.
- Popularity is near the ceiling. List share orders the Tyranid consensus right on 85.8% (ρ 0.88), the ledger 77.6% (ρ 0.72); ledger and share correlate at ρ 0.66. A blend (rank of share × 0.7 + rank of ledger × 0.3) reaches 87.6% on the five lists, in sample, one parameter: the two are complementary.
- Points. The consensus barely tracks price (ρ 0.15 with the best form's points). Making Value more "per point" (Value per 100 × points^−0.25) lifts the five-list mean 78.3% → 79.9%; making it more absolute lowers it (points^0.5: 72.9%, points^1: 69.8%). One interpretable number, a points exponent, moves the yardstick as much as a typical dial.
- Who isn't graded. Within Tyranids, units graded by 3 or 4 lists average 2.8 to 2.9 (on S 5 … D 1) against 3.0 for units graded by all five: a mild "skip the weak ones" effect, small here.
- A toy of the T-662 grid, and a null world (section 3, pitfalls 1 and 2): eight crude families from the 4 Oct feature table (defensive, arrival, forward deployment, Lone Operative, Fights First, Synapse, damage keywords, Precision or Indirect), each on {0, 0.5, 1, 2, 4}, multiplying Track B's Value for carriers, all 390,625 combinations, 5 s a sweep.
- On the real letters: 1× 78.3% → best 81.7% (+3.4 points); 2,180 combinations within 2 points of the best, 8,820 within one bootstrap SE (3.5 points), and that band contains the all-1× setting itself.
- In eight simulated worlds where every dial is truly 1× (letters drawn from the ledger plus a shared and a personal noise, cut at each reviewer's own letter shares, tuned so a reviewer agrees with the others' consensus 84 to 89%, as the real ones do, and 1× scores 74 to 84%), the same grid "finds" +0.0, +0.7, +1.0, +1.2, +1.6, +1.8, +3.8 and +4.2 points. The real +3.4 sits at about the 75th percentile of the null. The null winning sets also push dials away from 1× (one world puts 90% of its set at Synapse 0.5; another all of it at forward deployment ≥ 2; another 84% at damage keywords 2), so a winning-set range that excludes 1× is not, on its own, evidence.
- It is a toy (Value, not a line, is multiplied; my families are rough; eight families, not 15), but it says the plan's controls must be calibrated this way before any number from it is published.
3. Pitfalls we are missing, ranked by how much each could mislead us
- The grid can manufacture the gain it reports (selection on a tiny target). Effective sample: ~52 units, five correlated lists. With tens of millions of combinations the best is biased upward by about the size of a real effect (the toy null: 0 to 4 points from eight families; more families, more bias). The shuffled-letters control doesn't measure this well: shuffled letters start at ~50%, the real ones at ~78%, and gains near a ceiling aren't comparable to gains from chance level. The 3 Oct crack already showed the problem with 11 knobs (ρ 0.63 on shuffled letters, 0.70 on real), and the 5 Oct dials report admits one more layer (the four dials were chosen from one-at-a-time sweeps on the same 52 units). Check: a calibrated null world (simulation-based calibration): simulate letters from the 1× ledger with reviewer-like noise matched to the real agreement, run the entire T-662 pipeline (grid, band, winning set, family split, Tyranid shortlist) on 50 such worlds, and report the real gain's percentile. Publish nothing unless it clears the null's 95th percentile. Also run the pipeline inside unit folds (nested), and on the Track B history (the choice of Track B's members is itself a forking path).
- The winning set's ranges are mostly the grid's prior (identification). A family carried by almost every unit is a copy of its line's weight; rank metrics ignore overall scale; a family with 3 to 7 Tyranid carriers is fitted by those few units' letters. The toy shows the signature: a universal family comes back 20/20/20/20/20 (a 10 to 90% range of 0 to 4×, which reads like a finding), and few-carrier families come back wide and skewed in null worlds too. Fix: for each family print carriers (per army), the winning set's concentration against (a) the grid's uniform prior and (b) the null worlds' concentration, and the units that move it (drop-one-unit influence). Add knockoff families (the same number of carriers, assigned at random within unit type) to the grid; a real family whose range is no tighter than its knockoffs' is "not identified", and the site says so instead of showing a range.
- Dial semantics: a dial on the carrier's line is not a perk's value. If "defensive perks × 2" doubles a carrier's whole Soak, the dial confounds the perk with the carrier's other traits and stacks (a Stealth + Feel No Pain + invulnerable unit gets three multipliers on one line). Fix: a dial scales the perk's engine increment, Δ = line with the perk − line with it removed in the engine; perks the engine can't remove (Deep Strike, Scouts, Lone Operative, actions) enter as an additive credit in points per 100, with the carrier count printed beside it.
- Multiplicative stacking and the grid's corners. {0, 4} on several families multiplies to 0 or 64× on units carrying several. 0 isn't "a smaller effect"; it is a different hypothesis (the perk is worthless, or the line is switched off), and a grid that mixes them reports a range spanning "worthless to quadruple". Fix: dials in log space with a prior (log-dial ~ Normal(0, 0.35), so ×0.5 and ×2 are each about 2 SD away); drop 0 from the grid or treat it as a separate model; report the largest combined multiplier any unit gets in the winning set.
- What the letters measure (construct). On Tyranids, list share alone is within about 3 points of the reviewers' own ceiling. Reviewers are, to a large degree, reporting what the meta takes. A dial fitted to them learns "what the meta rewards", including army context (Synapse, the best detachment, the role a unit fills in a list), and labels it as the worth of a generic perk. That's the opposite of "true worth per perk". Check: fit the same grid with list share as the target (one army out at a time) and with the consensus residual on list share as the target; a perk value that survives only on raw letters is popularity, not mechanism. Report the partial agreement with share held fixed (the hedonic report's 0.33 on Tyranids) as the headline of "what the ledger adds".
- Reviewers are not five independent truths, and two of them graded a different game. Maelstrom and Astrategas are 10th edition and cluster together. Flagging units whose points changed misses rule changes, and the pairwise mean double-counts the 10th-edition view. Fix: a Thurstonian rater model with a shared factor and a date/edition facet; report every fit on the three 11th-edition lists alone beside the five; compute the ceiling within the 11th-edition cluster. Weight lists by estimated reliability, not one each.
- Too little power for the exam. The 14 armies carry 1 to 3 lists each, some near or below chance for any model (Thousand Sons 38.5%, Space Wolves 51%). A third of them held out is about five panels, each with a unit-bootstrap SE of roughly 4 to 6 points, pooled perhaps 2 to 2.5. Choosing among a Tyranid shortlist whose members differ by under 2 points on those armies is choosing by noise. Fix: before running it, compute the minimum detectable difference on the held-out third (paired bootstrap over units and lists) and say in advance that the shortlist is chosen by prior (closest to 1×) unless a member beats it by more than that.
- Points-to-tier non-linearity and replacement level. Value per 100 points assumes tier ∝ worth per point. The check above says the reviewers lean even further toward cheap units (points^−0.25 helps 1.6 points), which is what a replacement-level model predicts (a cheap unit that fills a slot frees points). One exponent or a replacement level is one interpretable parameter; fit it before family dials, or the families will soak it up (cheap units carry certain perks more often).
- Best-of-many selection inside the ledger. Each unit is scored at its best solo form over sizes, loadouts and detachments. Units with more options get a higher maximum by chance of measurement, not worth. Check: regress the consensus residual on the number of forms per unit; if it's negative, score the form players field (the mode) or shrink the max.
- A metric with a surrogate. The search steers by Spearman or an ordinal likelihood and reports pairwise. Fine, but pick one surrogate in advance; trying both and keeping the better is one more fork. Ties: the model never ties, the reviewers often do; keep the ceiling computed on exactly the model's rule.
4. Directions not yet tried
| # | Idea (field it comes from) | What it needs | Do we have it | Effort | Expected payoff |
|---|---|---|---|---|---|
| 1 | The designers' hedonic price (economics, Rosen). Regress log(points) for every fieldable unit in all 36 armies (~1,100) on the engine's capabilities (Kill, Soak, OC, Move) and perk flags (Stealth, Feel No Pain, Deep Strike, Scouts, Lone Operative, Fights First, Lethal/Sustained/Devastating, Anti, Precision, Indirect), monotone, with army fixed effects. The coefficients are what the game's designers charge for each perk, with standard errors from n ≈ 1,100, not 52. A unit's residual is its mispricing, which is what a tier list is about. | points and datasheets | yes (build/units.json, BSData; points are numbers the ledger already reads; our coefficients are our own) | 1 to 2 days | High. The first perk values with real statistical power, independent of reviewers; a natural prior for every dial; a headline the site can show ("the game charges about X points for Stealth on infantry; our engine says it's worth Y"). |
| 2 | Calibrated null worlds and knockoff families (simulation-based calibration, model-X knockoffs). As in my toy: letters simulated from the 1× ledger with reviewer-like noise; fake families with matched carrier counts. Run the whole T-662 pipeline on them. | the ledger, the letters | yes | half a day (the toy is most of it) | High as a guard: it turns "within the band" into "beats chance by this much", per perk. |
| 3 | A many-facet Thurstonian model of the reviewers (psychometrics; Dawid–Skene; sensory panels). Each unit has a latent worth = engine lines × dials + residual; each list has its own cut points, noise and an edition facet; a shared factor for what all reviewers see. Fit by Laplace approximation or a small sampler in Node, pooled over the 16 armies with panels. Gives honest posterior intervals for each dial, reviewer reliabilities (who to trust), and replaces the fit-weighted grid. | the panels | yes | 2 to 3 days | High methodologically; medium in agreement (it fixes the arithmetic, not the missing lines). |
| 4 | A discrete-choice model of list building (conjoint, industrial organisation). Each tournament list is a basket chosen under a 2,000-point budget: a conditional logit on which datasheets and how many copies, with points as the price, slot rules and army fixed effects. Its unit utilities are revealed preference with the "tax unit" bias modelled instead of named; units with similar utility but different reviewer letters are where the reviewers know something. | per-list contents (not only counts) | the corpus has them (CC BY 4.0, outside listhammer/); the repo keeps counts only, so it's a build-time read | 3 days | Medium-high: a second target with ~190 lists × ~10 choices, for all 16 armies, not only those with reviewers. |
| 5 | Co-occurrence and a within-list Shapley value (recommenders; sabermetrics' replacement level). Lift of unit pairs in real lists (tools/colist.js counts already) finds roles and synergy (who goes with the Neurotyrant). Then value real winning lists with the ledger and split each list's worth among its units by Shapley value against a replacement-level unit. The support units the ledger misses (Neurolictor, Neurotyrant) are exactly those whose worth shows up only at list level. | lists, the ledger | yes | 2 days | Medium: a principled route to "support" and "role" without a multiplier on a unit's own damage. |
| 6 | Points changes as natural experiments (difference-in-differences). BSData's history gives each unit's points at each dataslate; the inclusion months (July to September) and reviewers' dates give before/after. A price cut that raises inclusion a lot says the unit was near the line. Calibrates the points exponent (pitfall 8). | BSData history, monthly counts | yes, thin (three months) | 1 to 2 days | Medium, growing with each dataslate. |
| 7 | Army-level outcomes as an independent target (sports ratings). Sum the ledger over the best legal 2,000-point list per army and compare with public army win rates (36 armies). Thin, but it owes nothing to reviewers or popularity. | army win rates | public; the repo mentions them | 1 day | Medium as a sanity test of the line weights, not of perks. |
| 8 | Targeted elicitation (active learning; RLHF-style pairwise labels). Show Jordan 60 to 100 pairwise questions chosen where the winning set disagrees most (e.g. the Fights First and Lone Operative carriers), with a Bradley–Terry fit. Thirty minutes of his time gives information exactly where the dials are unidentified. | Jordan's time | yes | half a day to build | Medium; label it as his view, not independent truth. |
5. Top three for the next week, in plain English
- Prove the search beats luck before trusting any number it prints. Build the "fake world" test: invent reviewer letters from our own model with realistic disagreement, where every perk is worth exactly what the engine says, and run the whole T-662 search on 50 of them, plus fake perks given to random units. Only perk values that come out tighter and further from 1× than they do in the fake worlds get published; the others say "the reviewers can't tell us". Make the pull toward 1× the default, and show every result with a paired margin of error (on Tyranids a single setting's error is about 3.5 points).
- Make each dial mean what its label says, and fit game-wide dials on the whole game. A perk's dial should scale what the engine says that perk adds (with it minus without it), not the carrier's whole line, and dials should combine gently (in log space, no zero). Fit game-wide perks pooled over every army with a panel, with Tyranids weighted for their five lists, because on Tyranids alone several perks have three to seven carriers. Use Tyranids to find mechanisms; use the whole game to price them.
- Use the biggest dataset we have: the points themselves. Fit what the game charges for each perk across all ~1,100 units (a hedonic price list, our own numbers), add one "points exponent" to the ledger, and use both as the starting point (prior) for the dials. Then ask Jordan one product question: may tournament inclusion count as evidence of worth when estimating perks (not as a popularity bonus on the letter)? On Tyranids it already explains most of what the reviewers see, and a 70/30 blend with the ledger comes within about 1.5 points of the reviewers' own ceiling.