4 Oct 2026. A review: no model, number, setting, benchmark, data file or frozen input changed. Every number below was run read-only on the confirmed baseline, tools/ledger-benchmarks/2026-10-04-step26-baseline.json (burst Kill + threat-gated Soak: seven headline lists 72.3% pairwise, ρ 0.520; 16 armies 59.3%; the 14 never-fitted 57.8%). Our own numbers and words; unit names only. Guesses are marked as guesses. The scratch scripts that produced the sweeps are described in section 8 so anyone can repeat them with the repo's tools.
The short answer
- The ledger today is, in practice, "Kill per point". Kill is about 91% of a median unit's Value in both armies (Soak 2%, Hold 2–4%, Score 1–2%, the rest under 1%). Value and its Kill part order the units almost identically (ρ 0.91 on Tyranids, 0.97 on Space Marines). The other six lines and most sliders act as tie-breakers.
- What the other lines add doesn't carry to new armies. Ranked on its Kill line alone, the ledger scores 69.0% on the seven headline lists (3.3 points under the full ledger), but 58.4% on the 14 armies it was never tuned on, better than the full ledger's 57.8%. So, so far, every line and slider beyond Kill has learnt our seven lists rather than the game.
- The reviewers are seeing roughly what tournament players see. How often a unit appears in the 23 Tyranid tournament lists of July to September orders the reviewers' consensus right on 86% of pairs, close to the reviewers' agreement with each other (about 90%), against the ledger's 75–76%. Over all 16 armies, list share scores 70.9% against the ledger's 59.4%, and wins in 14 of the 16 (the exceptions: Space Marines, with only 2 tournament lists, and Blood Angels, a near-tie). The gap between the ledger and the reviewers is real game knowledge that players share (a unit's role in a list, missions, support, being hard to pin down), or shared reputation; this data can't tell the two apart.
- That explains why every model misses the same units. The Neurolictor (in 74% of Tyranid lists, graded S by every reviewer, 28th by us), the Neurotyrant (74%, 41st by us), Gargoyles and Tyrant Guard are worth what they do for the rest of the list. The Zoanthropes, Von Ryan's Leapers and the Winged Hive Tyrant are top of our list because their Kill per point is the highest. A model built on one unit's own dice will keep missing in both directions until it prices support, scoring and role.
- There are more tuned numbers than the data can support. With 52 Tyranid and 84 Space Marine units, five fairly agreeing Tyranid reviewers and two Space Marine ones, the data supports about 6 to 9 freely fitted numbers. The ledger has about 20 that were set by searching against these same lists: 7 line weights, the melee lands scale, 11 rule multipliers and the Biovore dial. Some multipliers were fitted on 3 units (Fights First ×2.64, healing ×2.61).
- The best single change measured today is a step back toward neutral: setting the mortal-wounds multiplier from 1.69 back to 1. It gives +0.8 points on the headline (± 0.6), +1.1 on the 14 never-fitted armies (± 0.8; better in 95% of resamples), and the largest drop in the typical rank gap of any knob (17.3 → 16.8 places). It moves the Zoanthropes from 1st to 9th (consensus 19th). Moving four of the fitted multipliers halfway or all the way back to 1 gives +1.0 on the headline and +1.9 on the 14 armies. That is the classic sign of over-fitting being undone.
- Second best: the melee lands scale from 1.5 to 1.25. It was set at the search's upper limit before Kill became Burst. At 1.25: +0.9 on the headline (± 0.7), every Tyranid list up (Tyrannofex 26th → 22nd, Neurolictor 28th → 24th), flat on the 14 armies.
- What would most likely close the gap is a second target for the weights: tournament inclusion, a revealed preference fitted the way economists price a house's features from sale prices, with the reviewers kept as the check. Our reviewers are, in effect, already a noisy copy of that signal. The mission-card model being built (T-647) is the other route, and the only one that explains why, but steps 12 and 27 show its "alive to score" term must be per point, or it rewards size.
- Tonight, in order: (1) the mortal-wounds multiplier to 1; (2) the other fitted multipliers shrunk toward 1 by one pre-set factor; (3) lands to 1.25; (4) freeze every knob that moves less than a quarter of a point; (5) role scarcity built for Space Marines (it is silently off there); then the larger cards (sections 5 and 6).
1. The problem, stated precisely
What is estimated. For each unit of an army, one number: its worth to a 2,000-point list per point it costs, in its best legal form (size, loadout, detachment, alone or led), against one snapshot of the rules (BSData 374f505) and of the meta (a typical enemy army built from top-quarter tournament lists). The tier is that number cut into bands. Only the order within an army is judged.
From what. The engine's measurements per form: Kill (enemy points removed in one activation, since step 25 as Burst: only targets finished), Soak (enemy points that must be spent to remove it, since step 26 scaled by its own threat), OC per point, actions, held objectives (OC × Soak), spawned points, bodies for the unit-count cards, carrying, and rule multipliers per kind of rule. A linear sum of the lines, each divided by its army's median, times the multipliers.
Judged how. Against reviewers' tier letters: of the pairs a reviewer puts in different tiers, the share we order the same way (50% chance); beside it ρ, pairs two or more tiers apart, the movement view (tools/stepmoves.js), held-out units (tools/diagnose.js), a re-fit (tools/ledgerrefit.js) and the 14 armies never tuned on.
The information budget.
| Units | Lists | Cross-tier pairs per list | Reviewers' own agreement | |
|---|---|---|---|---|
| Tyranids | 52 (44–52 graded per list) | 5 | 744–1,044 | 75–96% pairwise between lists; ceiling 90.4% |
| Space Marines | 84 | 2 | 2,375–2,440 | 91.6% between the two (on pairs both split) |
| Headline | 136 | 7 | about 9,000 pairs, but highly dependent | |
| All 16 armies | about 600 graded | 1–3 per army outside the headline |
- Thousands of pairs sound like a lot, but they come from 136 units graded on five or six steps. A tier letter carries about 2 bits. Five Tyranid reviewers who agree 85–95% of the time are worth perhaps one and a half independent lists, not five. Realistically the headline holds a few hundred bits.
- The rule of thumb the repo already adopted (one free number per 15–20 graded units) gives about 3 for Tyranids and 4–5 for Space Marines: 7 to 9 in all. That is the budget for anything fitted to these lists, including numbers set by a search, a hand dial chosen while looking at the letters, or a switch kept because it raised the headline.
- The noise floor of a one-knob change (paired bootstrap over units, 300 resamples, this review): the standard error of a change in mean headline pairwise is 0.2 to 1.4 points, about 0.6 for a knob that moves a dozen units. The keep rule's bar of 1 point is therefore about 1.5 standard errors. A change of pure noise clears it roughly 1 time in 10 to 20, and passes test (b) of the provisional rule (more units closer than farther) about half the time.
2. Methods from other fields, compared
design/ranking-methods.md (3 Oct) and design/results-methods.md (T-631) surveyed most of these. Here each is set against what has actually been built and measured since, and graded on whether it could plausibly close the 72% → 90% gap with this data.
| Field | What it does | How it maps here | What we already do | What we lack | Likely to close the gap? |
|---|---|---|---|---|---|
| Sports value over replacement (WAR, VORP) | Worth = output above what a freely available player gives, per game or per salary | Worth = Value above the cheapest unit that does the same job, per point | Lines are divided by the army's median (a crude "average" baseline, not "replacement") | A replacement level per job: a Neurolictor is worth what the list loses if the next-best Synapse/action piece fills its slot | Partly. It is the honest form of "scarce answers" (step 19 failed because Tyranids aren't short of Kill). Replacement on support and scoring jobs is untried. |
| Plus-minus with ridge (RAPM, RAPTOR) | Regress team results on who was on the floor; shrink each player toward a prior built from his box-score stats | Lists' placings regressed on the units in them, shrunk toward the ledger | Nothing fitted to results; T-631 designed it (the "two layers") | The outcome layer | Not alone. T-631 measured army explaining 2.4% of placings and only 41 unit rows clearing a loose bar. It is too noisy to grade units; useful for one or two line weights. |
| Fantasy auction values (points above replacement) | Projected stats → points above the replacement at each position → dollars at a fixed budget | Exactly our ledger's design: lines → worth in points → per point | The ledger is this, with points as the budget | Position-specific replacement (as above), and a check that values sum to the budget | Already the frame; the missing part is the replacement level. |
| Paired comparisons (Bradley–Terry, Thurstone, Elo/TrueSkill) | Latent strength from who beat whom, with a sharpness per judge | Each reviewer's cross-tier pairs as comparisons | Yes: tools/ledgerfit.js is a Bradley–Terry fit with a sharpness per list and a reliability weight per reviewer, ridge to the prior | Ties used as information (an ordinal "same tier" likelihood); reviewer bias per unit type | No. It is the right loss and it is in place; a better loss won't add information the features don't carry (T-632: every model lands at about 70% on held-out units). |
| Learning to rank (RankNet, LambdaMART) | Fit a scorer with a pairwise or listwise loss, often with trees | The flexible models of T-632/T-635/T-636 | Yes, tested: trees, kernels, neighbours, 341 variables | Nothing | No, measured: they memorise 52 units and fall to 64–73% held out; across armies 58–63%. |
| Hedonic pricing | A product's price = the sum of implicit prices of its features; prices read from market transactions | Market = tournament list-building. A unit's inclusion rate (or its share of points spent) regressed on its ledger lines, within each army | Not done. List share has been a yardstick (step 1: held-out ρ 0.41) | The fit itself: weights chosen to explain what players buy, across ~11 armies with 12+ tournament lists | The most likely single gain (section 3: list share already orders the Tyranid consensus 86% right). The weights would learn from about 900 unit rows instead of 136. |
| Conjoint / discrete choice | People choose between bundles; a conditional logit reads each feature's worth | Each tournament list is a bundle chosen under a 2,000-point budget; units compete for points within an army | Nothing | A choice model with the budget: each unit's chance of being picked against the army's other options | Strong in principle, the right shape for "worth per point in a list". Heavier to build than hedonic; do hedonic first. |
| Expert elicitation and aggregation (Dawid–Skene, many-facet Rasch) | Estimate each rater's bias and noise together with the item's true score | Reviewer severity (Auspex's 15 S on Space Marines against Tobias's 6), and per-reviewer leanings | Reliability weights per list; pairwise ignores severity by design | Per-reviewer cut points (ordinal model); a "consensus" built from the model rather than the mean letter | Small. The reviewers agree at about 90%, so modelling them better buys a point at most. It mainly sharpens the yardstick. |
| Bayesian hierarchical models with partial pooling | Each group's parameters drawn from a shared distribution; thin groups shrink to the mean | A rule multiplier fitted on 3 Tyranid carriers borrows strength from 75 carriers across 16 armies; army-specific weights shrink to shared ones | Not done; weights are shared and fixed, rule multipliers are Tyranid/Marine fits | Shrinkage of multipliers toward 1 (or toward a 16-army estimate) | Yes, cheaply. Measured in this review: moving fitted multipliers toward 1 gains on both the headline and the 14 armies (section 4). |
| Structural / game-simulation models | Value from simulating the game (e.g. win probability added in chess or baseball) | The mission-card model (T-647): each primary and secondary scored per unit; the charge model; Kill over the game | Several parts built; steps 12, 18 and 27 tried game-length terms and lost | A unit's worth as the change in a typical list's expected VP when it is swapped in | The only route to the WHY, highest ceiling, highest risk. Steps 12 and 27 show the trap: a survival term with fixed enemy fire rewards size. Step 27b's per-point fire is the right ruler. |
Which would most likely close the gap, given this data.
- Hedonic pricing on tournament inclusion, judged on the reviewers. Expected: the line weights move toward what players pay for (guess: Score/Hold and support up, raw Kill down), with a gain on the 14 armies. It is the one source of more labelled units, and the evidence below says it is the same signal the reviewers carry.
- Partial pooling (shrinkage) of the rule multipliers. Measured: +1 to +2 points on never-fitted armies.
- The structural model (T-647), for the units no datasheet number explains, on condition that it is judged on the 14 armies first and its survival term is per point.
- Better loss functions and flexible learners: already exhausted (T-632, T-635, T-636).
3. What we do well, and what is wrong or missing
Done well
- The frame is the right one. Worth in the game's own currency, line by line, per point, against a stated meta. Fantasy auction pricing works the same way, and it is what lets a player check a number.
- The discipline is unusually good: one change per step, a written prediction, scrambled letters, a re-fit, held-out units, a movement view, the 14 never-fitted armies as a gauge, switches kept so every step can be unwound. Most teams in sports analytics don't do half of this.
- Bradley–Terry with reviewer reliability and ridge is the textbook fitter, and the unit-weight reference ("Dawes") is reported beside it.
- The diagnoses (T-632, T-635, T-636) were the right experiments and their reading is correct: the features, not the method, are the limit.
Wrong or missing
- The target is closer to "what players take" than "worth from the rules". Read-only numbers from this review (docs/data/inclusion.json, July to September):
| | Tournament lists | Ledger vs reviewers | List share vs reviewers | |---|---|---|---| | Tyranids | 23 | 74.8% pairwise against the consensus; ρ 0.66 | 85.8%; ρ 0.88 | | Space Marines | 2 | 64.1% | 57.5% (2 lists: no signal) | | 16 armies, mean over each army's lists | 1–23 | 59.4% | 70.9% (14 of 16 better) | | The 11 armies with 12+ tournament lists | | 62.0% | 71.3% |
On Tyranids, list share and the ledger agree with each other at ρ 0.61; an equal blend of the two (84.7%) is no better than list share alone. Either the reviewers report the meta, or the meta follows the reviewers. Either way fitting the reviewers is very nearly fitting popularity, and a rules-only model judged against them has a ceiling well under the reviewers' 90% (guess: high 70s on Tyranids). That doesn't make the ledger wrong; it makes the yardstick ambiguous. The repo's own plan (design/ranking-methods.md, move 4: popularity as a nuisance term) was never run.
- Value is Kill; the other lines are decoration with fitted weights. Median shares of Value: Kill 91%, Soak 2%, Hold 2–4%, Score 1–2%. Halving or doubling the Kill weight changes the headline by 0.2–0.4 points: the weight sits on a flat ridge, so it is not identified, and neither are the others relative to it. Kill alone transfers better to the 14 armies (58.4%) than the full ledger (57.8%).
- Too many fitted numbers for the budget (about 20 against 7–9). The rule multipliers were set by the 3 Oct joint search on a box from 1 to 3, on the headline's best rows, with very few carriers: Fights First 3, healing 3, Lone Operative 5, Better Overwatch 6, mortal wounds 12, buffs 15, command points 15, Synapse 18. A multiplier fitted on 3 to 6 units is noise by the repo's own rule ("an ability fitted on 3 carriers is noise"). The measured pattern confirms it: pulling them toward 1 helps the never-fitted armies.
- Some knobs are stale: they were fitted to an earlier model. The melee lands scale sits at 1.5, the search's upper limit, fitted before Kill became Burst (step 25). Burst already counts only finished targets, so the old lift on melee is now counted on top. Measured: 1.25 is better on every Tyranid list.
- The re-fit can't really move anything. ledgerfit's ridge pulls each weight toward the base benchmark with λ 0.01 per relative size, and the rule multipliers are hibernated. A re-fit therefore re-confirms the old weights (wSoak 14.97 → 15.23), and "kept after a re-fit" is a weaker test than it sounds. Worse, the re-fitted base scores below the as-is weights (71.9% against 72.3%), so the as-is weights carry some luck. ledgerrefit.js also refuses the confirmed base (it requires every step switch off), so a step on top of step 26 must be run as
--on "burstKill+threatSoak;burstKill+threatSoak+<step>". - The two headline armies don't run the same model. Space Marines' file carries no role scarcity and no Soak effects (their meta has no medianFactor), so those switches do nothing there; Tyranids and the other 14 armies have them. Role scarcity is worth 4.7 points on Tyranids and 2.0 on the 14 armies. The weakest headline army is running without one of the strongest parts. (design/ledger.md section 16 notes "Space Marines' file carries no scarcity"; it has not been fixed.)
- The Biovore dial has a side effect. Placed-anywhere 523 also lifts the Harpy (its spawned Spore Mines are placed anywhere too): the Harpy is 30th of 52 against a consensus of 50th, one of our worst "too high" units; with the dial at 1 it drops 13 places. A hand fix for one unit should be scoped to that unit.
- The yardstick over-weights one army. Five of the seven lists are Tyranid and they agree with each other, so the mean is roughly 70% Tyranid. The per-army guard in the keep rule handles falls, but gains are still judged mostly on Tyranids.
- The provisional keep rule lets noise stack. Its tests (a) and (b) are close to coin flips for a change of pure noise, and (c) only stops falls past noise (Tyranids 0.039, Space Marines 0.098 in ρ: large). The confirmation step protects against this, but only if confirmation is judged on data the steps didn't see: the 14 armies, and held-out units.
- "Per point" is right, and the joined-unit handling is honest, but "best form across detachments" is generous to units that are good in one detachment nobody plays; that is stated in design/ledger.md and accepted. Not a priority.
- The "same units missed by every model" signal is correct and actionable: what those units have in common is support, scoring and list role (Neurolictor, Neurotyrant, Gargoyles, Tyrant Guard, Termagants), or being pure Kill per point that the reviewers discount (Zoanthropes, Von Ryan's Leapers, Winged Hive Tyrant, Deathleaper). More datasheet variables cannot fix that (T-636); a model of the list and the mission can.
4. The knobs (T-649)
How measured. One knob at a time from the step-26 base, every other setting held, through docs/lab/ledger/ledger.js compute() exactly as tools/ledgerstep.js applies --on key=value. For each value: the change in mean headline pairwise (7 lists), each army's change, the movement view (tools/stepmoves.js movesOf: closer/farther, typical gap, in band, mean places moved), and the change in mean pairwise over the 14 never-fitted armies (as tools/ledgercarry.js scores them). Swing = the largest change in headline pairwise over the knob's sweep. Paired standard errors (bootstrap over units) for the promising settings are in the second table. Base: headline 72.29%, ρ 0.520, Tyranids 75.6%, Space Marines 64.0%, 16 armies 59.34%, the 14 57.85%, typical gap 17.32 places, 54 of 136 units in the reviewers' tier.
4.1 Inventory, with sensitivity
"Fitted" means set by a search or fit against the headline lists (or their predecessors). "Reasoned" means set before looking. "Call" means one of Jordan's calls. Sensible ranges are the reviewer's judgement.
| Knob | In the game | Now | Sensible range | Set by | Swing (headline) | Best measured (headline; the 14) | Units it moves most |
|---|---|---|---|---|---|---|---|
| wKill | Exchange rate of a point killed in one activation | 498 | 250–1000 (the scale is the other lines') | Fitted (joint search, rescaled) | 9.5 at 0; 0.2–0.4 within ×½–×2 | none: flat ridge | at 250: Tactical Squad, Neurogaunts, Rhino up; at 2000 Biovores −27 |
| wSoak | A point of enemy fire absorbed (threat-gated) | 15.0 | 0–120 | Fitted | 0.7 | 120: +0.71; −0.11 | Land Raiders, Terminator Assault Squad, Repulsor up; Tyrannofex |
| wScore | OC per point | 9.65 | 0–40 | Fitted | 0.4 | none | Neurogaunts, Termagants, Tactical Squad |
| wActions | Actions per point | 1.31 | 0–20 | Fitted (at the floor) | 0.1 | none | almost nothing |
| wHold | OC × Soak: holding an objective | 18.9 | 0–75 | Fitted | 0.6 | 75.6: +0.21; −0.34 | Tactical Squad, Rhino, Neurogaunts up; Outriders, Intercessors down at 0 |
| wPresence | Bodies for unit-count cards | 2.52 | 0–10 | Fitted | 1.9 | none (0 costs −1.9: Biovores) | Biovores, Harpy |
| wSpawn | Points placed a round | 15.7 | 0–60 | Fitted | 0.2 | 0: +0.16; 0 | Tervigon, Biovores |
| wCarry | Models carried forward | 0 | 0–1 | Call (hibernated) | 5.1 | none (every value loses) | Rhino, Drop Pod, Tyrannocyte, Land Raider Crusader |
| prem | Premium on Kill | 1 | – | Fixed | same as wKill | – | (duplicates wKill: retire) |
| lands | Scale on the charge model's chance melee lands | 1.5 | 1–1.5 | Fitted (at the search's upper bound, before Burst) | 1.3 | 1.25: +0.90; +0.08 (Tyranids +1.3) | Judiciar, Bladeguard, Terminators down; Tyrannofex, Neurolictor, Broodlord up |
| anywhere | Spawned units placed anywhere (Presence) | 523 | 1 or a unit-scoped dial | Call (Biovore dial) | 1.9 | none (300: −0.06) | Biovores, Harpy (side effect) |
| reach | Own bodies counted × when it arrives far forward | 1 | 1–2 | Fitted (at the floor) | 0.1 | – | Neurolictor, Raveners |
| tunnel | Bodies per tunnel marker | 0 | 0–1 | Call | 0.0 | – | none (no form carries tunnels at the ranking rows) |
| threat | Old threat weighting on Soak | 0 | – | Superseded by threatSoak | 0.0 | – | none while threatSoak is on: retire |
| solo | Rank units alone (not joined) | on | on | Reasoned | 3.4 (off: −3.4; the 14: −5.6) | keep on | every Space Marine character |
| scarcity | Role scarcity on Kill | on | on | Reasoned (measurement) | 3.4 (off: Tyranids −4.7; the 14: −2.0) | keep on; missing for Space Marines | Exocrine, Tyrannofex, Trygon (down when off) |
| soakfx | Exposure and own healing on Soak | on | on | Reasoned | 0.0 | – | The Red Terror |
| beta | How sharply fire picks its best targets | 6 | 3–∞ | Reasoned | 0.4 | 12: +0.15; −0.04 | Trygon, Tyrannofex, Mawloc |
| combo, cutS–cutC, raw | Display (combos shown, band cuts, raw lines) | – | – | – | 0 on pairwise | – | bands only |
| fireB | Enemy fire on each unit a turn | 167 | per point (step 27b) | Reasoned | inactive at base | – | only with killOverGame or secondaryVP |
| r.mortal | Mortal wounds outside attacks | 1.69 | 1–1.5 | Fitted (12 headline carriers) | 2.3 | 1: +0.79; +1.13 | Zoanthropes, Toxicrene, Mawloc, Chaplain with Jump Pack, Brutalis down |
| r.overwatch | Better Overwatch | 1 | 0.75–1.5 | Call | 1.0 | 0.75: +0.18 (Marines +1.4); −0.01 | Firestrike Servo-Turrets, Invictor, Hammerfall Bunker |
| r.first | Fights First (on melee Kill) | 2.64 | 1–1.5 | Fitted (3 carriers) | 0.3 | 1.82: +0.08; −0.11 (at 1: −0.92 on the 14) | Deathleaper, Lictor, Von Ryan's Leapers |
| r.movement | Extra movement (on melee lands) | 1 | 1–1.5 | Fitted (at the floor) | 0.4 | 0.5: +0.44; −0.54 | Reiver Squad, Lieutenant, Toxicrene |
| r.detachment | Detachment rules (on the total) | 1.13 | 1–1.2 | Fitted | 0.1 | 1.065: +0.02; +0.14 | Gladiators, Land Raider Redeemer |
| r.lone | Lone Operative (on Soak) | 1.93 | 1–2 | Fitted (5 carriers) | 0.1 | none | Neurolictor |
| r.synapse | Projects Synapse (on the total) | 2.18 | 1.5–2.5 | Fitted (18 carriers) | 1.5 | none (keep) | Ranged Warriors, Norn Emissary, Tervigon, Maleceptor |
| r.debuff | Hinders enemy units (on the total) | 1 | 1–1.25 | Fitted (at the floor) | 2.1 | none (0.75: −0.85; +0.13) | Reiver, Infernus, Suppressor Squads |
| r.buff | Auras on friends (on the total) | 1.24 | 1–1.25 | Fitted (15 carriers) | 2.1 | 1.12: +0.10; +0.13 | Storm Speeders, Techmarine; Razorback, Hammerfall at 2.48 |
| r.heal | Heals or returns models (on the total) | 2.61 | 1–2 | Fitted (3 carriers) | 0.6 | 1.805: +0.11; +0.46 (at 1: +0.60 on the 14) | The Red Terror, Old One Eye, Rhino |
| r.cp | Command points (on the total) | 1.88 | 1–2 | Fitted (15 carriers) | 1.1 | 2.82: +0.58; −0.13 (at 1: −2.26 on the 14) | Captains, Infiltrator Squad, Neurolictor |
| burstKill | Kill as Burst (step 25) | on | on | Kept step | 1.5 (off) | keep | Apothecary, Rhino, Firestrike Servo-Turrets |
| threatSoak | Soak gated by threat (step 26) | on | on | Kept step | 0.15 (off) | keep | Neurogaunts, Rhino, Neurolictor |
| other switches (charsJoined, synapseMarginal, landsNoClamp, killMargin, condShares, leadersLeading, killOverGame, joinedMarginal, joinedShapley, killByJob, arrivalKill, scarceAnswers, carryDedicated, scoreFloor, killClassWeight, secondaryVP) | Reverted steps 6–27 | off | – | Reverted | see reports/tiers/ledger-steps.md | – | – |
| Code constants (not swept: they need a code change) | THREAT_SOAK_K 1 and CAP 2; LANDS_FALLBACK and ENEMY_LANDS 0.6; CLASS_TOUGH_W 2; SCORE_LO/HI 0.25/0.75; ARRIVE_W ⅓; ANSWER_BENCH 3; JOB_STRONG/STEP/TIE; ROUNDS 5; LIST_UNITS 12; the secondary-VP constants (30 + 35 VP, 2,000 ÷ 65 points a VP); data: burstScale (2.21, 2.90), metaKillRate 0.444, the line scales, CUT_SHARE | Reasoned | – | – | – |
4.2 Sensitivity, the top 10 (headline swing within a sensible range)
| # | Knob | Swing | Direction that helps the headline | Does it help the 14 armies? |
|---|---|---|---|---|
| 1 | wCarry | 5.1 | none: any carrying loses | no (flat to −0.4) |
| 2 | scarcity (switch) | 3.4 | keep on | yes (off: −2.0); missing for Space Marines |
| 3 | solo (switch) | 3.4 | keep on | yes (off: −5.6) |
| 4 | r.mortal | 2.3 | toward 1 (+0.79) | yes (+1.13) |
| 5 | r.debuff | 2.1 | stay at 1 | stay at 1 (0.75: +0.13, inside noise) |
| 6 | r.buff | 2.1 | slightly toward 1 (+0.10) | slightly (+0.13) |
| 7 | wPresence / anywhere (one lever: the Biovores) | 1.9 | stay | n/a (Tyranids only) |
| 8 | burstKill (switch) | 1.5 | keep on | yes (off: −1.4) |
| 9 | r.synapse | 1.5 | stay at 2.18 | n/a (Tyranids only) |
| 10 | lands | 1.3 | to 1.25 (+0.90) | flat (+0.08) |
Close behind: r.cp (1.1; raising it helps Tyranids, lowering it costs the 14 armies 2.3 points: a real effect worth pricing properly), r.overwatch (1.0; Marines only), wSoak (0.7), wHold (0.6), r.heal (0.6). Everything else moves the headline by under 0.45 points over its whole range: wScore, wActions, wSpawn, reach, tunnel, threat, soakfx, beta, r.first, r.lone, r.detachment, r.movement, and wKill within ×½ to ×2. Those are not identified by this data and should be frozen.
Combinations and their noise (paired bootstrap over units, 300 resamples; ± one standard error; "P>0" the share of resamples in which the change is positive):
| Setting | Headline Δ | 16 armies Δ | The 14 Δ (P>0) | Typical gap |
|---|---|---|---|---|
| r.mortal 1 | +0.79 ± 0.62 | +1.07 ± 0.69 | +1.13 ± 0.79 (95%) | 17.32 → 16.80 |
| r.mortal 1.345 (halfway) | +0.62 ± 0.44 | +0.59 ± 0.42 | +0.60 ± 0.47 (92%) | 17.04 |
| lands 1.25 | +0.90 ± 0.73 | +0.14 ± 0.49 | +0.08 ± 0.54 (57%) | 17.29 |
| lands 1.3 | +0.85 ± 0.66 | +0.33 ± 0.42 | +0.28 ± 0.47 (73%) | – |
| r.mortal 1 + lands 1.25 | +1.57 ± 0.90 | +1.29 ± 0.79 | +1.32 ± 0.90 (92%) | – |
| r.mortal 1, r.heal 1.805, r.buff 1.12, r.detachment 1.065 | +0.95 ± 0.77 | +1.74 ± 0.84 | +1.87 ± 0.95 (98%) | – |
| the same + lands 1.25 | +1.27 ± 1.03 | +1.61 ± 0.93 | +1.69 ± 1.06 (94%) | – |
| r.mortal 1, every other fitted multiplier halfway to 1 | +0.80 ± 1.40 | +1.50 ± 0.89 | +1.61 ± 1.02 (97%) | – |
| every fitted multiplier halfway to 1 | +0.53 ± 1.33 | +0.79 ± 0.69 | +0.82 ± 0.79 (86%) | – |
| every fitted multiplier at 1 | −2.74 ± 2.75 | −0.53 ± 1.43 | −0.33 ± 1.60 (42%) | – |
| scarcity off (for scale) | −3.39 ± 2.10 | −2.03 ± 1.17 | −1.99 ± 1.30 (7%) | 18.07 |
After a re-fit (tools/ledgerrefit.js, the seven weights and lands re-fitted on the step-6 procedure, run as --on "burstKill+threatSoak;burstKill+threatSoak+r.mortal=1"): r.mortal at 1 gives 71.9% → 71.9% (+0.1), ρ 0.504 → 0.506, left out 0.506 → 0.503; Tyranids +0.003, Marines −0.001: a wash on the headline once the weights absorb it, KEEP by the tool's hint. So the evidence for it is the 14 armies and the carrier count, not the headline.
Reading. All the way to 1 is too far (Synapse, command points and buffs carry real signal), but shrinking toward 1 helps the never-fitted armies every time. That is what a fitted multiplier with too few carriers looks like.
5. Tuning methods, in order of expected value, each with its guard
Common rules for all six. Every run is fixed before it starts: the knobs, the grid, the yardstick and the stopping rule. The judge is the 14 never-fitted armies plus held-out Tyranid units; the headline is a guard (no army falls past noise). Write a benchmark file for each result (tools/ledger-benchmarks/, with a note naming the method), a row in reports/tiers/ledger-steps.md and an entry in design/decisions.md. Never look at a candidate's per-unit letters before the numbers are in.
(A) Shrink the fitted rule multipliers toward 1 (hierarchical shrinkage by hand). Expected value: highest, measured.
- Why: 8 multipliers were fitted on 3 to 18 headline carriers; the 16 armies have 15 to 348 carriers of each kind. With a prior of 1 (the rule does nothing beyond its dice), partial pooling means m ← 1 + s × (m − 1) with one shared shrink factor s.
- Run: a grid s ∈ {1, 0.75, 0.5, 0.25, 0} fixed in advance, mortal wounds first on its own (it is the outlier), then the rest together:
node tools/ledgerstep.js --base tools/ledger-benchmarks/2026-10-04-step26-baseline.json --on r.mortal=1node tools/ledgercarry.js --on burstKill,threatSoak,r.mortal=1 --w 0(the 16 armies and the 14; its base is step 6, so the two kept switches ride along)node tools/stepmoves.js --base tools/ledger-benchmarks/2026-10-04-step26-baseline.json --on r.mortal=1 --aim Zoanthropes,Toxicrenenode tools/ledgerrefit.js --on "burstKill+threatSoak;burstKill+threatSoak+r.mortal=1", andnode tools/diagnose.js --on burstKill+threatSoak+r.mortal=1against--on burstKill+threatSoakfor held-out units. - Judge: pick the s that is best on the 14 armies, provided the headline doesn't fall and neither army's ρ falls past noise. Stop at the first s where the 14 armies stop improving.
- Better, as a card: a true hierarchical fit of each kind's multiplier on all 16 armies (each army's own estimate pulled toward the 16-army one by its carriers), so multipliers come from ~600 units rather than 136.
(B) Sensitivity-ranked selection: freeze what the data can't see. Expected value: high (it prevents losses), free.
- Why: section 4.2 finds 12 knobs that move the headline by under half a point over their whole range. Tuning them is fitting noise.
- Run: nothing new; adopt section 4's table. Record in a benchmark note the active set: wKill (its ratio to the rest), wSoak, wHold, lands, r.mortal, r.synapse, r.cp, r.buff, plus the kept switches. Everything else is frozen at its reasoned value or 1. Repeat the sweep (the review's script, section 8) after every confirmed step, because a model change re-ranks the knobs: lands became stale exactly this way.
- Stop: a knob re-enters the active set only when a new sweep shows a swing over 1 point.
(C) One knob at a time on the active set, with the movement view and the never-fitted armies. Expected value: medium-high.
- Run: for each active knob in order of swing, a grid fixed in advance (e.g. lands {1.0, 1.1, 1.2, 1.25, 1.3, 1.4, 1.5}). Commands as in (A), plus
--unitsfor the aimed units. Save each grid's table. - Judge: choose by the 14 armies and the typical gap; break ties toward the reasoned value. Keep only if the headline gain clears 1.5 paired standard errors (about 1 point) or the 14 armies gain at P>0 of 90% or more, with no army falling past noise. Stop after one pass over the active set; a second pass only after a model change.
(D) Replace knobs with game-derived constants. Expected value: highest in the long run, medium risk.
- Why: every line weight is an exchange rate into points. The game fixes one: about 65 VP decide a game of 2,000 points, so a VP is worth about 31 points (the T-641 bridge). If Score, Hold, Actions and Presence are converted into expected VP by the mission-card model (T-647) and Kill into the VP it denies, the weights disappear and only the conversion remains.
- Run: T-647's per-unit VP as a line priced at 2,000 ÷ 65 points a VP, with fire on each unit in proportion to its points (step 27b's fix); judge with the full battery, the 14 armies first.
- Guard: no constant is chosen after looking at the letters; the survival term is checked for a size lean (ρ of the line with points should be near zero or negative, step 27's diagnosis) before any scoring.
- Stop: if it fails on the 14 armies, it is a measurement to report, not a step.
(E) Expert priors from Jordan. Expected value: medium, for the thin knobs only.
- Why: a multiplier on 3 to 6 carriers (Fights First, healing, Lone Operative, Better Overwatch) can't be fitted; a player's judgement is a better prior than a search.
- Run: ask Jordan, per kind, one question in his terms: "is a unit that fights first worth 1×, 1.25×, 1.5× or 2× its damage?" (the CLAUDE.md format: up to 4 questions, recommended answer first). Record answers as calls (the CALLS list) and never fit them afterwards.
- Judge: the 14 armies, as a check; the call stands unless it loses more than a point there.
(F) A Bayesian or CMA-ES search over a small set, scored by leave-one-list-out and the 14 armies. Expected value: low to medium; last.
- Why last: with 7–9 knobs and gains of 1–2 points against a standard error of about 0.6–1, a global optimiser mostly finds noise; the repo's own first re-search (step 3) found a worse optimum than its start.
- Run: tools/ledgerfit.js already does the right likelihood:
node tools/ledgerfit.js --base tools/ledger-benchmarks/2026-10-04-step26-baseline.json --fit wKill,wSoak,wHold,lands,r.mortal,r.cp,r.synapse,r.buff --prior ones --lambda 0.1 --label review(--prior ones: ridge toward 1, not toward today's values; λ 10 times today's so it can't wander). Score the fit on the 14 armies with ledgercarry.js (--onthe fitted values), and nest it: leave one list out and one army out. - Stop: keep the fit only if it beats (A)+(C) on the 14 armies by more than one standard error.
The order for tonight: (B) first, since it costs nothing and fixes the rules of the game, then (A) mortal wounds on its own, (A) the rest with one shrink factor, (C) lands, then stop and confirm the stack by the old keep rule. (E) can be asked in parallel. (D) is T-647's judgement, and (F) waits until (A)–(C) are done.
6. Ranked recommendations
| # | Step | Kind | Expected effect (measured where shown) | How we'd know |
|---|---|---|---|---|
| 1 | r.mortal 1.69 → 1 | one knob | Headline +0.8 ± 0.6; the 14 +1.1 ± 0.8; typical gap 17.3 → 16.8; Zoanthropes 1 → 9, Toxicrene 31 → 42 (consensus 46), Chaplain with Jump Pack and Brutalis Dreadnought down about 40 places, toward their consensus (letters A/D, A/C). Re-fitted: a wash on the headline | ledgerstep, stepmoves, ledgercarry --w 0 for the 14, ledgerrefit, diagnose --on |
| 2 | The other fitted multipliers shrunk toward 1 by one factor (heal, buff, detachment first; then first, lone, synapse, cp at s = 0.5) | one shrink factor | With 1: headline +1.0 ± 0.8, the 14 +1.9 ± 1.0 (98% of resamples positive) | the same battery; the 14 armies decide s |
| 3 | lands 1.5 → 1.25 | one knob (stale) | Headline +0.9 ± 0.7, every Tyranid list up (ρ +0.015 to +0.034), Space Marine Auspex −0.024; the 14 flat. Tyrannofex 26 → 22, Neurolictor 28 → 24, Broodlord up; Judiciar, Bladeguard down | the same battery; 1.3 is a near-equal alternative with a little more on the 14 |
| 4 | Freeze the 12 knobs with no measurable effect, retire prem and the old threat slider | governance | No number moves; the active set drops from about 20 to 8, inside the budget | a benchmark note listing the active set; the sweep repeated after each confirmed step |
| 5 | Build role scarcity (and Soak effects) for Space Marines | data, one switch's coverage | Guess: Space Marines +1 to +3 points pairwise (scarcity is worth 4.7 on Tyranids, 2.0 on the 14) | tools/scarcity.js and soakfx.js for the army, tools/ledger-merge.js; as its own step, prediction first. Inputs move, so it may need a re-freeze (T-603) |
| 6 | Scope the Biovore dial to the Biovores | code, one call | Harpy about 30 → 43 (consensus 50); Biovores unchanged; guess +0.2 to +0.4 headline | ledgerstep --units Harpy,Biovores |
| 7 | Tournament inclusion as a second target (hedonic fit) | a card | The line weights fitted to within-army inclusion over the 11 armies with 12+ lists, judged on the reviewers; guess +2 to +4 points on the 14 armies, and a straight answer to "are we fitting popularity?" | a new tool beside ledgerfit.js (its likelihood with inclusion as the label); report list share as a yardstick column meanwhile |
| 8 | T-647 judged on the 14 armies first, with per-point fire | the model being built | Unknown; the only route that can lift the support and scoring units (Neurolictor, Neurotyrant, Gargoyles, Tyrant Guard) | its aimed units move toward the consensus and the 14 armies rise; ρ(its line, points) checked first |
Small fixes to the tools, no number changed: let ledgerrefit.js take a confirmed base with switches on; give ledgercarry.js (or a twin) a --base so the 16-army gauge reads the confirmed benchmark directly; add the 14-army mean and the paired standard error to ledgerstep.js's output, so the keep rule's 1-point bar can be read against the noise.
7. What this review did not do
- It changed nothing and judged no step: the supervisor decides.
- The code constants (section 4.1's last row) were not swept; each needs a code edit.
- Held-out Tyranid units (tools/diagnose.js, about 15 minutes per run) were not run for the recommended settings; they are part of each step's battery above.
- The list-share comparison uses three months of one source (grimstat-corpus, CC BY 4.0, lists published by players and organisers on MiniHeadQuarters); Space Marines have only 2 lists, and inclusion also counts tax units and cheap fillers. It is a yardstick, not a verdict on any unit.
8. How to repeat the numbers
- Baseline:
node tools/ledgerstep.js --base tools/ledger-benchmarks/2026-10-04-step26-baseline.json(ρ 0.520, 72.3%). - One knob:
node tools/ledgerstep.js --base tools/ledger-benchmarks/2026-10-04-step26-baseline.json --on <key>=<value>(rule kinds asr.mortal=1); 16 armies:node tools/ledgercarry.js --on burstKill,threatSoak,<key>=<value> --w 0. - The sweep in section 4 is that pair of measurements looped over each knob's grid, with tools/stepmoves.js
movesOffor the movement columns; the paired standard errors resample units within each army (300 draws) and recompute every list's pairwise for base and setting on the same draw. The scripts were scratch files (not in the repo); a tool for it would be tools/knobsweep.js, a small card. - Value shares and the line-only rankings: each unit's ranking row's
contribfromcompute(); list share from docs/data/inclusion.json (lists that include the unit ÷ the army's lists, July to September).