Tactical Reroll
⋯

An independent review of the points ledger, and its knobs (T-648, T-649)

From reports/tiers/2026-10-04-review.md , rendered when the site is built.

4 Oct 2026. A review: no model, number, setting, benchmark, data file or frozen input changed. Every number below was run read-only on the confirmed baseline, tools/ledger-benchmarks/2026-10-04-step26-baseline.json (burst Kill + threat-gated Soak: seven headline lists 72.3% pairwise, ρ 0.520; 16 armies 59.3%; the 14 never-fitted 57.8%). Our own numbers and words; unit names only. Guesses are marked as guesses. The scratch scripts that produced the sweeps are described in section 8 so anyone can repeat them with the repo's tools.

The short answer


1. The problem, stated precisely

What is estimated. For each unit of an army, one number: its worth to a 2,000-point list per point it costs, in its best legal form (size, loadout, detachment, alone or led), against one snapshot of the rules (BSData 374f505) and of the meta (a typical enemy army built from top-quarter tournament lists). The tier is that number cut into bands. Only the order within an army is judged.

From what. The engine's measurements per form: Kill (enemy points removed in one activation, since step 25 as Burst: only targets finished), Soak (enemy points that must be spent to remove it, since step 26 scaled by its own threat), OC per point, actions, held objectives (OC × Soak), spawned points, bodies for the unit-count cards, carrying, and rule multipliers per kind of rule. A linear sum of the lines, each divided by its army's median, times the multipliers.

Judged how. Against reviewers' tier letters: of the pairs a reviewer puts in different tiers, the share we order the same way (50% chance); beside it ρ, pairs two or more tiers apart, the movement view (tools/stepmoves.js), held-out units (tools/diagnose.js), a re-fit (tools/ledgerrefit.js) and the 14 armies never tuned on.

The information budget.

UnitsListsCross-tier pairs per listReviewers' own agreement
Tyranids52 (44–52 graded per list)5744–1,04475–96% pairwise between lists; ceiling 90.4%
Space Marines8422,375–2,44091.6% between the two (on pairs both split)
Headline1367about 9,000 pairs, but highly dependent
All 16 armiesabout 600 graded1–3 per army outside the headline

2. Methods from other fields, compared

design/ranking-methods.md (3 Oct) and design/results-methods.md (T-631) surveyed most of these. Here each is set against what has actually been built and measured since, and graded on whether it could plausibly close the 72% → 90% gap with this data.

FieldWhat it doesHow it maps hereWhat we already doWhat we lackLikely to close the gap?
Sports value over replacement (WAR, VORP)Worth = output above what a freely available player gives, per game or per salaryWorth = Value above the cheapest unit that does the same job, per pointLines are divided by the army's median (a crude "average" baseline, not "replacement")A replacement level per job: a Neurolictor is worth what the list loses if the next-best Synapse/action piece fills its slotPartly. It is the honest form of "scarce answers" (step 19 failed because Tyranids aren't short of Kill). Replacement on support and scoring jobs is untried.
Plus-minus with ridge (RAPM, RAPTOR)Regress team results on who was on the floor; shrink each player toward a prior built from his box-score statsLists' placings regressed on the units in them, shrunk toward the ledgerNothing fitted to results; T-631 designed it (the "two layers")The outcome layerNot alone. T-631 measured army explaining 2.4% of placings and only 41 unit rows clearing a loose bar. It is too noisy to grade units; useful for one or two line weights.
Fantasy auction values (points above replacement)Projected stats → points above the replacement at each position → dollars at a fixed budgetExactly our ledger's design: lines → worth in points → per pointThe ledger is this, with points as the budgetPosition-specific replacement (as above), and a check that values sum to the budgetAlready the frame; the missing part is the replacement level.
Paired comparisons (Bradley–Terry, Thurstone, Elo/TrueSkill)Latent strength from who beat whom, with a sharpness per judgeEach reviewer's cross-tier pairs as comparisonsYes: tools/ledgerfit.js is a Bradley–Terry fit with a sharpness per list and a reliability weight per reviewer, ridge to the priorTies used as information (an ordinal "same tier" likelihood); reviewer bias per unit typeNo. It is the right loss and it is in place; a better loss won't add information the features don't carry (T-632: every model lands at about 70% on held-out units).
Learning to rank (RankNet, LambdaMART)Fit a scorer with a pairwise or listwise loss, often with treesThe flexible models of T-632/T-635/T-636Yes, tested: trees, kernels, neighbours, 341 variablesNothingNo, measured: they memorise 52 units and fall to 64–73% held out; across armies 58–63%.
Hedonic pricingA product's price = the sum of implicit prices of its features; prices read from market transactionsMarket = tournament list-building. A unit's inclusion rate (or its share of points spent) regressed on its ledger lines, within each armyNot done. List share has been a yardstick (step 1: held-out ρ 0.41)The fit itself: weights chosen to explain what players buy, across ~11 armies with 12+ tournament listsThe most likely single gain (section 3: list share already orders the Tyranid consensus 86% right). The weights would learn from about 900 unit rows instead of 136.
Conjoint / discrete choicePeople choose between bundles; a conditional logit reads each feature's worthEach tournament list is a bundle chosen under a 2,000-point budget; units compete for points within an armyNothingA choice model with the budget: each unit's chance of being picked against the army's other optionsStrong in principle, the right shape for "worth per point in a list". Heavier to build than hedonic; do hedonic first.
Expert elicitation and aggregation (Dawid–Skene, many-facet Rasch)Estimate each rater's bias and noise together with the item's true scoreReviewer severity (Auspex's 15 S on Space Marines against Tobias's 6), and per-reviewer leaningsReliability weights per list; pairwise ignores severity by designPer-reviewer cut points (ordinal model); a "consensus" built from the model rather than the mean letterSmall. The reviewers agree at about 90%, so modelling them better buys a point at most. It mainly sharpens the yardstick.
Bayesian hierarchical models with partial poolingEach group's parameters drawn from a shared distribution; thin groups shrink to the meanA rule multiplier fitted on 3 Tyranid carriers borrows strength from 75 carriers across 16 armies; army-specific weights shrink to shared onesNot done; weights are shared and fixed, rule multipliers are Tyranid/Marine fitsShrinkage of multipliers toward 1 (or toward a 16-army estimate)Yes, cheaply. Measured in this review: moving fitted multipliers toward 1 gains on both the headline and the 14 armies (section 4).
Structural / game-simulation modelsValue from simulating the game (e.g. win probability added in chess or baseball)The mission-card model (T-647): each primary and secondary scored per unit; the charge model; Kill over the gameSeveral parts built; steps 12, 18 and 27 tried game-length terms and lostA unit's worth as the change in a typical list's expected VP when it is swapped inThe only route to the WHY, highest ceiling, highest risk. Steps 12 and 27 show the trap: a survival term with fixed enemy fire rewards size. Step 27b's per-point fire is the right ruler.

Which would most likely close the gap, given this data.

  1. Hedonic pricing on tournament inclusion, judged on the reviewers. Expected: the line weights move toward what players pay for (guess: Score/Hold and support up, raw Kill down), with a gain on the 14 armies. It is the one source of more labelled units, and the evidence below says it is the same signal the reviewers carry.
  2. Partial pooling (shrinkage) of the rule multipliers. Measured: +1 to +2 points on never-fitted armies.
  3. The structural model (T-647), for the units no datasheet number explains, on condition that it is judged on the 14 armies first and its survival term is per point.
  4. Better loss functions and flexible learners: already exhausted (T-632, T-635, T-636).

3. What we do well, and what is wrong or missing

Done well

Wrong or missing

  1. The target is closer to "what players take" than "worth from the rules". Read-only numbers from this review (docs/data/inclusion.json, July to September):

| | Tournament lists | Ledger vs reviewers | List share vs reviewers | |---|---|---|---| | Tyranids | 23 | 74.8% pairwise against the consensus; ρ 0.66 | 85.8%; ρ 0.88 | | Space Marines | 2 | 64.1% | 57.5% (2 lists: no signal) | | 16 armies, mean over each army's lists | 1–23 | 59.4% | 70.9% (14 of 16 better) | | The 11 armies with 12+ tournament lists | | 62.0% | 71.3% |

On Tyranids, list share and the ledger agree with each other at ρ 0.61; an equal blend of the two (84.7%) is no better than list share alone. Either the reviewers report the meta, or the meta follows the reviewers. Either way fitting the reviewers is very nearly fitting popularity, and a rules-only model judged against them has a ceiling well under the reviewers' 90% (guess: high 70s on Tyranids). That doesn't make the ledger wrong; it makes the yardstick ambiguous. The repo's own plan (design/ranking-methods.md, move 4: popularity as a nuisance term) was never run.

  1. Value is Kill; the other lines are decoration with fitted weights. Median shares of Value: Kill 91%, Soak 2%, Hold 2–4%, Score 1–2%. Halving or doubling the Kill weight changes the headline by 0.2–0.4 points: the weight sits on a flat ridge, so it is not identified, and neither are the others relative to it. Kill alone transfers better to the 14 armies (58.4%) than the full ledger (57.8%).
  2. Too many fitted numbers for the budget (about 20 against 7–9). The rule multipliers were set by the 3 Oct joint search on a box from 1 to 3, on the headline's best rows, with very few carriers: Fights First 3, healing 3, Lone Operative 5, Better Overwatch 6, mortal wounds 12, buffs 15, command points 15, Synapse 18. A multiplier fitted on 3 to 6 units is noise by the repo's own rule ("an ability fitted on 3 carriers is noise"). The measured pattern confirms it: pulling them toward 1 helps the never-fitted armies.
  3. Some knobs are stale: they were fitted to an earlier model. The melee lands scale sits at 1.5, the search's upper limit, fitted before Kill became Burst (step 25). Burst already counts only finished targets, so the old lift on melee is now counted on top. Measured: 1.25 is better on every Tyranid list.
  4. The re-fit can't really move anything. ledgerfit's ridge pulls each weight toward the base benchmark with λ 0.01 per relative size, and the rule multipliers are hibernated. A re-fit therefore re-confirms the old weights (wSoak 14.97 → 15.23), and "kept after a re-fit" is a weaker test than it sounds. Worse, the re-fitted base scores below the as-is weights (71.9% against 72.3%), so the as-is weights carry some luck. ledgerrefit.js also refuses the confirmed base (it requires every step switch off), so a step on top of step 26 must be run as --on "burstKill+threatSoak;burstKill+threatSoak+<step>".
  5. The two headline armies don't run the same model. Space Marines' file carries no role scarcity and no Soak effects (their meta has no medianFactor), so those switches do nothing there; Tyranids and the other 14 armies have them. Role scarcity is worth 4.7 points on Tyranids and 2.0 on the 14 armies. The weakest headline army is running without one of the strongest parts. (design/ledger.md section 16 notes "Space Marines' file carries no scarcity"; it has not been fixed.)
  6. The Biovore dial has a side effect. Placed-anywhere 523 also lifts the Harpy (its spawned Spore Mines are placed anywhere too): the Harpy is 30th of 52 against a consensus of 50th, one of our worst "too high" units; with the dial at 1 it drops 13 places. A hand fix for one unit should be scoped to that unit.
  7. The yardstick over-weights one army. Five of the seven lists are Tyranid and they agree with each other, so the mean is roughly 70% Tyranid. The per-army guard in the keep rule handles falls, but gains are still judged mostly on Tyranids.
  8. The provisional keep rule lets noise stack. Its tests (a) and (b) are close to coin flips for a change of pure noise, and (c) only stops falls past noise (Tyranids 0.039, Space Marines 0.098 in ρ: large). The confirmation step protects against this, but only if confirmation is judged on data the steps didn't see: the 14 armies, and held-out units.
  9. "Per point" is right, and the joined-unit handling is honest, but "best form across detachments" is generous to units that are good in one detachment nobody plays; that is stated in design/ledger.md and accepted. Not a priority.
  10. The "same units missed by every model" signal is correct and actionable: what those units have in common is support, scoring and list role (Neurolictor, Neurotyrant, Gargoyles, Tyrant Guard, Termagants), or being pure Kill per point that the reviewers discount (Zoanthropes, Von Ryan's Leapers, Winged Hive Tyrant, Deathleaper). More datasheet variables cannot fix that (T-636); a model of the list and the mission can.

4. The knobs (T-649)

How measured. One knob at a time from the step-26 base, every other setting held, through docs/lab/ledger/ledger.js compute() exactly as tools/ledgerstep.js applies --on key=value. For each value: the change in mean headline pairwise (7 lists), each army's change, the movement view (tools/stepmoves.js movesOf: closer/farther, typical gap, in band, mean places moved), and the change in mean pairwise over the 14 never-fitted armies (as tools/ledgercarry.js scores them). Swing = the largest change in headline pairwise over the knob's sweep. Paired standard errors (bootstrap over units) for the promising settings are in the second table. Base: headline 72.29%, ρ 0.520, Tyranids 75.6%, Space Marines 64.0%, 16 armies 59.34%, the 14 57.85%, typical gap 17.32 places, 54 of 136 units in the reviewers' tier.

4.1 Inventory, with sensitivity

"Fitted" means set by a search or fit against the headline lists (or their predecessors). "Reasoned" means set before looking. "Call" means one of Jordan's calls. Sensible ranges are the reviewer's judgement.

KnobIn the gameNowSensible rangeSet bySwing (headline)Best measured (headline; the 14)Units it moves most
wKillExchange rate of a point killed in one activation498250–1000 (the scale is the other lines')Fitted (joint search, rescaled)9.5 at 0; 0.2–0.4 within ×½–×2none: flat ridgeat 250: Tactical Squad, Neurogaunts, Rhino up; at 2000 Biovores −27
wSoakA point of enemy fire absorbed (threat-gated)15.00–120Fitted0.7120: +0.71; −0.11Land Raiders, Terminator Assault Squad, Repulsor up; Tyrannofex
wScoreOC per point9.650–40Fitted0.4noneNeurogaunts, Termagants, Tactical Squad
wActionsActions per point1.310–20Fitted (at the floor)0.1nonealmost nothing
wHoldOC × Soak: holding an objective18.90–75Fitted0.675.6: +0.21; −0.34Tactical Squad, Rhino, Neurogaunts up; Outriders, Intercessors down at 0
wPresenceBodies for unit-count cards2.520–10Fitted1.9none (0 costs −1.9: Biovores)Biovores, Harpy
wSpawnPoints placed a round15.70–60Fitted0.20: +0.16; 0Tervigon, Biovores
wCarryModels carried forward00–1Call (hibernated)5.1none (every value loses)Rhino, Drop Pod, Tyrannocyte, Land Raider Crusader
premPremium on Kill1–Fixedsame as wKill–(duplicates wKill: retire)
landsScale on the charge model's chance melee lands1.51–1.5Fitted (at the search's upper bound, before Burst)1.31.25: +0.90; +0.08 (Tyranids +1.3)Judiciar, Bladeguard, Terminators down; Tyrannofex, Neurolictor, Broodlord up
anywhereSpawned units placed anywhere (Presence)5231 or a unit-scoped dialCall (Biovore dial)1.9none (300: −0.06)Biovores, Harpy (side effect)
reachOwn bodies counted × when it arrives far forward11–2Fitted (at the floor)0.1–Neurolictor, Raveners
tunnelBodies per tunnel marker00–1Call0.0–none (no form carries tunnels at the ranking rows)
threatOld threat weighting on Soak0–Superseded by threatSoak0.0–none while threatSoak is on: retire
soloRank units alone (not joined)ononReasoned3.4 (off: −3.4; the 14: −5.6)keep onevery Space Marine character
scarcityRole scarcity on KillononReasoned (measurement)3.4 (off: Tyranids −4.7; the 14: −2.0)keep on; missing for Space MarinesExocrine, Tyrannofex, Trygon (down when off)
soakfxExposure and own healing on SoakononReasoned0.0–The Red Terror
betaHow sharply fire picks its best targets63–∞Reasoned0.412: +0.15; −0.04Trygon, Tyrannofex, Mawloc
combo, cutS–cutC, rawDisplay (combos shown, band cuts, raw lines)–––0 on pairwise–bands only
fireBEnemy fire on each unit a turn167per point (step 27b)Reasonedinactive at base–only with killOverGame or secondaryVP
r.mortalMortal wounds outside attacks1.691–1.5Fitted (12 headline carriers)2.31: +0.79; +1.13Zoanthropes, Toxicrene, Mawloc, Chaplain with Jump Pack, Brutalis down
r.overwatchBetter Overwatch10.75–1.5Call1.00.75: +0.18 (Marines +1.4); −0.01Firestrike Servo-Turrets, Invictor, Hammerfall Bunker
r.firstFights First (on melee Kill)2.641–1.5Fitted (3 carriers)0.31.82: +0.08; −0.11 (at 1: −0.92 on the 14)Deathleaper, Lictor, Von Ryan's Leapers
r.movementExtra movement (on melee lands)11–1.5Fitted (at the floor)0.40.5: +0.44; −0.54Reiver Squad, Lieutenant, Toxicrene
r.detachmentDetachment rules (on the total)1.131–1.2Fitted0.11.065: +0.02; +0.14Gladiators, Land Raider Redeemer
r.loneLone Operative (on Soak)1.931–2Fitted (5 carriers)0.1noneNeurolictor
r.synapseProjects Synapse (on the total)2.181.5–2.5Fitted (18 carriers)1.5none (keep)Ranged Warriors, Norn Emissary, Tervigon, Maleceptor
r.debuffHinders enemy units (on the total)11–1.25Fitted (at the floor)2.1none (0.75: −0.85; +0.13)Reiver, Infernus, Suppressor Squads
r.buffAuras on friends (on the total)1.241–1.25Fitted (15 carriers)2.11.12: +0.10; +0.13Storm Speeders, Techmarine; Razorback, Hammerfall at 2.48
r.healHeals or returns models (on the total)2.611–2Fitted (3 carriers)0.61.805: +0.11; +0.46 (at 1: +0.60 on the 14)The Red Terror, Old One Eye, Rhino
r.cpCommand points (on the total)1.881–2Fitted (15 carriers)1.12.82: +0.58; −0.13 (at 1: −2.26 on the 14)Captains, Infiltrator Squad, Neurolictor
burstKillKill as Burst (step 25)ononKept step1.5 (off)keepApothecary, Rhino, Firestrike Servo-Turrets
threatSoakSoak gated by threat (step 26)ononKept step0.15 (off)keepNeurogaunts, Rhino, Neurolictor
other switches (charsJoined, synapseMarginal, landsNoClamp, killMargin, condShares, leadersLeading, killOverGame, joinedMarginal, joinedShapley, killByJob, arrivalKill, scarceAnswers, carryDedicated, scoreFloor, killClassWeight, secondaryVP)Reverted steps 6–27off–Revertedsee reports/tiers/ledger-steps.md––
Code constants (not swept: they need a code change)THREAT_SOAK_K 1 and CAP 2; LANDS_FALLBACK and ENEMY_LANDS 0.6; CLASS_TOUGH_W 2; SCORE_LO/HI 0.25/0.75; ARRIVE_W ⅓; ANSWER_BENCH 3; JOB_STRONG/STEP/TIE; ROUNDS 5; LIST_UNITS 12; the secondary-VP constants (30 + 35 VP, 2,000 ÷ 65 points a VP); data: burstScale (2.21, 2.90), metaKillRate 0.444, the line scales, CUT_SHAREReasoned–––

4.2 Sensitivity, the top 10 (headline swing within a sensible range)

#KnobSwingDirection that helps the headlineDoes it help the 14 armies?
1wCarry5.1none: any carrying losesno (flat to −0.4)
2scarcity (switch)3.4keep onyes (off: −2.0); missing for Space Marines
3solo (switch)3.4keep onyes (off: −5.6)
4r.mortal2.3toward 1 (+0.79)yes (+1.13)
5r.debuff2.1stay at 1stay at 1 (0.75: +0.13, inside noise)
6r.buff2.1slightly toward 1 (+0.10)slightly (+0.13)
7wPresence / anywhere (one lever: the Biovores)1.9stayn/a (Tyranids only)
8burstKill (switch)1.5keep onyes (off: −1.4)
9r.synapse1.5stay at 2.18n/a (Tyranids only)
10lands1.3to 1.25 (+0.90)flat (+0.08)

Close behind: r.cp (1.1; raising it helps Tyranids, lowering it costs the 14 armies 2.3 points: a real effect worth pricing properly), r.overwatch (1.0; Marines only), wSoak (0.7), wHold (0.6), r.heal (0.6). Everything else moves the headline by under 0.45 points over its whole range: wScore, wActions, wSpawn, reach, tunnel, threat, soakfx, beta, r.first, r.lone, r.detachment, r.movement, and wKill within ×½ to ×2. Those are not identified by this data and should be frozen.

Combinations and their noise (paired bootstrap over units, 300 resamples; ± one standard error; "P>0" the share of resamples in which the change is positive):

SettingHeadline Δ16 armies ΔThe 14 Δ (P>0)Typical gap
r.mortal 1+0.79 ± 0.62+1.07 ± 0.69+1.13 ± 0.79 (95%)17.32 → 16.80
r.mortal 1.345 (halfway)+0.62 ± 0.44+0.59 ± 0.42+0.60 ± 0.47 (92%)17.04
lands 1.25+0.90 ± 0.73+0.14 ± 0.49+0.08 ± 0.54 (57%)17.29
lands 1.3+0.85 ± 0.66+0.33 ± 0.42+0.28 ± 0.47 (73%)–
r.mortal 1 + lands 1.25+1.57 ± 0.90+1.29 ± 0.79+1.32 ± 0.90 (92%)–
r.mortal 1, r.heal 1.805, r.buff 1.12, r.detachment 1.065+0.95 ± 0.77+1.74 ± 0.84+1.87 ± 0.95 (98%)–
the same + lands 1.25+1.27 ± 1.03+1.61 ± 0.93+1.69 ± 1.06 (94%)–
r.mortal 1, every other fitted multiplier halfway to 1+0.80 ± 1.40+1.50 ± 0.89+1.61 ± 1.02 (97%)–
every fitted multiplier halfway to 1+0.53 ± 1.33+0.79 ± 0.69+0.82 ± 0.79 (86%)–
every fitted multiplier at 1−2.74 ± 2.75−0.53 ± 1.43−0.33 ± 1.60 (42%)–
scarcity off (for scale)−3.39 ± 2.10−2.03 ± 1.17−1.99 ± 1.30 (7%)18.07

After a re-fit (tools/ledgerrefit.js, the seven weights and lands re-fitted on the step-6 procedure, run as --on "burstKill+threatSoak;burstKill+threatSoak+r.mortal=1"): r.mortal at 1 gives 71.9% → 71.9% (+0.1), ρ 0.504 → 0.506, left out 0.506 → 0.503; Tyranids +0.003, Marines −0.001: a wash on the headline once the weights absorb it, KEEP by the tool's hint. So the evidence for it is the 14 armies and the carrier count, not the headline.

Reading. All the way to 1 is too far (Synapse, command points and buffs carry real signal), but shrinking toward 1 helps the never-fitted armies every time. That is what a fitted multiplier with too few carriers looks like.

5. Tuning methods, in order of expected value, each with its guard

Common rules for all six. Every run is fixed before it starts: the knobs, the grid, the yardstick and the stopping rule. The judge is the 14 never-fitted armies plus held-out Tyranid units; the headline is a guard (no army falls past noise). Write a benchmark file for each result (tools/ledger-benchmarks/, with a note naming the method), a row in reports/tiers/ledger-steps.md and an entry in design/decisions.md. Never look at a candidate's per-unit letters before the numbers are in.

(A) Shrink the fitted rule multipliers toward 1 (hierarchical shrinkage by hand). Expected value: highest, measured.

(B) Sensitivity-ranked selection: freeze what the data can't see. Expected value: high (it prevents losses), free.

(C) One knob at a time on the active set, with the movement view and the never-fitted armies. Expected value: medium-high.

(D) Replace knobs with game-derived constants. Expected value: highest in the long run, medium risk.

(E) Expert priors from Jordan. Expected value: medium, for the thin knobs only.

(F) A Bayesian or CMA-ES search over a small set, scored by leave-one-list-out and the 14 armies. Expected value: low to medium; last.

The order for tonight: (B) first, since it costs nothing and fixes the rules of the game, then (A) mortal wounds on its own, (A) the rest with one shrink factor, (C) lands, then stop and confirm the stack by the old keep rule. (E) can be asked in parallel. (D) is T-647's judgement, and (F) waits until (A)–(C) are done.

6. Ranked recommendations

#StepKindExpected effect (measured where shown)How we'd know
1r.mortal 1.69 → 1one knobHeadline +0.8 ± 0.6; the 14 +1.1 ± 0.8; typical gap 17.3 → 16.8; Zoanthropes 1 → 9, Toxicrene 31 → 42 (consensus 46), Chaplain with Jump Pack and Brutalis Dreadnought down about 40 places, toward their consensus (letters A/D, A/C). Re-fitted: a wash on the headlineledgerstep, stepmoves, ledgercarry --w 0 for the 14, ledgerrefit, diagnose --on
2The other fitted multipliers shrunk toward 1 by one factor (heal, buff, detachment first; then first, lone, synapse, cp at s = 0.5)one shrink factorWith 1: headline +1.0 ± 0.8, the 14 +1.9 ± 1.0 (98% of resamples positive)the same battery; the 14 armies decide s
3lands 1.5 → 1.25one knob (stale)Headline +0.9 ± 0.7, every Tyranid list up (ρ +0.015 to +0.034), Space Marine Auspex −0.024; the 14 flat. Tyrannofex 26 → 22, Neurolictor 28 → 24, Broodlord up; Judiciar, Bladeguard downthe same battery; 1.3 is a near-equal alternative with a little more on the 14
4Freeze the 12 knobs with no measurable effect, retire prem and the old threat slidergovernanceNo number moves; the active set drops from about 20 to 8, inside the budgeta benchmark note listing the active set; the sweep repeated after each confirmed step
5Build role scarcity (and Soak effects) for Space Marinesdata, one switch's coverageGuess: Space Marines +1 to +3 points pairwise (scarcity is worth 4.7 on Tyranids, 2.0 on the 14)tools/scarcity.js and soakfx.js for the army, tools/ledger-merge.js; as its own step, prediction first. Inputs move, so it may need a re-freeze (T-603)
6Scope the Biovore dial to the Biovorescode, one callHarpy about 30 → 43 (consensus 50); Biovores unchanged; guess +0.2 to +0.4 headlineledgerstep --units Harpy,Biovores
7Tournament inclusion as a second target (hedonic fit)a cardThe line weights fitted to within-army inclusion over the 11 armies with 12+ lists, judged on the reviewers; guess +2 to +4 points on the 14 armies, and a straight answer to "are we fitting popularity?"a new tool beside ledgerfit.js (its likelihood with inclusion as the label); report list share as a yardstick column meanwhile
8T-647 judged on the 14 armies first, with per-point firethe model being builtUnknown; the only route that can lift the support and scoring units (Neurolictor, Neurotyrant, Gargoyles, Tyrant Guard)its aimed units move toward the consensus and the 14 armies rise; ρ(its line, points) checked first

Small fixes to the tools, no number changed: let ledgerrefit.js take a confirmed base with switches on; give ledgercarry.js (or a twin) a --base so the 16-army gauge reads the confirmed benchmark directly; add the 14-army mean and the paired standard error to ledgerstep.js's output, so the keep rule's 1-point bar can be read against the noise.

7. What this review did not do

8. How to repeat the numbers