Dry run: not a step, nothing adopted. No setting was written.
The printout of node tools/ledgerfit.js (no options: the step-6 baseline, every switch as it, the default knobs fitted), 3 Oct. A reading by the builder follows the printout.
- Base: tools/ledger-benchmarks/2026-10-03-step6-baseline.json; lists: tyr.auspex, tyr.hivemind, tyr.second, tyr.maelstrom, sm.auspex (Tista Minis excluded).
- Fitted: wKill, wSoak, wScore, wActions, wHold, wPresence, wSpawn, lands; prior: base; ridge λ 0.01 (the cost, in NLL per pair, of moving a knob by its own size).
- Floors: wKill ≥ 0, wSoak ≥ 0, wScore ≥ 0.5, wActions ≥ 0.5, wHold ≥ 1, wPresence ≥ 0.5, wSpawn ≥ 0, lands ≥ 0.
- Switches and sliders set for the run: none (as the base).
- Hibernated (fixed at the base): prem 1, anywhere 523, threat 0, solo 1, scarcity 1, soakfx 1, tunnel 0, raw 0, reach 1, beta 6, combo 10, r.mortal 1.69, r.overwatch 1, r.first 2.64, r.movement 1, r.detachment 1.13, r.lone 1.93, r.synapse 2.18, r.debuff 1, r.buff 1.24, r.heal 2.61, r.cp 1.88.
- Run time: 24.5 s (the fit itself 1.07 s, 1667 iterations; 200 bootstrap fits, 5 folds, 10 scrambles on 11 threads). Basis against valueForm(): max relative difference 5.5e-16; ρ and pairwise checked against compute() and shape().
The lists
| List | Reliability (mean ρ with the others) | Pair weight | Scale s | s × SD(Value) | Cross-tier pairs |
|---|---|---|---|---|---|
| tyr.auspex | 0.746 (hivemind 0.734, second 0.667, maelstrom 0.836) | 0.746 | 1.683e-3 | 1.444 | 742 |
| tyr.hivemind | 0.699 (auspex 0.734, second 0.542, maelstrom 0.822) | 0.699 | 1.459e-3 | 1.254 | 743 |
| tyr.second | 0.648 (auspex 0.667, hivemind 0.542, maelstrom 0.735) | 0.648 | 1.151e-3 | 1.042 | 630 |
| tyr.maelstrom | 0.798 (auspex 0.836, hivemind 0.822, second 0.735) | 0.798 | 1.089e-3 | 0.904 | 920 |
| sm.auspex | 0.746 (proxy: tyr.auspex (the same reviewer on another army)) | 0.746 | 3.033e-3 | 0.753 | 778 |
The fitted knobs (standard errors: bootstrap over units, 200 resamples)
| Knob | Base | Prior | Fit | SE | 95% interval | |
|---|---|---|---|---|---|---|
| wKill | 498.2 | 498.2 | 449.3 | 176.7 | 48.41 to 677.1 | |
| wSoak | 14.97 | 14.97 | 15.24 | 2.234 | 11.88 to 20.28 | |
| wScore | 9.652 | 9.652 | 9.829 | 1.616 | 6.433 to 12.95 | |
| wActions | 1.306 | 1.306 | 1.295 | 0.023 | 1.256 to 1.327 | |
| wHold | 18.89 | 18.89 | 17.34 | 3.868 | 8.651 to 23.29 | |
| wPresence | 2.515 | 2.515 | 2.900 | 0.764 | 0.500 to 3.399 | |
| wSpawn | 15.65 | 15.65 | 14.97 | 1.261 | 11.13 to 16.06 | |
| lands | 1.500 | 1.500 | 1.451 | 0.342 | 0.638 to 2.355 |
The overall size of the weights is set by the ridge alone (the scales absorb it): read the weights against each other.
The yardstick
ρ / pairwise accuracy / pairs 2+ tiers apart, per list:
| List | Base (step 6) | Every weight 1 | Fit |
|---|---|---|---|
| tyr.auspex | 0.662 / 80% / 87% | 0.475 / 72% / 76% | 0.656 / 79% / 86% |
| tyr.hivemind | 0.622 / 79% / 91% | 0.370 / 67% / 77% | 0.620 / 79% / 90% |
| tyr.second | 0.587 / 77% / 88% | 0.304 / 62% / 75% | 0.567 / 76% / 87% |
| tyr.maelstrom | 0.571 / 75% / 84% | 0.343 / 64% / 70% | 0.558 / 74% / 83% |
| sm.auspex | 0.406 / 68% / 80% | 0.278 / 62% / 71% | 0.390 / 67% / 79% |
| mean | 0.569 / 75% / 86% | 0.354 / 65% / 74% | 0.558 / 75% / 85% |
| NLL per pair (scales fitted) | 0.528 | 0.663 | 0.525 |
Leave one list out (fit on the other lists, score the one left out)
| Left out | Fold fit (held out) | Full fit (in sample) | Base | Every weight 1 |
|---|---|---|---|---|
| tyr.auspex | 0.645 / 79% / 86% | 0.656 / 79% / 86% | 0.662 / 80% / 87% | 0.475 / 72% / 76% |
| tyr.hivemind | 0.620 / 79% / 90% | 0.620 / 79% / 90% | 0.622 / 79% / 91% | 0.370 / 67% / 77% |
| tyr.second | 0.565 / 76% / 87% | 0.567 / 76% / 87% | 0.587 / 77% / 88% | 0.304 / 62% / 75% |
| tyr.maelstrom | 0.562 / 74% / 84% | 0.558 / 74% / 83% | 0.571 / 75% / 84% | 0.343 / 64% / 70% |
| sm.auspex | 0.380 / 67% / 79% | 0.390 / 67% / 79% | 0.406 / 68% / 80% | 0.278 / 62% / 71% |
| mean | 0.555 / 75% / 85% | 0.558 | 0.569 | 0.354 |
How far the knobs move with each list dropped:
| Left out | wKill | wSoak | wScore | wActions | wHold | wPresence | wSpawn | lands |
|---|---|---|---|---|---|---|---|---|
| tyr.auspex | 458.1 | 14.53 | 9.664 | 1.294 | 15.76 | 3.103 | 15.00 | 1.451 |
| tyr.hivemind | 450.7 | 14.93 | 10.10 | 1.291 | 18.44 | 2.768 | 15.07 | 1.451 |
| tyr.second | 457.7 | 15.27 | 10.14 | 1.297 | 17.81 | 2.729 | 15.14 | 1.289 |
| tyr.maelstrom | 475.1 | 15.27 | 9.402 | 1.296 | 16.07 | 2.988 | 15.00 | 1.451 |
| sm.auspex | 418.4 | 16.53 | 9.761 | 1.296 | 18.64 | 2.703 | 14.78 | 1.235 |
Scrambled letters (10 fits, each list's letters shuffled, scored against them)
- Mean ρ: -0.013 (SD 0.059, -0.116 to 0.070); mean pairwise 49% (SD 2%).
- The real fit's mean ρ 0.558 is 9.637 SDs above the scrambled mean; the fit's gain over the base, -0.011, against the scrambled SD 0.059.
Reading (the builder's, not part of the printout)
- The likelihood barely moves off the baseline. At λ 0.01 the fit lands within one standard error of the step-6 weights on every knob (Kill 449 against 498, Presence 2.90 against 2.52, lands 1.45 against 1.5), and its NLL per pair is 0.525 against the baseline's 0.528. The baseline was searched on ρ, so it already sits near the likelihood's optimum; the fit's mean ρ (0.558) is 0.011 under the baseline's (0.569), well inside the scrambled spread (SD 0.059): the same ledger as far as this yardstick can tell. Every weight at 1 is far behind (0.354, NLL 0.663): the baseline's weights carry real information beyond the unit-weight ledger.
- The surface is flat; the ridge decides. A side run at λ 0.001 (not in the printout) reaches NLL 0.523 with Hold at its floor, Spawn 6.3 and Kill 327, mean ρ 0.555: a different setting, no better on either yardstick. The bootstrap's wide Kill interval (48 to 677) says the same: the lists pin the weights' order more than their sizes. The tight Actions interval (±0.02) is the ridge, not the data.
- Leave one list out holds: each held-out list scores within 0.011 of its in-sample ρ (mean 0.555 against 0.558), and the knobs move little with any list dropped (Kill 418 to 475, Hold 15.8 to 18.6, lands 1.24 to 1.45).
- Scrambled letters: the fitter on shuffled tiers scores ρ −0.013 (SD 0.059, −0.116 to 0.070) against the shuffled letters. With the ridge holding it near the prior it can't chase noise far, so this check bounds the fitter at this λ only.
- The floors are in the setting's own units. With
--prior ones(side runs) the floors bind (Hold 1, Presence 0.5; at λ 0.001 Score, Actions and Spawn too) because the unit-weight prior puts Kill near 1 where the baseline has 498; a fit toward unit weights would need the floors restated on the normalised scale first. - The reliability weights: Maelstrom 0.80, Auspex 0.75, Hivemind 0.70, the second creator 0.65 (mean ρ with the other Tyranid lists); the Space Marines' Auspex list has no second Marine list in the fit (Tista Minis excluded) and borrows Auspex's Tyranid figure. The scales: Auspex the sharpest of the Tyranid lists (s × SD 1.44), Maelstrom the softest (0.90), Marine Auspex 0.75.