Tactical Reroll
⋯

T-632 tables (tools/diagnose.js output)

From reports/tiers/2026-10-04-diagnosis-tables.md , rendered when the site is built.

Ceiling (each list against the mean of the other four): 90.4%

ModelIn-sample vs consensus (tuned)In-sample, most flexibleIn-sample vs listsHeld out vs consensus (± SE)Held out vs lists (± SE)Held out, pooled across folds (consensus / lists)Scrambled vs consensus (SD)Scrambled vs lists (SD)
(a) the ledger as it is (step 6)75.3%75.3%75.4%74.5% ± 0.675.0% ± 0.672.2% / 72.4%51.1% (4.3)51.4% (4.5)
(a1) the ledger's form, its seven weights re-fitted75.3%75.9%75.4%71.8% ± 1.171.9% ± 1.269.7% / 69.6%50.2% (4.5)50.5% (4.7)
(b) linear on every variable (ridge or lasso)95.0%99.3%91.8%70.1% ± 1.069.2% ± 0.869.1% / 68.4%45.6% (7.7)46.3% (8.0)
(b2) linear on a short core list (25 variables), ridge74.4%81.4%74.0%67.4% ± 0.965.7% ± 0.864.6% / 63.3%49.4% (7.5)49.7% (8.1)
(c) linear plus every pairwise interaction (kernel, ridge)87.9%98.5%86.9%73.1% ± 0.972.0% ± 0.871.6% / 70.8%44.1% (7.0)44.3% (7.1)
(d) boosted trees, monotone where the sign is obvious96.0%99.9%92.2%68.6% ± 0.868.0% ± 0.767.4% / 67.2%51.0% (3.5)51.9% (3.1)
(e) nearest neighbours (similar units)100.0%100.0%93.3%72.5% ± 0.871.6% ± 1.070.2% / 69.3%44.9% (6.9)44.8% (7.1)

Learning curve (trained on m units, scored on the pairs among the rest; ± SE over splits):

Per list, held out:

ModelauspexFullhivemindFullsecondFullmaelstromastrategas
ceiling92.1%92.9%87.7%95.8%83.7%
a80.9%80.1%78.3%70.8%65.1%
a178.3%76.8%75.1%67.2%62.0%
b73.5%72.3%71.9%64.6%63.5%
b272.1%70.7%66.4%58.7%60.7%
c76.9%75.5%73.3%69.4%64.9%
d74.0%72.5%73.2%60.5%60.0%
e75.2%75.1%73.9%69.5%64.2%

Settings chosen in the folds (most often):

Flexible against the ledger, held out (paired; SE over repeats; bootstrap over units SD and the share of resamples at or below 0):

Pairvs consensusvs listsbootstrap SDP(≤ 0)scrambled SD
a1 − a-2.6 ± 0.8-3.1 ± 0.91.71.002.1
b − a-4.4 ± 1.1-5.9 ± 1.04.50.8210.1
b2 − a-7.1 ± 1.2-9.3 ± 1.25.50.899.2
c − a-1.4 ± 0.8-3.0 ± 0.74.50.639.8
d − a-5.8 ± 1.0-7.0 ± 0.93.90.945.6
e − a-2.0 ± 0.8-3.5 ± 1.15.00.649.9
b − a1-1.8 ± 1.2-2.7 ± 1.43.70.679.7
b2 − a1-4.5 ± 1.4-6.2 ± 1.64.80.838.6
c − a11.3 ± 1.00.1 ± 1.23.70.399.5
d − a1-3.2 ± 1.4-3.8 ± 1.63.00.865.1
e − a10.6 ± 1.3-0.3 ± 1.74.30.489.8

Permutation importance, held out, model c (drop in pairwise vs consensus, points):

VariableDropSE
k6_light_vehicle1.640.19
kw_is_fly1.340.22
kw_is_transport1.130.11
kw_heavy1.080.18
kw_is_burrowers0.950.18
kw_precision0.940.17
lands0.910.20
valueCheap0.800.24
inv0.670.19
kw_is_titanic0.610.10
value0.580.29
k6_monster0.550.15
kw_is_battleline0.510.19
bestR_A0.490.19
kw_psychic0.470.15
FamilyDropSE
flags5.790.63
ledger3.370.67
weapons2.740.50
killclass1.300.34
datasheet0.660.47
interaction0.440.20

Permutation importance, held out, model d (drop in pairwise vs consensus, points):

VariableDropSE
c_score2.410.59
value2.010.48
lands0.970.30
valueCheap0.680.26
k6_light_vehicle0.670.24
c_hold0.560.15
k6_character0.270.13
ix_arrivexMelee0.270.23
k6_monster0.260.20
kw_heavy0.230.18
kw_psychic0.150.10
line_killMelee0.140.14
modelsCheap0.130.05
M0.110.13
ocPer1000.100.11
FamilyDropSE
ledger12.501.02
killclass0.720.50
datasheet-0.040.33
flags-0.070.21
interaction-0.500.35
weapons-0.660.30

Permutation importance, held out, model a1 (drop in pairwise vs consensus, points):

VariableDropSE
c_kill19.781.25
c_presence2.160.56
c_score0.910.18
c_hold0.780.22
c_soak0.030.27
c_actions0.010.02
M0.000.00
T0.000.00
Sv0.000.00
inv0.000.00
W0.000.00
Ld0.000.00
OC0.000.00
models0.000.00
pts0.000.00
FamilyDropSE
ledger22.411.40
datasheet0.000.00
weapons0.000.00
flags0.000.00
killclass0.000.00
interaction0.000.00

Partial dependence, model c (fit on all 52; consensus scale; biggest step):

Partial dependence, model d (fit on all 52; consensus scale; biggest step):

The trees' splits (fit on all 52): variable, gain, median threshold, uses:

The trees' interactions (a depth-3 fit on all 52, 100 trees: one split under the other, by the child's gain):

Linear (b) fitted on all 52, largest standardised weights:

The ledger's seven weights re-fitted on all 52 (a1):

Units the best model (c) gets right that the ledger gets wrong (held-out share of the unit's pairs ordered right):

UnitConsensusConsensus rankLedger rankBest modelLedgerGainTop variables (percentile)
Toxicrene1.33346.52388.0%53.4%35k6_light_vehicle 33, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0
Harpy0.850.533100.0%65.6%34k6_light_vehicle 29, kw_is_fly 76, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0
Gargoyles3.6194759.6%41.2%18k6_light_vehicle 24, kw_is_fly 76, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0
Neurotyrant3.667173970.3%52.3%18k6_light_vehicle 12, kw_is_fly 76, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0
Tyranid Warriors with Ranged Bio-Weapons1.8431359.6%44.0%16k6_light_vehicle 18, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0
Norn Emissary4.2111976.5%61.9%15k6_light_vehicle 51, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0
Tyrant Guard2.832.54979.3%65.4%14k6_light_vehicle 37, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0
Tervigon2401653.4%40.0%13k6_light_vehicle 53, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0
Genestealers3.8143466.3%55.3%11k6_light_vehicle 49, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0
Carnifexes2403082.3%71.9%10k6_light_vehicle 96, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0

And the reverse (the ledger right, the model wrong):

Every flexible model misses the same way (mean held-out prediction minus consensus, letters):