Ceiling (each list against the mean of the other four): 90.4%
| Model | In-sample vs consensus (tuned) | In-sample, most flexible | In-sample vs lists | Held out vs consensus (± SE) | Held out vs lists (± SE) | Held out, pooled across folds (consensus / lists) | Scrambled vs consensus (SD) | Scrambled vs lists (SD) |
|---|---|---|---|---|---|---|---|---|
| (a) the ledger as it is (step 6) | 75.3% | 75.3% | 75.4% | 74.5% ± 0.6 | 75.0% ± 0.6 | 72.2% / 72.4% | 51.1% (4.3) | 51.4% (4.5) |
| (a1) the ledger's form, its seven weights re-fitted | 75.3% | 75.9% | 75.4% | 71.8% ± 1.1 | 71.9% ± 1.2 | 69.7% / 69.6% | 50.2% (4.5) | 50.5% (4.7) |
| (b) linear on every variable (ridge or lasso) | 95.0% | 99.3% | 91.8% | 70.1% ± 1.0 | 69.2% ± 0.8 | 69.1% / 68.4% | 45.6% (7.7) | 46.3% (8.0) |
| (b2) linear on a short core list (25 variables), ridge | 74.4% | 81.4% | 74.0% | 67.4% ± 0.9 | 65.7% ± 0.8 | 64.6% / 63.3% | 49.4% (7.5) | 49.7% (8.1) |
| (c) linear plus every pairwise interaction (kernel, ridge) | 87.9% | 98.5% | 86.9% | 73.1% ± 0.9 | 72.0% ± 0.8 | 71.6% / 70.8% | 44.1% (7.0) | 44.3% (7.1) |
| (d) boosted trees, monotone where the sign is obvious | 96.0% | 99.9% | 92.2% | 68.6% ± 0.8 | 68.0% ± 0.7 | 67.4% / 67.2% | 51.0% (3.5) | 51.9% (3.1) |
| (e) nearest neighbours (similar units) | 100.0% | 100.0% | 93.3% | 72.5% ± 0.8 | 71.6% ± 1.0 | 70.2% / 69.3% | 44.9% (6.9) | 44.8% (7.1) |
Learning curve (trained on m units, scored on the pairs among the rest; ± SE over splits):
- c: 16: 67.0% ± 1.4, 21: 69.0% ± 1.8, 26: 71.2% ± 1.1, 31: 70.2% ± 1.5, 36: 73.2% ± 1.6, 41: 70.2% ± 3.0
- a1: 16: 73.3% ± 1.1, 21: 72.0% ± 1.6, 26: 71.9% ± 1.0, 31: 72.0% ± 1.5, 36: 72.3% ± 2.3, 41: 73.0% ± 2.7
- a: 16: 75.7% ± 0.9, 21: 75.6% ± 0.8, 26: 73.3% ± 1.1, 31: 76.3% ± 1.9, 36: 75.6% ± 1.6, 41: 73.7% ± 2.7
- b2: 16: 58.8% ± 2.3, 21: 62.2% ± 1.8, 26: 63.7% ± 1.5, 31: 62.5% ± 2.6, 36: 66.2% ± 2.4, 41: 66.7% ± 2.9
Per list, held out:
| Model | auspexFull | hivemindFull | secondFull | maelstrom | astrategas |
|---|---|---|---|---|---|
| ceiling | 92.1% | 92.9% | 87.7% | 95.8% | 83.7% |
| a | 80.9% | 80.1% | 78.3% | 70.8% | 65.1% |
| a1 | 78.3% | 76.8% | 75.1% | 67.2% | 62.0% |
| b | 73.5% | 72.3% | 71.9% | 64.6% | 63.5% |
| b2 | 72.1% | 70.7% | 66.4% | 58.7% | 60.7% |
| c | 76.9% | 75.5% | 73.3% | 69.4% | 64.9% |
| d | 74.0% | 72.5% | 73.2% | 60.5% | 60.0% |
| e | 75.2% | 75.1% | 73.9% | 69.5% | 64.2% |
Settings chosen in the folds (most often):
- a: in-sample {}; folds {} ×50
- a1: in-sample {"lam":10}; folds {"lam":10} ×32, {"lam":0} ×3, {"lam":0.003} ×3, {"lam":0.03} ×3, {"lam":1} ×3
- b: in-sample {"kind":"lasso","lam":0.01}; folds {"kind":"ridge","lam":0.1} ×9, {"kind":"ridge","lam":0.3} ×8, {"kind":"ridge","lam":3} ×8, {"kind":"ridge","lam":1} ×6, {"kind":"lasso","lam":0.1} ×4
- b2: in-sample {"kind":"ridge","lam":0.3}; folds {"kind":"ridge","lam":3} ×26, {"kind":"ridge","lam":1} ×11, {"kind":"ridge","lam":0.3} ×7, {"kind":"ridge","lam":0.1} ×4, {"kind":"ridge","lam":0.003} ×1
- c: in-sample {"lam":0.03}; folds {"lam":0.0003} ×18, {"lam":0.001} ×12, {"lam":0.03} ×5, {"lam":0.003} ×5, {"lam":0.01} ×5
- d: in-sample {"depth":1,"colsample":0.3,"trees":400}; folds {"depth":1,"colsample":0.3,"trees":25} ×10, {"depth":3,"colsample":0.3,"trees":25} ×5, {"depth":1,"colsample":1,"trees":400} ×4, {"depth":3,"colsample":1,"trees":25} ×3, {"depth":1,"colsample":1,"trees":100} ×3
- e: in-sample {"k":3}; folds {"k":3} ×22, {"k":2} ×15, {"k":5} ×7, {"k":8} ×3, {"k":12} ×3
Flexible against the ledger, held out (paired; SE over repeats; bootstrap over units SD and the share of resamples at or below 0):
| Pair | vs consensus | vs lists | bootstrap SD | P(≤ 0) | scrambled SD |
|---|---|---|---|---|---|
| a1 − a | -2.6 ± 0.8 | -3.1 ± 0.9 | 1.7 | 1.00 | 2.1 |
| b − a | -4.4 ± 1.1 | -5.9 ± 1.0 | 4.5 | 0.82 | 10.1 |
| b2 − a | -7.1 ± 1.2 | -9.3 ± 1.2 | 5.5 | 0.89 | 9.2 |
| c − a | -1.4 ± 0.8 | -3.0 ± 0.7 | 4.5 | 0.63 | 9.8 |
| d − a | -5.8 ± 1.0 | -7.0 ± 0.9 | 3.9 | 0.94 | 5.6 |
| e − a | -2.0 ± 0.8 | -3.5 ± 1.1 | 5.0 | 0.64 | 9.9 |
| b − a1 | -1.8 ± 1.2 | -2.7 ± 1.4 | 3.7 | 0.67 | 9.7 |
| b2 − a1 | -4.5 ± 1.4 | -6.2 ± 1.6 | 4.8 | 0.83 | 8.6 |
| c − a1 | 1.3 ± 1.0 | 0.1 ± 1.2 | 3.7 | 0.39 | 9.5 |
| d − a1 | -3.2 ± 1.4 | -3.8 ± 1.6 | 3.0 | 0.86 | 5.1 |
| e − a1 | 0.6 ± 1.3 | -0.3 ± 1.7 | 4.3 | 0.48 | 9.8 |
Permutation importance, held out, model c (drop in pairwise vs consensus, points):
| Variable | Drop | SE |
|---|---|---|
| k6_light_vehicle | 1.64 | 0.19 |
| kw_is_fly | 1.34 | 0.22 |
| kw_is_transport | 1.13 | 0.11 |
| kw_heavy | 1.08 | 0.18 |
| kw_is_burrowers | 0.95 | 0.18 |
| kw_precision | 0.94 | 0.17 |
| lands | 0.91 | 0.20 |
| valueCheap | 0.80 | 0.24 |
| inv | 0.67 | 0.19 |
| kw_is_titanic | 0.61 | 0.10 |
| value | 0.58 | 0.29 |
| k6_monster | 0.55 | 0.15 |
| kw_is_battleline | 0.51 | 0.19 |
| bestR_A | 0.49 | 0.19 |
| kw_psychic | 0.47 | 0.15 |
| Family | Drop | SE |
|---|---|---|
| flags | 5.79 | 0.63 |
| ledger | 3.37 | 0.67 |
| weapons | 2.74 | 0.50 |
| killclass | 1.30 | 0.34 |
| datasheet | 0.66 | 0.47 |
| interaction | 0.44 | 0.20 |
Permutation importance, held out, model d (drop in pairwise vs consensus, points):
| Variable | Drop | SE |
|---|---|---|
| c_score | 2.41 | 0.59 |
| value | 2.01 | 0.48 |
| lands | 0.97 | 0.30 |
| valueCheap | 0.68 | 0.26 |
| k6_light_vehicle | 0.67 | 0.24 |
| c_hold | 0.56 | 0.15 |
| k6_character | 0.27 | 0.13 |
| ix_arrivexMelee | 0.27 | 0.23 |
| k6_monster | 0.26 | 0.20 |
| kw_heavy | 0.23 | 0.18 |
| kw_psychic | 0.15 | 0.10 |
| line_killMelee | 0.14 | 0.14 |
| modelsCheap | 0.13 | 0.05 |
| M | 0.11 | 0.13 |
| ocPer100 | 0.10 | 0.11 |
| Family | Drop | SE |
|---|---|---|
| ledger | 12.50 | 1.02 |
| killclass | 0.72 | 0.50 |
| datasheet | -0.04 | 0.33 |
| flags | -0.07 | 0.21 |
| interaction | -0.50 | 0.35 |
| weapons | -0.66 | 0.30 |
Permutation importance, held out, model a1 (drop in pairwise vs consensus, points):
| Variable | Drop | SE |
|---|---|---|
| c_kill | 19.78 | 1.25 |
| c_presence | 2.16 | 0.56 |
| c_score | 0.91 | 0.18 |
| c_hold | 0.78 | 0.22 |
| c_soak | 0.03 | 0.27 |
| c_actions | 0.01 | 0.02 |
| M | 0.00 | 0.00 |
| T | 0.00 | 0.00 |
| Sv | 0.00 | 0.00 |
| inv | 0.00 | 0.00 |
| W | 0.00 | 0.00 |
| Ld | 0.00 | 0.00 |
| OC | 0.00 | 0.00 |
| models | 0.00 | 0.00 |
| pts | 0.00 | 0.00 |
| Family | Drop | SE |
|---|---|---|
| ledger | 22.41 | 1.40 |
| datasheet | 0.00 | 0.00 |
| weapons | 0.00 | 0.00 |
| flags | 0.00 | 0.00 |
| killclass | 0.00 | 0.00 |
| interaction | 0.00 | 0.00 |
Partial dependence, model c (fit on all 52; consensus scale; biggest step):
- k6_light_vehicle: range 0.13 letters; biggest step +0.02 between 17.922 and 22.099; curve 0→2.93, 0.96→2.94, 2.045→2.95, 4.396→2.96, 7.602→2.98, 9.287→2.99, 12.061→3.00, 14.028→3.01, 16.005→3.02, 17.922→3.04, 22.099→3.06
- kw_is_fly: range 0.11 letters; biggest step -0.11 between 0 and 1; curve 0→3.02, 1→2.91
- kw_is_transport: range 0.17 letters; biggest step -0.17 between 0 and 1; curve 0→3.00, 1→2.83
- kw_heavy: range 0.13 letters; biggest step +0.13 between 0 and 1; curve 0→2.97, 1→3.10
- kw_is_burrowers: range 0.14 letters; biggest step +0.14 between 0 and 1; curve 0→2.98, 1→3.12
- kw_precision: range 0.12 letters; biggest step +0.12 between 0 and 1; curve 0→2.97, 1→3.09
- lands: range 0.21 letters; biggest step +0.14 between 0 and 0.66; curve 0→2.81, 0.66→2.95, 0.742→2.97, 0.786→2.98, 0.924→3.01, 0.988→3.02, 1→3.02
- valueCheap: range 0.20 letters; biggest step +0.05 between 2133.618 and 3017.884; curve 10.796→2.91, 441.848→2.94, 617.075→2.95, 676.313→2.95, 876.142→2.96, 1006.336→2.97, 1135.23→2.98, 1481.212→3.00, 2133.618→3.04, 3017.884→3.09, 3422.068→3.11
- inv: range 0.10 letters; biggest step -0.03 between 4 and 5; curve 4→3.06, 5→3.03, 6→2.99, 7→2.96
- kw_is_titanic: range 0.13 letters; biggest step -0.13 between 0 and 1; curve 0→3.00, 1→2.87
- value: range 0.21 letters; biggest step +0.06 between 2133.618 and 3076.76; curve 10.796→2.91, 479.387→2.94, 644.868→2.95, 731.117→2.95, 960.284→2.96, 1019.608→2.97, 1135.23→2.98, 1481.212→3.00, 2133.618→3.04, 3076.76→3.09, 3422.068→3.12
- k6_monster: range 0.08 letters; biggest step +0.02 between 12.058 and 30.133; curve 0→2.96, 0.039→2.96, 0.187→2.96, 0.534→2.97, 1.829→2.98, 2.475→2.99, 3.56→2.99, 4.62→3.00, 8.905→3.01, 12.058→3.02, 30.133→3.04
Partial dependence, model d (fit on all 52; consensus scale; biggest step):
- c_score: range 0.57 letters; biggest step +0.33 between 12.606 and 13.904; curve 0→2.61, 9.455→2.85, 11.587→2.85, 12.606→2.85, 13.904→3.17, 17.176→3.17, 22.034→3.17, 26.182→3.17, 33.279→3.17, 85.881→3.17, 209.932→3.17
- value: range 0.54 letters; biggest step +0.30 between 3076.76 and 3422.068; curve 10.796→2.82, 479.387→2.82, 644.868→2.82, 731.117→2.91, 960.284→2.91, 1019.608→3.02, 1135.23→3.06, 1481.212→3.06, 2133.618→3.06, 3076.76→3.06, 3422.068→3.36
- lands: range 0.41 letters; biggest step +0.26 between 0.924 and 0.988; curve 0→2.75, 0.66→2.75, 0.742→2.75, 0.786→2.89, 0.924→2.89, 0.988→3.16, 1→3.16
- valueCheap: range 0.52 letters; biggest step +0.31 between 3017.884 and 3422.068; curve 10.796→2.87, 441.848→2.87, 617.075→2.87, 676.313→2.87, 876.142→2.92, 1006.336→2.92, 1135.23→3.03, 1481.212→3.03, 2133.618→3.03, 3017.884→3.07, 3422.068→3.39
- k6_light_vehicle: range 0.32 letters; biggest step +0.15 between 12.061 and 14.028; curve 0→2.86, 0.96→2.86, 2.045→2.86, 4.396→2.86, 7.602→2.86, 9.287→2.89, 12.061→3.03, 14.028→3.18, 16.005→3.18, 17.922→3.18, 22.099→3.18
- c_hold: range 0.00 letters; biggest step +0.00 between 0 and 9.496; curve 0→2.99, 9.496→2.99, 16.777→2.99, 22.551→2.99, 28.013→2.99, 32.485→2.99, 43.27→2.99, 49.041→2.99, 72.508→2.99, 120.975→2.99, 344.469→2.99
- k6_character: range 0.17 letters; biggest step +0.17 between 0 and 28.046; curve 0→2.83, 28.046→3.00, 33.848→3.00, 44.935→3.00, 50.724→3.00, 61.857→3.00, 68.975→3.00, 72.744→3.00, 88.538→3.00, 95.222→3.00, 122.294→3.00
- ix_arrivexMelee: range 0.20 letters; biggest step +0.20 between 0 and 16.4; curve 0→2.85, 16.4→3.05, 21.033→3.05, 30.979→3.05, 62.253→3.05, 82.392→3.05, 92.285→3.05, 100.223→3.05, 130.103→3.05, 134.006→3.05, 170.194→3.05
- k6_monster: range 0.16 letters; biggest step +0.16 between 0.187 and 0.534; curve 0→2.87, 0.039→2.87, 0.187→2.87, 0.534→3.03, 1.829→3.03, 2.475→3.03, 3.56→3.03, 4.62→3.03, 8.905→3.03, 12.058→3.03, 30.133→3.03
- kw_heavy: range 0.39 letters; biggest step +0.39 between 0 and 1; curve 0→2.95, 1→3.34
- kw_psychic: range 0.40 letters; biggest step +0.40 between 0 and 1; curve 0→2.95, 1→3.35, 2→3.35, 3→3.35
- line_killMelee: range 1.00 letters; biggest step +0.63 between 0 and 17.172; curve 0→2.37, 17.172→2.99, 23.332→2.99, 40.942→2.99, 62.272→2.99, 82.392→2.99, 92.285→2.99, 110.161→3.00, 130.103→3.12, 134.006→3.37, 170.194→3.37
The trees' splits (fit on all 52): variable, gain, median threshold, uses:
- valueCheap: gain 1045.3, threshold ~3047.439 (685.486 to 3047.439), 14 splits
- value: gain 919.5, threshold ~3076.877 (685.486 to 3110.314), 16 splits
- c_score: gain 723.5, threshold ~13.613 (6.201 to 13.613), 26 splits
- lands: gain 474.1, threshold ~0.956 (0.764 to 0.956), 21 splits
- ix_bodiesxOC: gain 346.1, threshold ~30.404 (0.694 to 30.404), 6 splits
- ix_MxMelee: gain 278.6, threshold ~5.198 (4.28 to 7.877), 15 splits
- line_killMelee: gain 250.7, threshold ~105.192 (5.25 to 133.593), 42 splits
- OC: gain 249.4, threshold ~0.5 (0.5 to 0.5), 1 splits
- k6_light_vehicle: gain 232.0, threshold ~13.027 (8.759 to 13.027), 21 splits
- kw_heavy: gain 214.6, threshold ~0.5 (0.5 to 0.5), 13 splits
- k6_character: gain 131.4, threshold ~27.106 (27.106 to 27.106), 4 splits
- c_presence: gain 130.4, threshold ~14.037 (14.037 to 14.037), 32 splits
The trees' interactions (a depth-3 fit on all 52, 100 trees: one split under the other, by the child's gain):
- c_score × value: 735.7
- c_score × k6_cavalry_and_beasts: 158.4
- lands × value: 139.1
- k6_monster × value: 112.5
- c_score × valueCheap: 108.7
- kw_heavy × lands: 87.5
- ix_invxT × k6_light_vehicle: 82.3
- inv × k6_light_vehicle: 82.1
- c_score × ix_invxT: 63.4
- ix_MxMelee × kw_heavy: 61.2
- c_score × lands: 60.8
- ptsPerModel × value: 57.2
Linear (b) fitted on all 52, largest standardised weights:
- k6_light_vehicle: 0.708
- lands: 0.630
- kw_is_battleline: 0.629
- bestM_skill: 0.590
- kw_indirect: 0.554
- inv: -0.473
- kw_is_transport: -0.433
- fx_list: -0.413
- kw_is_fly: -0.403
- value: 0.361
- kw_precision: 0.342
- kw_heavy: 0.331
- kw_psychic: 0.323
- meleeVeh100: -0.304
- kw_ignoresCover: 0.292
- rangedAttacks100: -0.224
- kw_is_burrowers: 0.224
- kw_anti: -0.216
- fx_mark: 0.193
- kw_is_infantry: 0.178
The ledger's seven weights re-fitted on all 52 (a1):
- kill: step 6 498.17 → 497.97 (×1.00)
- soak: step 6 14.97 → 14.97 (×1.00)
- score: step 6 9.65 → 9.65 (×1.00)
- actions: step 6 1.31 → 1.31 (×1.00)
- hold: step 6 18.89 → 18.89 (×1.00)
- spawn: step 6 15.65 → 15.65 (×1.00)
- presence: step 6 2.52 → 2.52 (×1.00)
Units the best model (c) gets right that the ledger gets wrong (held-out share of the unit's pairs ordered right):
| Unit | Consensus | Consensus rank | Ledger rank | Best model | Ledger | Gain | Top variables (percentile) |
|---|---|---|---|---|---|---|---|
| Toxicrene | 1.333 | 46.5 | 23 | 88.0% | 53.4% | 35 | k6_light_vehicle 33, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0 |
| Harpy | 0.8 | 50.5 | 33 | 100.0% | 65.6% | 34 | k6_light_vehicle 29, kw_is_fly 76, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0 |
| Gargoyles | 3.6 | 19 | 47 | 59.6% | 41.2% | 18 | k6_light_vehicle 24, kw_is_fly 76, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0 |
| Neurotyrant | 3.667 | 17 | 39 | 70.3% | 52.3% | 18 | k6_light_vehicle 12, kw_is_fly 76, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0 |
| Tyranid Warriors with Ranged Bio-Weapons | 1.8 | 43 | 13 | 59.6% | 44.0% | 16 | k6_light_vehicle 18, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0 |
| Norn Emissary | 4.2 | 11 | 19 | 76.5% | 61.9% | 15 | k6_light_vehicle 51, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0 |
| Tyrant Guard | 2.8 | 32.5 | 49 | 79.3% | 65.4% | 14 | k6_light_vehicle 37, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0 |
| Tervigon | 2 | 40 | 16 | 53.4% | 40.0% | 13 | k6_light_vehicle 53, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0 |
| Genestealers | 3.8 | 14 | 34 | 66.3% | 55.3% | 11 | k6_light_vehicle 49, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0 |
| Carnifexes | 2 | 40 | 30 | 82.3% | 71.9% | 10 | k6_light_vehicle 96, kw_is_fly 0, kw_is_transport 0, kw_heavy 0, kw_is_burrowers 0 |
And the reverse (the ledger right, the model wrong):
- Biovores: consensus 5 (rank 1.5), ledger rank 5; model 6.2% against ledger 92.4%
- Parasite of Mortrex: consensus 2.2 (rank 36.5), ledger rank 43; model 53.0% against ledger 75.2%
- Psychophage: consensus 3 (rank 28.5), ledger rank 27; model 65.7% against ledger 87.3%
- Zoanthropes: consensus 3.6 (rank 19), ledger rank 4; model 61.6% against ledger 76.5%
- Termagants: consensus 3 (rank 28.5), ledger rank 36; model 65.0% against ledger 79.8%
- Lictor: consensus 4.8 (rank 3), ledger rank 2; model 77.1% against ledger 91.2%
- The Red Terror: consensus 4.333 (rank 9), ledger rank 12; model 69.8% against ledger 83.9%
- Exocrine: consensus 4.4 (rank 6.5), ledger rank 20; model 57.2% against ledger 71.2%
Every flexible model misses the same way (mean held-out prediction minus consensus, letters):
- Biovores: -3.372 (consensus 5; S/S/S/S/S)
- Neurolictor: -1.844 (consensus 5; S/S/S/–/S)
- Tyranid Warriors with Ranged Bio-Weapons: +1.598 (consensus 1.8; C/C/C/D/C)
- Tervigon: +1.491 (consensus 2; B/C/C/C/D)
- Toxicrene: +1.474 (consensus 1.333; C/–/–/D/D)
- Tyrannofex: -1.454 (consensus 4.4; S/A/A/A/S)
- Exocrine: -1.414 (consensus 4.4; A/A/A/S/S)
- Hierophant: +1.333 (consensus 1.5; D/C/C/D/–)
- Hormagaunts: -1.271 (consensus 4.2; S/A/B/A/S)
- Tyrannocyte: +1.195 (consensus 1.333; C/–/–/D/D)
- Parasite of Mortrex: +1.193 (consensus 2.2; B/B/C/C/D)
- Lictor: -1.175 (consensus 4.8; S/A/S/S/S)