Ceiling (each list against the mean of the other four): 88.5%
| Model | In-sample vs consensus (tuned) | In-sample, most flexible | In-sample vs lists | Held out vs consensus (± SE) | Held out vs lists (± SE) | Held out, pooled across folds (consensus / lists) | Scrambled vs consensus (SD) | Scrambled vs lists (SD) |
|---|---|---|---|---|---|---|---|---|
| (a) the ledger as it is (step 6) | 73.3% | 73.3% | 74.5% | 72.9% ± 0.4 | 74.1% ± 0.6 | 70.2% / 71.4% | 50.8% (4.5) | 51.7% (3.7) |
| (a1) the ledger's form, its seven weights re-fitted | 73.3% | 73.8% | 74.5% | 69.7% ± 1.0 | 70.1% ± 1.0 | 66.8% / 67.5% | 49.1% (3.8) | 50.7% (3.9) |
| (b) linear on every variable (ridge or lasso) | 89.1% | 98.8% | 87.7% | 69.7% ± 1.3 | 68.6% ± 0.9 | 69.2% / 68.2% | 45.9% (6.9) | 46.5% (7.5) |
| (b2) linear on a short core list (25 variables), ridge | 72.2% | 81.7% | 71.0% | 66.0% ± 0.9 | 64.2% ± 0.8 | 64.1% / 62.4% | 46.5% (8.4) | 47.1% (9.2) |
| (c) linear plus every pairwise interaction (kernel, ridge) | 97.2% | 97.2% | 92.2% | 73.3% ± 0.9 | 71.4% ± 1.0 | 71.5% / 70.6% | 44.8% (6.6) | 45.8% (7.0) |
| (d) boosted trees, monotone where the sign is obvious | 95.3% | 99.7% | 91.5% | 68.3% ± 0.9 | 67.6% ± 0.8 | 66.6% / 66.5% | 49.3% (3.8) | 50.1% (4.5) |
| (e) nearest neighbours (similar units) | 100.0% | 100.0% | 92.3% | 71.9% ± 0.8 | 70.1% ± 1.1 | 69.7% / 68.4% | 45.2% (7.5) | 46.3% (7.5) |
Learning curve (trained on m units, scored on the pairs among the rest; ± SE over splits):
- c: 16: 65.5% ± 1.7, 21: 67.4% ± 2.0, 26: 70.3% ± 1.3, 31: 69.7% ± 1.7, 36: 75.0% ± 1.7, 41: 71.4% ± 2.6
- a1: 16: 71.8% ± 1.1, 21: 69.7% ± 1.7, 26: 69.8% ± 0.9, 31: 69.7% ± 1.9, 36: 70.7% ± 2.4, 41: 70.7% ± 2.9
- a: 16: 73.4% ± 1.0, 21: 73.8% ± 0.9, 26: 71.6% ± 1.2, 31: 74.2% ± 1.9, 36: 74.1% ± 1.5, 41: 73.3% ± 2.6
- b2: 16: 60.7% ± 1.8, 21: 61.6% ± 1.7, 26: 63.1% ± 1.5, 31: 62.2% ± 2.1, 36: 67.1% ± 2.7, 41: 65.8% ± 3.3
Per list, held out:
| Model | auspex | hivemind | second | maelstrom | astrategas |
|---|---|---|---|---|---|
| ceiling | 91.7% | 87.7% | 81.9% | 94.9% | 86.5% |
| a | 79.9% | 78.4% | 76.2% | 70.8% | 65.1% |
| a1 | 76.7% | 74.3% | 71.9% | 66.4% | 61.3% |
| b | 75.4% | 68.1% | 69.5% | 65.3% | 64.8% |
| b2 | 73.2% | 65.0% | 63.8% | 58.3% | 60.6% |
| c | 77.9% | 71.7% | 71.2% | 69.9% | 66.5% |
| d | 76.0% | 69.0% | 72.5% | 61.4% | 59.1% |
| e | 76.6% | 71.2% | 69.1% | 69.4% | 64.2% |
Settings chosen in the folds (most often):
- a: in-sample {}; folds {} ×50
- a1: in-sample {"lam":10}; folds {"lam":10} ×31, {"lam":0} ×5, {"lam":0.03} ×4, {"lam":0.0001} ×3, {"lam":0.003} ×3
- b: in-sample {"kind":"ridge","lam":0.3}; folds {"kind":"ridge","lam":1} ×12, {"kind":"ridge","lam":0.3} ×8, {"kind":"ridge","lam":0.1} ×6, {"kind":"ridge","lam":0.003} ×5, {"kind":"ridge","lam":0.03} ×4
- b2: in-sample {"kind":"ridge","lam":1}; folds {"kind":"ridge","lam":3} ×20, {"kind":"ridge","lam":1} ×10, {"kind":"ridge","lam":0.3} ×8, {"kind":"ridge","lam":0.1} ×6, {"kind":"ridge","lam":0.003} ×4
- c: in-sample {"lam":0.0003}; folds {"lam":0.0003} ×19, {"lam":0.001} ×10, {"lam":0.003} ×10, {"lam":0.01} ×5, {"lam":0.03} ×4
- d: in-sample {"depth":1,"colsample":1,"trees":400}; folds {"depth":1,"colsample":1,"trees":400} ×6, {"depth":2,"colsample":1,"trees":100} ×4, {"depth":3,"colsample":1,"trees":100} ×4, {"depth":1,"colsample":0.3,"trees":25} ×3, {"depth":1,"colsample":1,"trees":100} ×3
- e: in-sample {"k":3}; folds {"k":3} ×18, {"k":2} ×16, {"k":5} ×9, {"k":8} ×3, {"k":12} ×3
Flexible against the ledger, held out (paired; SE over repeats; bootstrap over units SD and the share of resamples at or below 0):
| Pair | vs consensus | vs lists | bootstrap SD | P(≤ 0) | scrambled SD |
|---|---|---|---|---|---|
| a1 − a | -3.2 ± 0.9 | -4.0 ± 1.0 | 2.2 | 0.99 | 2.3 |
| b − a | -3.2 ± 1.1 | -5.5 ± 1.0 | 4.3 | 0.74 | 9.9 |
| b2 − a | -7.0 ± 0.9 | -9.9 ± 0.8 | 5.8 | 0.89 | 10.2 |
| c − a | 0.3 ± 0.7 | -2.6 ± 0.7 | 4.7 | 0.47 | 9.5 |
| d − a | -4.6 ± 0.7 | -6.5 ± 0.6 | 4.0 | 0.91 | 6.1 |
| e − a | -1.1 ± 0.7 | -4.0 ± 0.8 | 4.8 | 0.63 | 11.1 |
| b − a1 | 0.0 ± 1.1 | -1.5 ± 1.1 | 3.4 | 0.54 | 8.6 |
| b2 − a1 | -3.7 ± 1.5 | -5.9 ± 1.5 | 5.0 | 0.76 | 9.1 |
| c − a1 | 3.6 ± 0.8 | 1.3 ± 0.9 | 3.6 | 0.17 | 8.7 |
| d − a1 | -1.4 ± 1.2 | -2.5 ± 1.3 | 3.0 | 0.70 | 5.9 |
| e − a1 | 2.2 ± 1.3 | -0.0 ± 1.4 | 3.9 | 0.36 | 10.4 |
Permutation importance, held out, model c (drop in pairwise vs consensus, points):
| Variable | Drop | SE |
|---|---|---|
| kw_is_battleline | 1.63 | 0.21 |
| kw_is_transport | 1.14 | 0.12 |
| kw_precision | 1.13 | 0.15 |
| kw_is_burrowers | 0.85 | 0.18 |
| k6_light_vehicle | 0.83 | 0.21 |
| inv | 0.81 | 0.15 |
| arrives | 0.81 | 0.22 |
| kw_psychic | 0.79 | 0.15 |
| kw_heavy | 0.71 | 0.17 |
| soakFight | 0.66 | 0.15 |
| valueCheap | 0.66 | 0.19 |
| lands | 0.60 | 0.19 |
| bestR_A | 0.59 | 0.12 |
| kw_is_fly | 0.57 | 0.24 |
| rangedAttacks100 | 0.43 | 0.13 |
| Family | Drop | SE |
|---|---|---|
| flags | 6.26 | 0.72 |
| weapons | 3.57 | 0.53 |
| ledger | 3.41 | 0.68 |
| killclass | 1.50 | 0.31 |
| datasheet | 0.98 | 0.37 |
| interaction | 0.57 | 0.23 |
Permutation importance, held out, model d (drop in pairwise vs consensus, points):
| Variable | Drop | SE |
|---|---|---|
| c_score | 4.70 | 0.46 |
| value | 1.80 | 0.42 |
| ix_arrivexMelee | 1.66 | 0.39 |
| k6_monster | 1.21 | 0.26 |
| valueCheap | 1.14 | 0.36 |
| lands | 1.11 | 0.29 |
| k6_light_vehicle | 0.88 | 0.20 |
| kw_precision | 0.51 | 0.22 |
| ix_MxMelee | 0.45 | 0.24 |
| kw_heavy | 0.35 | 0.20 |
| kw_psychic | 0.25 | 0.10 |
| k6_character | 0.23 | 0.10 |
| kw_ignoresCover | 0.20 | 0.07 |
| ocPer100 | 0.18 | 0.10 |
| c_kill | 0.15 | 0.16 |
| Family | Drop | SE |
|---|---|---|
| ledger | 12.41 | 1.05 |
| interaction | 1.92 | 0.47 |
| killclass | 0.74 | 0.45 |
| datasheet | 0.27 | 0.32 |
| weapons | 0.18 | 0.28 |
| flags | -0.18 | 0.14 |
Permutation importance, held out, model a1 (drop in pairwise vs consensus, points):
| Variable | Drop | SE |
|---|---|---|
| c_kill | 17.74 | 1.18 |
| c_score | 1.81 | 0.25 |
| c_presence | 1.65 | 0.53 |
| c_hold | 1.08 | 0.22 |
| c_soak | 0.39 | 0.24 |
| c_spawn | 0.10 | 0.07 |
| c_actions | 0.04 | 0.04 |
| M | 0.00 | 0.00 |
| T | 0.00 | 0.00 |
| Sv | 0.00 | 0.00 |
| inv | 0.00 | 0.00 |
| W | 0.00 | 0.00 |
| Ld | 0.00 | 0.00 |
| OC | 0.00 | 0.00 |
| models | 0.00 | 0.00 |
| Family | Drop | SE |
|---|---|---|
| ledger | 20.50 | 1.38 |
| datasheet | 0.00 | 0.00 |
| weapons | 0.00 | 0.00 |
| flags | 0.00 | 0.00 |
| killclass | 0.00 | 0.00 |
| interaction | 0.00 | 0.00 |
Partial dependence, model c (fit on all 52; consensus scale; biggest step):
- kw_is_battleline: range 0.36 letters; biggest step +0.36 between 0 and 1; curve 0→2.90, 1→3.26
- kw_is_transport: range 0.37 letters; biggest step -0.37 between 0 and 1; curve 0→2.96, 1→2.60
- kw_precision: range 0.23 letters; biggest step +0.23 between 0 and 1; curve 0→2.90, 1→3.13
- kw_is_burrowers: range 0.26 letters; biggest step +0.26 between 0 and 1; curve 0→2.92, 1→3.17
- k6_light_vehicle: range 0.20 letters; biggest step +0.04 between 17.922 and 22.099; curve 0→2.85, 0.96→2.86, 2.045→2.87, 4.396→2.89, 7.602→2.92, 9.287→2.93, 12.061→2.96, 14.028→2.98, 16.005→3.00, 17.922→3.01, 22.099→3.05
- inv: range 0.12 letters; biggest step -0.04 between 4 and 5; curve 4→3.03, 5→2.99, 6→2.94, 7→2.90
- arrives: range 0.15 letters; biggest step +0.15 between 0 and 1; curve 0→2.82, 1→2.98
- kw_psychic: range 0.25 letters; biggest step +0.12 between 0 and 1; curve 0→2.91, 1→3.04, 2→3.11, 3→3.17
- kw_heavy: range 0.23 letters; biggest step +0.23 between 0 and 1; curve 0→2.91, 1→3.13
- soakFight: range 0.29 letters; biggest step -0.11 between 240.496 and 371.127; curve 24.778→3.05, 85.552→3.00, 103.572→2.99, 124.392→2.97, 138.398→2.96, 152.874→2.95, 168.408→2.93, 191.973→2.91, 201.663→2.90, 240.496→2.87, 371.127→2.76
- valueCheap: range 0.22 letters; biggest step +0.06 between 2133.618 and 3017.884; curve 10.796→2.85, 441.848→2.88, 617.075→2.89, 676.313→2.90, 876.142→2.91, 1006.336→2.92, 1135.23→2.92, 1481.212→2.95, 2133.618→2.99, 3017.884→3.04, 3422.068→3.07
- lands: range 0.29 letters; biggest step +0.20 between 0 and 0.66; curve 0→2.69, 0.66→2.89, 0.742→2.91, 0.786→2.92, 0.924→2.96, 0.988→2.98, 1→2.98
Partial dependence, model d (fit on all 52; consensus scale; biggest step):
- c_score: range 0.96 letters; biggest step +0.46 between 0 and 9.455; curve 0→2.26, 9.455→2.72, 11.587→2.82, 12.606→2.82, 13.904→3.15, 17.176→3.22, 22.034→3.22, 26.182→3.22, 33.279→3.22, 85.881→3.22, 209.932→3.22
- value: range 0.73 letters; biggest step +0.41 between 3076.76 and 3422.068; curve 10.796→2.73, 479.387→2.73, 644.868→2.73, 731.117→2.83, 960.284→2.83, 1019.608→2.96, 1135.23→3.01, 1481.212→3.01, 2133.618→3.01, 3076.76→3.06, 3422.068→3.46
- ix_arrivexMelee: range 0.87 letters; biggest step +0.55 between 0 and 16.4; curve 0→2.52, 16.4→3.07, 21.033→3.07, 30.979→3.07, 62.253→3.07, 82.392→3.07, 92.285→3.07, 100.223→3.07, 130.103→3.20, 134.006→3.39, 170.194→3.39
- k6_monster: range 0.38 letters; biggest step +0.22 between 0.187 and 0.534; curve 0→2.74, 0.039→2.74, 0.187→2.74, 0.534→2.96, 1.829→2.96, 2.475→2.96, 3.56→2.96, 4.62→2.96, 8.905→2.96, 12.058→3.11, 30.133→3.11
- valueCheap: range 0.38 letters; biggest step +0.33 between 3017.884 and 3422.068; curve 10.796→2.90, 441.848→2.90, 617.075→2.90, 676.313→2.90, 876.142→2.90, 1006.336→2.90, 1135.23→2.90, 1481.212→2.90, 2133.618→2.90, 3017.884→2.95, 3422.068→3.28
- lands: range 0.23 letters; biggest step +0.11 between 0.742 and 0.786; curve 0→2.79, 0.66→2.79, 0.742→2.79, 0.786→2.91, 0.924→2.91, 0.988→3.02, 1→3.02
- k6_light_vehicle: range 0.23 letters; biggest step +0.11 between 9.287 and 12.061; curve 0→2.84, 0.96→2.84, 2.045→2.84, 4.396→2.84, 7.602→2.84, 9.287→2.86, 12.061→2.98, 14.028→3.07, 16.005→3.07, 17.922→3.07, 22.099→3.07
- kw_precision: range 0.31 letters; biggest step +0.31 between 0 and 1; curve 0→2.90, 1→3.21
- ix_MxMelee: range 0.20 letters; biggest step +0.18 between 6.089 and 7.623; curve 0→2.85, 0.65→2.85, 1.224→2.85, 1.96→2.85, 3.172→2.85, 4.174→2.85, 5.055→2.85, 6.089→2.87, 7.623→3.05, 8→3.05, 12→3.05
- kw_heavy: range 0.61 letters; biggest step +0.61 between 0 and 1; curve 0→2.88, 1→3.48
- kw_psychic: range 0.43 letters; biggest step +0.43 between 0 and 1; curve 0→2.89, 1→3.32, 2→3.32, 3→3.32
- k6_character: range 0.09 letters; biggest step +0.09 between 0 and 28.046; curve 0→2.85, 28.046→2.94, 33.848→2.94, 44.935→2.94, 50.724→2.94, 61.857→2.94, 68.975→2.94, 72.744→2.94, 88.538→2.94, 95.222→2.94, 122.294→2.94
The trees' splits (fit on all 52): variable, gain, median threshold, uses:
- c_score: gain 1835.9, threshold ~13.613 (6.201 to 14.189), 35 splits
- value: gain 1469.8, threshold ~3076.877 (685.486 to 3110.314), 19 splits
- ix_arrivexMelee: gain 619.5, threshold ~11.492 (5.25 to 131.936), 43 splits
- valueCheap: gain 283.5, threshold ~3047.439 (2421.886 to 3110.314), 11 splits
- kw_heavy: gain 267.2, threshold ~0.5 (0.5 to 0.5), 23 splits
- k6_monster: gain 220.5, threshold ~0.192 (0.187 to 9.43), 19 splits
- ix_MxMelee: gain 191.3, threshold ~6.457 (5.198 to 6.979), 11 splits
- k6_light_vehicle: gain 180.1, threshold ~13.027 (8.759 to 13.027), 14 splits
- kw_precision: gain 155.6, threshold ~0.5 (0.5 to 0.5), 11 splits
- kw_psychic: gain 88.6, threshold ~0.5 (0.5 to 0.5), 18 splits
- c_presence: gain 86.1, threshold ~14.037 (14.037 to 14.037), 26 splits
- inv: gain 85.1, threshold ~4.5 (4.5 to 4.5), 2 splits
The trees' interactions (a depth-3 fit on all 52, 100 trees: one split under the other, by the child's gain):
- c_score × value: 752.3
- k6_monster × value: 212.0
- c_score × lands: 170.9
- arrives × value: 166.1
- inv × k6_monster: 151.8
- ix_arrivexMelee × k6_monster: 138.5
- c_score × k6_monster: 115.9
- c_score × inv: 89.6
- ix_MxMelee × kw_heavy: 88.8
- c_score × k6_cavalry_and_beasts: 75.5
- lands × value: 75.0
- ix_durHeavy100 × value: 73.1
Linear (b) fitted on all 52, largest standardised weights:
- kw_indirect: 0.123
- kw_heavy: 0.111
- value: 0.101
- kw_is_transport: -0.101
- kw_is_battleline: 0.101
- valueCheap: 0.097
- lands: 0.095
- kw_is_fly: -0.093
- kw_precision: 0.093
- arrives: 0.091
- kw_is_burrowers: 0.082
- k6_light_vehicle: 0.079
- kw_anti: -0.078
- inv: -0.075
- kw_psychic: 0.074
- lr_movement: -0.074
- bestM_skill: 0.070
- soakFight: -0.070
- c_score: 0.068
- bestR_A: -0.067
The ledger's seven weights re-fitted on all 52 (a1):
- kill: step 6 498.17 → 497.82 (×1.00)
- soak: step 6 14.97 → 14.97 (×1.00)
- score: step 6 9.65 → 9.65 (×1.00)
- actions: step 6 1.31 → 1.31 (×1.00)
- hold: step 6 18.89 → 18.89 (×1.00)
- spawn: step 6 15.65 → 15.65 (×1.00)
- presence: step 6 2.52 → 2.52 (×1.00)
Units the best model (c) gets right that the ledger gets wrong (held-out share of the unit's pairs ordered right):
| Unit | Consensus | Consensus rank | Ledger rank | Best model | Ledger | Gain | Top variables (percentile) |
|---|---|---|---|---|---|---|---|
| Toxicrene | 1 | 48 | 23 | 89.2% | 50.7% | 39 | kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 33 |
| Harpy | 0.667 | 51.5 | 33 | 93.6% | 61.3% | 32 | kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 29 |
| Gargoyles | 3.75 | 15 | 47 | 65.6% | 35.2% | 30 | kw_is_battleline 96, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 24 |
| Tyranid Prime with Lash Whip | 2.5 | 35 | 14 | 71.0% | 51.3% | 20 | kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 82 |
| Tyrant Guard | 2.8 | 29.5 | 49 | 78.0% | 60.4% | 18 | kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 37 |
| Genestealers | 3.75 | 15 | 34 | 70.8% | 53.6% | 17 | kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 49 |
| Neurotyrant | 3.667 | 17 | 39 | 64.0% | 47.2% | 17 | kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 12 |
| Norn Emissary | 4.2 | 11 | 19 | 75.6% | 61.2% | 14 | kw_is_battleline 0, kw_is_transport 0, kw_precision 90, kw_is_burrowers 0, k6_light_vehicle 51 |
| Carnifexes | 2 | 40.5 | 30 | 89.2% | 74.7% | 14 | kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 96 |
| Tyranid Warriors with Ranged Bio-Weapons | 2 | 40.5 | 13 | 61.5% | 48.9% | 13 | kw_is_battleline 0, kw_is_transport 0, kw_precision 0, kw_is_burrowers 0, k6_light_vehicle 18 |
And the reverse (the ledger right, the model wrong):
- Biovores: consensus 5 (rank 1), ledger rank 5; model 7.2% against ledger 92.6%
- Psychophage: consensus 3 (rank 27.5), ledger rank 27; model 66.1% against ledger 83.1%
- Parasite of Mortrex: consensus 2.2 (rank 36.5), ledger rank 43; model 58.3% against ledger 75.2%
- Lictor: consensus 4.75 (rank 2), ledger rank 2; model 81.4% against ledger 96.7%
- Tyrannofex: consensus 4.4 (rank 6), ledger rank 25; model 54.4% against ledger 68.9%
- Hive Tyrant: consensus 4.4 (rank 6), ledger rank 3; model 81.4% against ledger 95.7%
- The Red Terror: consensus 4.333 (rank 8), ledger rank 12; model 67.9% against ledger 80.9%
- Winged Hive Tyrant: consensus 3.2 (rank 25), ledger rank 10; model 65.0% against ledger 75.6%
Every flexible model misses the same way (mean held-out prediction minus consensus, letters):
- Biovores: -3.357 (consensus 5; S/S/S/S/S)
- Toxicrene: +1.76 (consensus 1; –/–/–/D/D)
- Tyrannofex: -1.461 (consensus 4.4; S/A/A/A/S)
- Hierophant: +1.457 (consensus 1.333; D/C/–/D/–)
- Tervigon: +1.447 (consensus 2; B/C/C/C/D)
- Exocrine: -1.414 (consensus 4.4; A/A/A/S/S)
- Neurolictor: -1.372 (consensus 4.5; S/S/B/–/S)
- Lictor: -1.35 (consensus 4.75; –/A/S/S/S)
- Tyranid Warriors with Ranged Bio-Weapons: +1.288 (consensus 2; C/C/B/D/C)
- Deathleaper: +1.253 (consensus 3.2; A/S/B/B/D)
- Hormagaunts: -1.187 (consensus 4.2; S/A/B/A/S)
- Tyrannocyte: +1.17 (consensus 1.333; C/–/–/D/D)