Ceiling (each list against the mean of the other four): 88.5%
| Model | In-sample vs consensus (tuned) | In-sample, most flexible | In-sample vs lists | Held out vs consensus (± SE) | Held out vs lists (± SE) | Held out, pooled across folds (consensus / lists) | Scrambled vs consensus (SD) | Scrambled vs lists (SD) |
|---|---|---|---|---|---|---|---|---|
| (a) the ledger as it is (step 6) | 73.3% | 73.3% | 74.5% | 72.9% ± 0.4 | 74.1% ± 0.6 | 70.2% / 71.4% | 50.8% (4.5) | 51.7% (3.7) |
| (a1) the ledger's form, its seven weights re-fitted | 73.3% | 73.8% | 74.5% | 69.7% ± 1.0 | 70.1% ± 1.0 | 66.8% / 67.5% | 49.1% (3.8) | 50.7% (3.9) |
| (b) linear on every variable (ridge or lasso) | 99.1% | 99.1% | 92.6% | 70.0% ± 1.0 | 68.6% ± 1.1 | 69.1% / 67.8% | 45.8% (8.0) | 46.4% (8.2) |
| (b2) linear on a short core list (25 variables), ridge | 72.2% | 81.7% | 71.0% | 66.0% ± 0.9 | 64.2% ± 0.8 | 64.1% / 62.4% | 46.5% (8.4) | 47.1% (9.2) |
| (c) linear plus every pairwise interaction (kernel, ridge) | 96.5% | 96.5% | 92.1% | 72.6% ± 0.7 | 70.6% ± 0.9 | 70.8% / 69.5% | 43.0% (7.0) | 43.9% (7.3) |
| (d) boosted trees, monotone where the sign is obvious | 96.0% | 99.8% | 91.8% | 67.2% ± 0.8 | 66.7% ± 0.8 | 65.9% / 66.0% | 50.2% (4.9) | 50.8% (4.8) |
| (e) nearest neighbours (similar units) | 100.0% | 100.0% | 92.4% | 64.6% ± 1.4 | 62.0% ± 1.4 | 61.6% / 59.6% | 43.6% (8.5) | 44.6% (9.1) |
Learning curve (trained on m units, scored on the pairs among the rest; ± SE over splits):
- c: 16: 62.8% ± 1.8, 21: 64.1% ± 2.0, 26: 68.6% ± 1.2, 31: 68.8% ± 2.2, 36: 73.3% ± 2.1, 41: 70.9% ± 2.6
- a1: 16: 71.8% ± 1.1, 21: 69.7% ± 1.7, 26: 69.8% ± 0.9, 31: 69.7% ± 1.9, 36: 70.7% ± 2.4, 41: 70.7% ± 2.9
- a: 16: 73.4% ± 1.0, 21: 73.8% ± 0.9, 26: 71.6% ± 1.2, 31: 74.2% ± 1.9, 36: 74.1% ± 1.5, 41: 73.3% ± 2.6
- b2: 16: 60.7% ± 1.8, 21: 61.6% ± 1.7, 26: 63.1% ± 1.5, 31: 62.2% ± 2.1, 36: 67.1% ± 2.7, 41: 65.8% ± 3.3
Per list, held out:
| Model | auspex | hivemind | second | maelstrom | astrategas |
|---|---|---|---|---|---|
| ceiling | 91.7% | 87.7% | 81.9% | 94.9% | 86.5% |
| a | 79.9% | 78.4% | 76.2% | 70.8% | 65.1% |
| a1 | 76.7% | 74.3% | 71.9% | 66.4% | 61.3% |
| b | 75.8% | 67.7% | 71.0% | 65.3% | 63.4% |
| b2 | 73.2% | 65.0% | 63.8% | 58.3% | 60.6% |
| c | 77.4% | 70.5% | 72.5% | 68.7% | 64.1% |
| d | 75.2% | 68.1% | 72.4% | 60.4% | 57.6% |
| e | 69.2% | 61.9% | 64.0% | 58.7% | 56.1% |
Settings chosen in the folds (most often):
- a: in-sample {}; folds {} ×50
- a1: in-sample {"lam":10}; folds {"lam":10} ×31, {"lam":0} ×5, {"lam":0.03} ×4, {"lam":0.0001} ×3, {"lam":0.003} ×3
- b: in-sample {"kind":"ridge","lam":0.003}; folds {"kind":"ridge","lam":0.3} ×9, {"kind":"ridge","lam":0.03} ×7, {"kind":"ridge","lam":0.1} ×6, {"kind":"ridge","lam":1} ×6, {"kind":"ridge","lam":0.01} ×5
- b2: in-sample {"kind":"ridge","lam":1}; folds {"kind":"ridge","lam":3} ×20, {"kind":"ridge","lam":1} ×10, {"kind":"ridge","lam":0.3} ×8, {"kind":"ridge","lam":0.1} ×6, {"kind":"ridge","lam":0.003} ×4
- c: in-sample {"lam":0.0003}; folds {"lam":0.0003} ×22, {"lam":0.001} ×10, {"lam":0.01} ×7, {"lam":0.003} ×6, {"lam":0.03} ×3
- d: in-sample {"depth":1,"colsample":0.3,"trees":400}; folds {"depth":3,"colsample":0.3,"trees":25} ×4, {"depth":1,"colsample":1,"trees":200} ×4, {"depth":2,"colsample":0.3,"trees":50} ×4, {"depth":1,"colsample":1,"trees":400} ×3, {"depth":1,"colsample":1,"trees":50} ×3
- e: in-sample {"k":3}; folds {"k":2} ×16, {"k":12} ×15, {"k":3} ×10, {"k":8} ×5, {"k":1} ×3
Flexible against the ledger, held out (paired; SE over repeats; bootstrap over units SD and the share of resamples at or below 0):
| Pair | vs consensus | vs lists | bootstrap SD | P(≤ 0) | scrambled SD |
|---|---|---|---|---|---|
| a1 − a | -3.2 ± 0.9 | -4.0 ± 1.0 | 2.2 | 0.99 | 2.3 |
| b − a | -2.9 ± 0.8 | -5.4 ± 1.0 | 4.6 | 0.72 | 10.9 |
| b2 − a | -7.0 ± 0.9 | -9.9 ± 0.8 | 5.8 | 0.89 | 10.2 |
| c − a | -0.4 ± 0.4 | -3.4 ± 0.6 | 4.7 | 0.53 | 10.2 |
| d − a | -5.7 ± 0.5 | -7.3 ± 0.7 | 4.1 | 0.92 | 6.8 |
| e − a | -8.3 ± 1.3 | -12.1 ± 1.4 | 4.9 | 0.95 | 11.5 |
| b − a1 | 0.3 ± 1.1 | -1.5 ± 1.3 | 3.6 | 0.48 | 9.9 |
| b2 − a1 | -3.7 ± 1.5 | -5.9 ± 1.5 | 5.0 | 0.76 | 9.1 |
| c − a1 | 2.9 ± 0.9 | 0.5 ± 1.1 | 3.5 | 0.23 | 9.4 |
| d − a1 | -2.5 ± 0.8 | -3.4 ± 1.0 | 3.1 | 0.79 | 6.6 |
| e − a1 | -5.1 ± 1.5 | -8.1 ± 1.4 | 4.1 | 0.88 | 11.1 |
Permutation importance, held out, model c (drop in pairwise vs consensus, points):
| Variable | Drop | SE |
|---|---|---|
| kw_precision | 0.96 | 0.13 |
| arrives | 0.80 | 0.32 |
| kw_psychic | 0.73 | 0.14 |
| inv | 0.66 | 0.16 |
| value | 0.66 | 0.22 |
| kw_is_battleline | 0.62 | 0.24 |
| fx_arrival | 0.61 | 0.20 |
| fx_cp | 0.61 | 0.13 |
| ix_invxT | 0.60 | 0.20 |
| kw_is_fly | 0.55 | 0.18 |
| kw_is_burrowers | 0.53 | 0.14 |
| valueCheap | 0.49 | 0.17 |
| k6_chaff | 0.49 | 0.16 |
| lands | 0.48 | 0.17 |
| lr_movement | 0.48 | 0.15 |
The T-636 additions alone (top 12):
| Variable | Family | Drop | SE |
|---|---|---|---|
| modelProfiles | size | 0.42 | 0.12 |
| OC_max | size | 0.38 | 0.15 |
| scouts_inches | rules | 0.32 | 0.07 |
| anti_infantry | weaponkw | 0.28 | 0.16 |
| fx_partly | rules | 0.27 | 0.15 |
| kwu_Aircraft | keywords | 0.23 | 0.10 |
| transportCapacity | rules | 0.21 | 0.10 |
| fx_condEffects | rules | 0.18 | 0.15 |
| wk_twinLinkedAny | weaponkw | 0.18 | 0.12 |
| formShare | size | 0.17 | 0.10 |
| ppmDiscount | size | 0.17 | 0.14 |
| maxAP_any | loadout | 0.14 | 0.17 |
| Family | Drop | SE |
|---|---|---|
| flags | 4.42 | 0.64 |
| weapons | 2.96 | 0.34 |
| ledger | 2.74 | 0.54 |
| all T-636 additions | 2.07 | 0.46 |
| killclass | 0.98 | 0.23 |
| rules | 0.96 | 0.32 |
| datasheet | 0.71 | 0.30 |
| keywords | 0.41 | 0.16 |
| size | 0.39 | 0.27 |
| interaction | 0.14 | 0.23 |
| weaponkw | 0.02 | 0.20 |
| loadout | -0.48 | 0.33 |
Permutation importance, held out, model d (drop in pairwise vs consensus, points):
| Variable | Drop | SE |
|---|---|---|
| c_score | 4.12 | 0.75 |
| value | 2.60 | 0.49 |
| ix_arrivexMelee | 1.01 | 0.34 |
| k6_monster | 0.67 | 0.30 |
| ppmDiscount | 0.66 | 0.21 |
| inv | 0.59 | 0.16 |
| k6_light_vehicle | 0.53 | 0.25 |
| valueCheap | 0.49 | 0.24 |
| kw_heavy | 0.42 | 0.27 |
| lands | 0.36 | 0.24 |
| kw_precision | 0.35 | 0.17 |
| c_kill | 0.23 | 0.11 |
| ptsAtMax | 0.22 | 0.23 |
| c_hold | 0.20 | 0.11 |
| OC | 0.17 | 0.10 |
The T-636 additions alone (top 12):
| Variable | Family | Drop | SE |
|---|---|---|---|
| ppmDiscount | size | 0.66 | 0.21 |
| ptsAtMax | size | 0.22 | 0.23 |
| abilitiesOwn | rules | 0.17 | 0.08 |
| shortRanged | loadout | 0.10 | 0.08 |
| dsMelee | loadout | 0.07 | 0.08 |
| rulesCount | rules | 0.06 | 0.10 |
| sizeRatio | size | 0.06 | 0.03 |
| W_max | size | 0.04 | 0.02 |
| kwu_Transport | keywords | 0.03 | 0.03 |
| fx_partly | rules | 0.03 | 0.07 |
| defaultWeapons | loadout | 0.01 | 0.01 |
| maxS_any | loadout | 0.01 | 0.01 |
| Family | Drop | SE |
|---|---|---|
| ledger | 10.39 | 1.00 |
| size | 0.97 | 0.35 |
| killclass | 0.63 | 0.45 |
| interaction | 0.54 | 0.45 |
| datasheet | 0.45 | 0.27 |
| keywords | 0.00 | 0.00 |
| weaponkw | -0.01 | 0.01 |
| weapons | -0.06 | 0.46 |
| rules | -0.16 | 0.21 |
| loadout | -0.19 | 0.28 |
| flags | -0.25 | 0.19 |
| all T-636 additions | -0.34 | 0.33 |
Permutation importance, held out, model a1 (drop in pairwise vs consensus, points):
| Variable | Drop | SE |
|---|---|---|
| c_kill | 17.66 | 1.20 |
| c_score | 1.40 | 0.30 |
| c_presence | 1.06 | 0.43 |
| c_hold | 0.98 | 0.24 |
| c_soak | 0.85 | 0.24 |
| c_spawn | 0.18 | 0.06 |
| c_actions | 0.10 | 0.04 |
| M | 0.00 | 0.00 |
| T | 0.00 | 0.00 |
| Sv | 0.00 | 0.00 |
| inv | 0.00 | 0.00 |
| W | 0.00 | 0.00 |
| Ld | 0.00 | 0.00 |
| OC | 0.00 | 0.00 |
| models | 0.00 | 0.00 |
The T-636 additions alone (top 12):
| Variable | Family | Drop | SE |
|---|---|---|---|
| kwu_Vehicle | keywords | 0.00 | 0.00 |
| kwu_Walker | keywords | 0.00 | 0.00 |
| kwu_Mounted | keywords | 0.00 | 0.00 |
| kwu_Aircraft | keywords | 0.00 | 0.00 |
| kwu_Transport | keywords | 0.00 | 0.00 |
| kwu_Dedicated_Transport | keywords | 0.00 | 0.00 |
| kwu_Leader | keywords | 0.00 | 0.00 |
| kwu_Grenades | keywords | 0.00 | 0.00 |
| kwu_Smoke | keywords | 0.00 | 0.00 |
| kwu_Frame | keywords | 0.00 | 0.00 |
| kwu_Tacticus | keywords | 0.00 | 0.00 |
| kwu_Daemon | keywords | 0.00 | 0.00 |
| Family | Drop | SE |
|---|---|---|
| ledger | 19.04 | 1.03 |
| datasheet | 0.00 | 0.00 |
| weapons | 0.00 | 0.00 |
| flags | 0.00 | 0.00 |
| killclass | 0.00 | 0.00 |
| interaction | 0.00 | 0.00 |
| keywords | 0.00 | 0.00 |
| size | 0.00 | 0.00 |
| loadout | 0.00 | 0.00 |
| weaponkw | 0.00 | 0.00 |
| rules | 0.00 | 0.00 |
| all T-636 additions | 0.00 | 0.00 |
Partial dependence, model c (fit on all 52; consensus scale; biggest step):
- kw_precision: range 0.19 letters; biggest step +0.19 between 0 and 1; curve 0→2.91, 1→3.10
- arrives: range 0.14 letters; biggest step +0.14 between 0 and 1; curve 0→2.84, 1→2.97
- kw_psychic: range 0.20 letters; biggest step +0.10 between 0 and 1; curve 0→2.92, 1→3.02, 2→3.08, 3→3.12
- inv: range 0.10 letters; biggest step -0.03 between 4 and 5; curve 4→3.01, 5→2.98, 6→2.94, 7→2.91
- value: range 0.21 letters; biggest step +0.06 between 2133.618 and 3076.76; curve 10.796→2.86, 479.387→2.88, 644.868→2.89, 731.117→2.90, 960.284→2.91, 1019.608→2.92, 1135.23→2.92, 1481.212→2.94, 2133.618→2.98, 3076.76→3.04, 3422.068→3.06
- kw_is_battleline: range 0.27 letters; biggest step +0.27 between 0 and 1; curve 0→2.91, 1→3.18
- fx_arrival: range 0.14 letters; biggest step +0.14 between 0 and 1; curve 0→2.92, 1→3.06
- fx_cp: range 0.09 letters; biggest step +0.09 between 0 and 1; curve 0→2.92, 1→3.01
- ix_invxT: range 0.10 letters; biggest step +0.05 between 0 and 5; curve 0→2.91, 5→2.96, 8→2.97, 15→2.99, 18→2.99, 24→3.00, 27→3.00, 28→3.00, 30→3.01, 33→3.01
- kw_is_fly: range 0.13 letters; biggest step -0.13 between 0 and 1; curve 0→2.97, 1→2.84
- kw_is_burrowers: range 0.19 letters; biggest step +0.19 between 0 and 1; curve 0→2.92, 1→3.11
- valueCheap: range 0.20 letters; biggest step +0.05 between 2133.618 and 3017.884; curve 10.796→2.86, 441.848→2.89, 617.075→2.90, 676.313→2.90, 876.142→2.91, 1006.336→2.92, 1135.23→2.93, 1481.212→2.95, 2133.618→2.98, 3017.884→3.03, 3422.068→3.06
Partial dependence, model d (fit on all 52; consensus scale; biggest step):
- c_score: range 0.90 letters; biggest step +0.46 between 9.455 and 11.587; curve 0→2.27, 9.455→2.50, 11.587→2.96, 12.606→2.96, 13.904→3.17, 17.176→3.17, 22.034→3.17, 26.182→3.17, 33.279→3.17, 85.881→3.17, 209.932→3.17
- value: range 0.67 letters; biggest step +0.42 between 3076.76 and 3422.068; curve 10.796→2.75, 479.387→2.75, 644.868→2.75, 731.117→2.91, 960.284→2.91, 1019.608→2.95, 1135.23→2.95, 1481.212→2.95, 2133.618→2.95, 3076.76→3.00, 3422.068→3.42
- ix_arrivexMelee: range 0.62 letters; biggest step +0.45 between 0 and 16.4; curve 0→2.61, 16.4→3.06, 21.033→3.06, 30.979→3.06, 62.253→3.06, 82.392→3.06, 92.285→3.06, 100.223→3.06, 130.103→3.14, 134.006→3.23, 170.194→3.23
- k6_monster: range 0.23 letters; biggest step +0.14 between 8.905 and 12.058; curve 0→2.84, 0.039→2.84, 0.187→2.84, 0.534→2.93, 1.829→2.93, 2.475→2.93, 3.56→2.93, 4.62→2.93, 8.905→2.93, 12.058→3.07, 30.133→3.07
- ppmDiscount: range 0.63 letters; biggest step -0.30 between 0.969 and 1; curve 0.556→3.49, 0.778→3.49, 0.833→3.49, 0.857→3.49, 0.917→3.20, 0.933→3.20, 0.955→3.16, 0.969→3.16, 1→2.86, 1.056→2.86, 1.063→2.86
- inv: range 0.18 letters; biggest step -0.18 between 4 and 5; curve 4→3.08, 5→2.90, 6→2.90, 7→2.90
- k6_light_vehicle: range 0.24 letters; biggest step +0.24 between 9.287 and 12.061; curve 0→2.83, 0.96→2.83, 2.045→2.83, 4.396→2.83, 7.602→2.83, 9.287→2.83, 12.061→3.07, 14.028→3.07, 16.005→3.07, 17.922→3.07, 22.099→3.07
- valueCheap: range 0.57 letters; biggest step +0.39 between 3017.884 and 3422.068; curve 10.796→2.83, 441.848→2.83, 617.075→2.83, 676.313→2.83, 876.142→2.83, 1006.336→2.83, 1135.23→3.00, 1481.212→3.00, 2133.618→3.00, 3017.884→3.00, 3422.068→3.40
- kw_heavy: range 0.70 letters; biggest step +0.70 between 0 and 1; curve 0→2.87, 1→3.57
- lands: range 0.25 letters; biggest step +0.15 between 0.924 and 0.988; curve 0→2.79, 0.66→2.79, 0.742→2.79, 0.786→2.88, 0.924→2.88, 0.988→3.04, 1→3.04
- kw_precision: range 0.46 letters; biggest step +0.46 between 0 and 1; curve 0→2.88, 1→3.34
- c_kill: range 0.00 letters; biggest step +0.00 between 0 and 370.205; curve 0→2.93, 370.205→2.93, 508.789→2.93, 594.518→2.93, 707.168→2.93, 924.297→2.93, 958.704→2.93, 1215.177→2.93, 1688.716→2.93, 2746.972→2.93, 3293.715→2.93
The trees' splits (fit on all 52): variable, gain, median threshold, uses:
- c_score: gain 1105.7, threshold ~10.292 (6.201 to 13.613), 36 splits
- valueCheap: gain 964.8, threshold ~3047.439 (1096.886 to 3110.314), 15 splits
- value: gain 890.3, threshold ~3076.877 (685.486 to 3110.314), 19 splits
- ix_arrivexMelee: gain 375.4, threshold ~11.492 (5.25 to 131.936), 32 splits
- lands: gain 331.8, threshold ~0.956 (0.764 to 0.956), 11 splits
- ppmDiscount: gain 310.6, threshold ~0.887 (0.887 to 0.984), 23 splits
- OC: gain 280.4, threshold ~0.5 (0.5 to 0.5), 1 splits
- kw_heavy: gain 267.9, threshold ~0.5 (0.5 to 0.5), 27 splits
- k6_light_vehicle: gain 222.0, threshold ~11.083 (11.083 to 11.083), 13 splits
- kw_precision: gain 204.3, threshold ~0.5 (0.5 to 0.5), 17 splits
- OC_max: gain 152.3, threshold ~0.5 (0.5 to 0.5), 1 splits
- inv: gain 126.8, threshold ~4.5 (4.5 to 4.5), 8 splits
The trees' interactions (a depth-3 fit on all 52, 100 trees: one split under the other, by the child's gain):
- c_score × value: 741.9
- k6_monster × value: 210.4
- arrives × value: 168.9
- inv × k6_monster: 106.4
- c_score × inv: 89.6
- rangedVeh100 × value: 81.8
- lands × value: 79.0
- kw_precision × ppmDiscount: 77.2
- c_score × k6_cavalry_and_beasts: 74.3
- k6_light_vehicle × ppmDiscount: 71.7
- ix_durHeavy100 × value: 70.5
- ix_arrivexMelee × k6_monster: 67.7
Linear (b) fitted on all 52, largest standardised weights:
- kw_is_battleline: 0.535
- kw_indirect: 0.510
- fx_list: -0.442
- arrives: 0.391
- kw_precision: 0.387
- modelProfiles: -0.385
- kw_heavy: 0.356
- fx_arrival: 0.351
- fx_partly: 0.336
- ppmDiscount: -0.336
- kw_is_character: -0.325
- kw_psychic: 0.320
- lands: 0.314
- k6_light_vehicle: 0.301
- kw_is_fly: -0.297
- lr_movement: -0.293
- kw_anti: -0.290
- anti_infantry: -0.275
- maxAP_any: 0.270
- kw_is_transport: -0.265
The ledger's seven weights re-fitted on all 52 (a1):
- kill: step 6 498.17 → 497.82 (×1.00)
- soak: step 6 14.97 → 14.97 (×1.00)
- score: step 6 9.65 → 9.65 (×1.00)
- actions: step 6 1.31 → 1.31 (×1.00)
- hold: step 6 18.89 → 18.89 (×1.00)
- spawn: step 6 15.65 → 15.65 (×1.00)
- presence: step 6 2.52 → 2.52 (×1.00)
Units the best model (c) gets right that the ledger gets wrong (held-out share of the unit's pairs ordered right):
| Unit | Consensus | Consensus rank | Ledger rank | Best model | Ledger | Gain | Top variables (percentile) |
|---|---|---|---|---|---|---|---|
| Toxicrene | 1 | 48 | 23 | 86.2% | 50.7% | 36 | kw_precision 0, arrives 27, kw_psychic 0, inv 25, value 57 |
| Harpy | 0.667 | 51.5 | 33 | 93.6% | 61.3% | 32 | kw_precision 0, arrives 0, kw_psychic 0, inv 25, value 37 |
| Gargoyles | 3.75 | 15 | 47 | 65.5% | 35.2% | 30 | kw_precision 0, arrives 27, kw_psychic 0, inv 25, value 10 |
| Tyranid Warriors with Ranged Bio-Weapons | 2 | 40.5 | 13 | 69.7% | 48.9% | 21 | kw_precision 0, arrives 27, kw_psychic 0, inv 25, value 76 |
| Tyranid Prime with Lash Whip | 2.5 | 35 | 14 | 69.9% | 51.3% | 19 | kw_precision 0, arrives 0, kw_psychic 0, inv 25, value 75 |
| Genestealers | 3.75 | 15 | 34 | 71.2% | 53.6% | 18 | kw_precision 0, arrives 27, kw_psychic 0, inv 20, value 35 |
| Norn Emissary | 4.2 | 11 | 19 | 75.6% | 61.2% | 14 | kw_precision 90, arrives 27, kw_psychic 100, inv 0, value 65 |
| Tyrant Guard | 2.8 | 29.5 | 49 | 74.8% | 60.4% | 14 | kw_precision 0, arrives 27, kw_psychic 0, inv 25, value 6 |
| Neurotyrant | 3.667 | 17 | 39 | 61.2% | 47.2% | 14 | kw_precision 0, arrives 27, kw_psychic 92, inv 0, value 25 |
| Carnifexes | 2 | 40.5 | 30 | 86.2% | 74.7% | 12 | kw_precision 0, arrives 0, kw_psychic 0, inv 25, value 43 |
And the reverse (the ledger right, the model wrong):
- Biovores: consensus 5 (rank 1), ledger rank 5; model 7.2% against ledger 92.6%
- Tyrannofex: consensus 4.4 (rank 6), ledger rank 25; model 42.2% against ledger 68.9%
- Hyperadapted Raveners: consensus 4.2 (rank 11), ledger rank 9; model 63.9% against ledger 81.5%
- Lictor: consensus 4.75 (rank 2), ledger rank 2; model 79.3% against ledger 96.7%
- Psychophage: consensus 3 (rank 27.5), ledger rank 27; model 66.2% against ledger 83.1%
- Parasite of Mortrex: consensus 2.2 (rank 36.5), ledger rank 43; model 59.4% against ledger 75.2%
- Von Ryan's Leapers: consensus 3.4 (rank 20.5), ledger rank 6; model 60.6% against ledger 71.6%
- The Red Terror: consensus 4.333 (rank 8), ledger rank 12; model 70.0% against ledger 80.9%
Every flexible model misses the same way (mean held-out prediction minus consensus, letters):
- Biovores: -3.23 (consensus 5; S/S/S/S/S)
- Toxicrene: +1.901 (consensus 1; –/–/–/D/D)
- Tyrannofex: -1.614 (consensus 4.4; S/A/A/A/S)
- Tervigon: +1.61 (consensus 2; B/C/C/C/D)
- Hierophant: +1.591 (consensus 1.333; D/C/–/D/–)
- Lictor: -1.435 (consensus 4.75; –/A/S/S/S)
- Exocrine: -1.394 (consensus 4.4; A/A/A/S/S)
- Tyranid Warriors with Ranged Bio-Weapons: +1.328 (consensus 2; C/C/B/D/C)
- Neurolictor: -1.313 (consensus 4.5; S/S/B/–/S)
- Deathleaper: +1.157 (consensus 3.2; A/S/B/B/D)
- Neurogaunts: +1.132 (consensus 2.2; B/C/C/C/C)
- Hormagaunts: -1.105 (consensus 4.2; S/A/B/A/S)