6 Oct 2026, Pacific. tools/statlearn.js; results in build/statlearn/ (git-ignored: features.json and .csv, results.json). Snapshot 8dcd24032d, inputs hash dcd477d8a78c. Base: Track A (tools/ledger-benchmarks/2026-10-05-step34b-baseline.json), wCarry 0. A report: nothing in the ledger moves, no slider, setting or letter changes, and nothing is adopted.
In plain English
The question. For days we have tuned dials on one model: Value per 100 points, about 85–91% of it Kill. It sits near 62% agreement with the reviewers across 16 armies. Jordan asked for structurally different models instead, tested cheaply and honestly. So I built one table of 75 facts about each of 733 units (595 graded), and fitted 34 models of different kinds on it, plus every fact alone.
How it was judged.
- Every model was fitted on 15 armies and scored on the 16th, for each army in turn. No army's letters ever tune its own score.
- The score is the share of the held-out army's reviewer pairs put in the right order (cascade2's consensus pairs; ties count wrong).
- Every model was also refitted on 20 sets of shuffled letters. All 34 beat at least 19 of the 20 shuffles.
- Each is compared with the ledger by a bootstrap over armies.
- The ledger was tuned on Tyranids, so its Tyranid score (89.5) is in sample. The fairest column is 15 armies, no Tyranids, where the ledger scores 60.0.
The leaderboard (the main ones; the full table is below).
| Model | 16 armies | 15, no Tyranids | vs ledger on 15 [95%] | On Gemini's letters, vs ledger [95%] |
|---|---|---|---|---|
| List share (popularity; diagnostic only, never a letter) | 69.6 | 67.9 | +7.8 [−1.0, +15.9] | +10.4 [+1.2, +19.2] |
| Ensemble: ledger Value rank + datasheet model rank (declared before the run) | 65.5 | 63.9 | +3.9 [+0.4, +7.3] | +6.2 [+3.0, +9.8] |
| Pairwise logistic on each fact's rank within its army | 65.2 | 64.0 | +4.0 [−0.9, +9.2] | +6.5 [+2.0, +11.3] |
| Datasheet, keywords and support only, no ledger at all | 64.0 | 63.2 | +3.2 [−3.2, +9.6] | +7.4 [+1.3, +13.3] |
| Kill line alone | 62.6 | – | +0.8 [−1.2, +3.1] (16) | – |
| The ledger (Track A) | 61.9 | 60.0 | – | (56.7 on Gemini) |
| Every fact, pairwise logistic (L2 / L1) | 61.4 / 61.3 | 60.4 / 60.1 | +0.4 / +0.1 | +4.3 / +5.3 (CI touches 0) |
| Nearest neighbours across armies | 61.6 | 60.9 | +0.9 [−4.4, +6.2] | – |
| Ordinal logistic on the letters | 60.7 | 59.4 | −0.6 | – |
| Shallow boosted trees | 60.4 | 59.6 | −0.4 | +3.7 (CI touches 0) |
| Leaders at their best joined form | 60.1 | 58.2 | −1.8 [−4.9, +1.7] | – |
| The seven lines refitted / design/jobs.md's jobs only | 59.6 / 59.3 | 58.4 / 58.9 | −1.7 / −1.2 | – |
| Value within role (Jordan's eight roles, by rules) | 56.9 | 55.5 | −4.5 [−9.2, −0.4] | – |
| Value over replacement (less the best same-role unit) | 58.9 | 57.3 | −2.7 [−7.1, +1.4] | – |
| Replaceability (how many same-role units beat it) | 53.9 | 52.2 | −7.8 [−12.3, −4.1] | – |
What this says.
- Re-shaping the ledger's own information gets nowhere. Logistic (L1 and L2), ordinal, boosted trees, nearest neighbours, value over replacement and leaders at their joined form all land within about 2 points of the ledger. So does Value with every fact added. The ledger is about as good as its inputs allow. The Kill line alone does as well as the whole ledger (62.6).
- The gain comes from information the ledger does not use. A model that never sees the ledger beats it by about 3 points on 15 armies. It sees only the datasheet, keywords and support rules. Averaging its rank with the ledger's is the one model whose interval clears 0 on the human letters: +3.9 [+0.4, +7.3] on 15 armies, better in 9 of 16. It was declared before the run.
- A second, independent target agrees. The models were trained on human letters only. On Gemini's letters for the held-out army, the datasheet model beats the ledger by +7.4 [+1.3, +13.3] and the ensemble by +6.2 [+3.0, +9.8]. Caveat: Gemini says it draws on tournament results, and list share does just as well there. So some of that agreement may be popularity.
- Ranks beat raw values. The same every-fact model on each fact's rank within its army scores 65.2 against 61.4. Our facts are heavy-tailed, and a few outliers dominate a raw fit. Ranks are robust to that.
- Within-role scoring is worse, not better. Within role, within archetype, value over replacement and replaceability all lose 3 to 14 points. Reviewers rank across roles. A unit that is the best of a weak role is not graded up for it. (Our roles come from simple rules, not T-732's measured roles. A better role assignment might do better.)
- Multiplicity. I tried 34 models. One interval clearing 0 on the human letters is suggestive, not proof. The Gemini agreement is the stronger evidence, because Gemini's letters played no part in any fit. Nothing should be adopted before a confirmation run on lists none of these fits saw.
What reviewers seem to value. These signs hold in all 16 leave-one-out fits, across the datasheet model, the rank model, the ordinal model and the L1 model. Each is "one army-SD more of this, within an army".
- Up: Move (the largest and steadiest); Epic Hero; Psyker; Lone Operative; Infantry (hideable in 11e terrain); Soak (durability); Toughness and Wounds; Deep Strike; a command-point rule; an army-rule enabler; Kill into tanks, elites and hordes; points (dearer units are graded higher within an army, all else equal).
- Down: Fly (aircraft, skimmers and jump packs, all else equal); a debuff rule; Transport; Vehicle; long range; and a pile of support-rule kinds once the rest is known.
- Lines. A lines-only refit gives Soak and Kill about equal weight (31% and 28% of its weight), then Score 15% and Hold 11%. The ledger gives Kill 85% and Soak 3%. The scales differ (the fit is per army-SD of log(1 + line)), so read the order, not the sizes. Still, durability is badly under-weighted against killing.
- Cheapness does not help on its own. OC per point, OC kept, holding and screening alone are at chance (49–50, ties as half). Cheap activations are slightly negative.
- Gemini as training data. Human + 0.5 × Gemini pairs moves held-out human accuracy by +0.7 (every-fact logistic), −0.3 (lines) and +1.2 (trees). Small, as T-731's half weight suggests.
Prediction (written before any fit), checked.
- Wrong that nothing clears the ledger by 2 with an interval clear of 0: the ensemble does, by +3.6 to +3.9.
- Right that durability gets far more weight than the ledger gives it. Wrong on cost: dearer units are graded higher, and cheapness and OC per point carry nothing.
- Wrong that ranks would do about as well as raw values: ranks are better by 3.8.
- Wrong on within-role scoring: it is worse by 4 to 5, not even.
- Wrong on leaders at their joined form: −1.8, not +0.5 to +1.5.
- Right that trees do no better than linear models.
- Right that real models beat scrambled ones: 7 to 16 points above the scrambled mean.
- Right that Gemini training moves things little (+0.7 for the every-fact logistic; +1.2 for trees, just over the bound).
Recommendation
- Stop tuning dials on the present lines. Every reshaping of the ledger's own information, linear, ordinal, trees, neighbours, roles, lands within about 2 points of it.
- The missing signal is on the datasheet. Speed, named characters and psykers, Lone Operative, Infantry (Hidden), durability. A ledger-free datasheet model matches or beats the ledger.
- Run one confirmation of the declared ensemble (ledger rank + datasheet-model rank) on human lists none of these fits saw, such as new reviewer lists as they arrive. Adopt nothing before that.
- In the ledger itself, Soak deserves weight comparable to Kill. Speed and character or psyker status are mechanisms worth pricing as lines.
- Drop within-role, value-over-replacement and replaceability as letter models. Use ranks rather than raw values in any pooled model.
Method notes
- Features (75; build/statlearn/features.csv lists every value): each line's contribution per 100 points; Kill by target class; the datasheet (points, models, T, Sv, invulnerable save, W, wounds per 100 points, Move, OC, OC per 100 points, Ld, Feel No Pain, range); keywords (Character, Epic Hero, Vehicle, Monster, Infantry, Fly, Transport, Battleline, Psyker, Synapse, Deep Strike, Infiltrators, Scouts, Lone Operative, Stealth, Fights First, Leader); leading (the best joined form as leader and as squad, readable.js's package and lift); support (aura, CP, Battle-shock and debuff rules, how many kinds, enabler, transport capacity); and design/jobs.md's 22 job measures. The datasheet-only model uses the datasheet, keywords, support and Leader columns: no line, Kill class or job.
- Training pairs: each army's consensus pairs, weighted to total its graded units (halved for a single-list army). Features are standardised within each army. λ is chosen by an inner leave-army-group-out (4 groups of the training armies). Trees: depth 2, 100 rounds, learning rate 0.1, at least 20 units a leaf; fixed in advance, not tuned.
- Roles (Jordan's eight) come from simple rules in priority order: transport; character leader or support with low Burst; debuff, shock or character-hunter; cheap low-Kill bodies; action-doer at 110 points or less; anti-tank shooter; durable holder; else Hammer. Counts per army are at the end of the tables.
- Post hoc: three models were added after the first run had shown the datasheet and rank models on top. They are labelled POST HOC and are not counted as evidence: datasheet facts on ranks (64.0), datasheet + log Value (63.5), and the ledger plus the every-fact rank model (63.5).
- Ties: single flags tie on most pairs, so their strict score sits far below 50. The single-feature table sorts on ties as half a pair instead.
- The Actions line's coefficient in the lines-only fit comes out at about 0 (0.0001), so it rounds out of that table.
- Reproduce:
node --max-old-space-size=3000 tools/statlearn.js --features(about 10 s), thennode tools/statlearn.js --fit --scrambles 20 --boot 2000 --md <file>(about 14 min on one thread);--only <prefixes>refits some models into the same results file;--md-only <file>rewrites the tables.
The tables (tools/statlearn.js --md)
The leaderboard (held out, leave one army out)
Pooled = mean over the 16 armies of the held-out accuracy on each army's human consensus pairs (cascade2's definition); every pair = all armies' pairs together; lists = each human 11e list's own pairs (army mean). Scrambled: the identical pipeline on 20 letter sets permuted within each army. vs ledger: pooled difference from Track A's Value with a 95% bootstrap interval over armies (2000 draws); armies = how many of 16 it beats the ledger in. 15 (no Tyranids): the same without the army the ledger was tuned on (its Tyranid score is in sample), with its own bootstrap over the 15. Gemini = held-out accuracy on the held-out army's GEMINI AGGREGATE pairs (15 armies; Emperor's Children has none). Family: a ledger, c pairwise logistic, d ordinal, e within role, f value over replacement, g ranks, h boosting, i replaceability, j leaders joined, k diagnostic, plus/knn/ens other structural tries, gem trained with Gemini.
| Model | Family | Pooled, 16 | 15 (no Tyranids) | vs ledger on 15 [95%] | Every pair | Lists | Scrambled | Real − scrambled | Beats scrambled | vs ledger [95%] | Armies better | Gemini |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| List share alone (DIAGNOSTIC ONLY: popularity, never in a letter) | k | 69.6 | 67.9 | +7.8 [-1.0, 15.9] | 67.0 | 67.0 | 43.2 ± 2.2 | +26.4 | 20 of 20 | +7.7 [0.1, 15.7] | 10 | 67.1 |
| Ensemble: ledger Value rank + datasheet-only model rank | ens | 65.5 | 63.9 | +3.9 [0.4, 7.3] | 62.4 | 64.1 | 49.5 ± 2.9 | +15.9 | 20 of 20 | +3.6 [0.4, 7.1] | 9 | 62.9 |
| Pairwise logistic on within-army ranks, every feature, L2 | g | 65.2 | 64.0 | +4.0 [-0.9, 9.2] | 65.2 | 63.6 | 50.6 ± 2.7 | +14.6 | 20 of 20 | +3.4 [-1.3, 8.7] | 9 | 63.2 |
| POST HOC: datasheet-only model on within-army ranks, L2 | posthoc | 64.0 | 63.1 | +3.1 [-3.6, 9.9] | 58.9 | 62.6 | 50.7 ± 3.2 | +13.2 | 20 of 20 | +2.1 [-4.4, 9.2] | 8 | 66.3 |
| Pairwise logistic, datasheet + keywords + support only (no ledger), L2 | c | 64.0 | 63.2 | +3.2 [-3.2, 9.6] | 59.2 | 62.6 | 50.6 ± 2.6 | +13.3 | 20 of 20 | +2.1 [-4.1, 8.7] | 7 | 64.0 |
| POST HOC: datasheet-only features + log Value, L2 | posthoc | 63.5 | 62.6 | +2.6 [-2.2, 6.9] | 59.7 | 62.5 | 50.7 ± 2.9 | +12.9 | 20 of 20 | +1.7 [-2.9, 6.4] | 7 | 62.2 |
| POST HOC: ensemble, ledger Value rank + every-feature ranks model rank | ens | 63.5 | 61.8 | +1.8 [-0.9, 4.5] | 64.1 | 62.0 | 49.6 ± 3.2 | +14.0 | 20 of 20 | +1.7 [-0.8, 4.3] | 11 | 59.6 |
| c every feature L2, trained on human + 0.5 × Gemini | gem | 62.1 | 60.9 | +0.9 [-4.0, 5.6] | 60.4 | 60.7 | – | – | – | +0.3 [-4.3, 5.1] | 7 | 61.9 |
| Ledger Value (Track A) | a | 61.9 | 60.0 | +0.0 [0.0, 0.0] | 61.7 | 60.7 | 50.2 ± 2.3 | +11.7 | 20 of 20 | +0.0 [0.0, 0.0] | 0 | 56.7 |
| Pairwise logistic, every feature + log Value, L2 | c | 61.8 | 60.8 | +0.8 [-3.8, 5.2] | 60.6 | 60.5 | 50.7 ± 2.9 | +11.2 | 20 of 20 | -0.0 [-4.3, 4.6] | 8 | 60.8 |
| 25 nearest neighbours across armies (24 features) | knn | 61.6 | 60.9 | +0.9 [-4.4, 6.2] | 60.4 | 60.3 | 50.3 ± 2.7 | +11.3 | 20 of 20 | -0.2 [-5.7, 5.2] | 9 | 58.7 |
| h boosting, trained on human + 0.5 × Gemini | h | 61.6 | 60.5 | +0.4 [-4.4, 4.9] | 62.7 | 60.4 | – | – | – | -0.2 [-4.8, 4.3] | 7 | 59.0 |
| Absolute output: Value × points | plus | 61.5 | 60.4 | +0.3 [-3.8, 4.2] | 62.5 | 60.4 | 50.1 ± 2.7 | +11.4 | 20 of 20 | -0.4 [-4.4, 3.7] | 8 | 55.6 |
| Pairwise logistic, every feature, L2 | c | 61.4 | 60.4 | +0.4 [-4.4, 4.9] | 60.3 | 60.2 | 50.8 ± 3.1 | +10.6 | 20 of 20 | -0.4 [-4.8, 4.3] | 8 | 61.0 |
| Pairwise logistic, every feature, L1 | c | 61.3 | 60.1 | +0.1 [-4.0, 4.3] | 60.8 | 60.0 | 50.9 ± 3.0 | +10.4 | 20 of 20 | -0.6 [-4.6, 3.7] | 7 | 62.0 |
| Ensemble: ledger Value rank + boosting rank | ens | 61.2 | 59.7 | -0.3 [-2.5, 1.9] | 62.3 | 60.0 | 49.4 ± 2.8 | +11.8 | 20 of 20 | -0.6 [-2.9, 1.7] | 7 | 57.5 |
| Value + value over replacement (pairwise fit) | f | 61.2 | 59.3 | -0.7 [-1.9, 0.0] | 61.4 | 60.0 | 49.8 ± 3.2 | +11.4 | 20 of 20 | -0.7 [-1.8, 0.0] | 5 | 57.4 |
| Every unit at max(solo, best joined as leader or squad) | j | 61.2 | 59.3 | -0.8 [-4.8, 3.3] | 58.9 | 60.1 | 50.8 ± 2.1 | +10.4 | 20 of 20 | -0.7 [-4.5, 3.4] | 7 | 56.8 |
| Ordinal logistic on letters, every feature, L2 | d | 60.7 | 59.4 | -0.6 [-5.3, 4.3] | 59.5 | 59.5 | 50.8 ± 3.2 | +9.9 | 20 of 20 | -1.2 [-5.7, 4.0] | 5 | 58.9 |
| Shallow pairwise boosting (depth 2, 100 rounds) | h | 60.4 | 59.6 | -0.4 [-4.8, 4.2] | 61.9 | 59.3 | 50.2 ± 2.4 | +10.2 | 20 of 20 | -1.4 [-6.0, 3.4] | 7 | 60.4 |
| Leaders at max(solo, best joined) | j | 60.1 | 58.2 | -1.8 [-5.0, 1.7] | 58.6 | 59.3 | 50.7 ± 2.2 | +9.4 | 20 of 20 | -1.7 [-4.8, 1.8] | 4 | 56.5 |
| Leaders at their best joined form's Value | j | 60.1 | 58.2 | -1.8 [-4.9, 1.7] | 58.8 | 59.2 | 50.8 ± 2.1 | +9.3 | 20 of 20 | -1.8 [-4.7, 1.6] | 5 | 56.9 |
| Pairwise logistic, the seven lines only, L2 | c | 59.6 | 58.4 | -1.7 [-6.0, 2.8] | 59.8 | 58.1 | 49.8 ± 3.0 | +9.7 | 20 of 20 | -2.3 [-6.5, 2.2] | 6 | 56.5 |
| Pairwise logistic, design/jobs.md's jobs only, L2 | c | 59.3 | 58.9 | -1.2 [-4.6, 1.8] | 61.0 | 58.2 | 50.7 ± 2.6 | +8.6 | 20 of 20 | -2.6 [-6.7, 1.1] | 6 | 57.0 |
| c lines L2, trained on human + 0.5 × Gemini | gem | 59.3 | 58.2 | -1.8 [-6.4, 2.9] | 59.9 | 57.7 | – | – | – | -2.6 [-7.1, 2.3] | 5 | 56.0 |
| Role intercepts + Value within role (pairwise fit) | e | 59.2 | 57.6 | -2.4 [-6.0, 0.4] | 60.3 | 58.4 | 49.6 ± 3.6 | +9.6 | 20 of 20 | -2.7 [-6.0, 0.1] | 5 | 54.9 |
| Value over replacement: less the best other same-role unit | f | 58.9 | 57.3 | -2.7 [-7.1, 1.4] | 60.8 | 58.1 | 50.0 ± 3.0 | +8.9 | 20 of 20 | -3.0 [-7.2, 0.9] | 3 | 53.8 |
| Value within archetype (T-716) | e | 57.8 | 56.3 | -3.7 [-7.2, 0.0] | 53.9 | 57.7 | 50.3 ± 2.6 | +7.4 | 20 of 20 | -4.1 [-7.5, -0.5] | 5 | 56.4 |
| Value over replacement: less the role's median | f | 57.5 | 56.1 | -4.0 [-8.4, -0.2] | 60.1 | 56.9 | 50.0 ± 2.7 | +7.5 | 20 of 20 | -4.3 [-8.2, -0.9] | 5 | 53.6 |
| Archetype intercepts + Value within archetype (pairwise fit) | e | 57.5 | 57.2 | -2.9 [-6.3, 0.6] | 57.1 | 57.0 | 50.4 ± 2.5 | +7.1 | 20 of 20 | -4.4 [-8.7, -0.3] | 4 | 57.9 |
| Value within role (Jordan's eight, by rules) | e | 56.9 | 55.5 | -4.5 [-9.2, -0.4] | 59.6 | 56.3 | 50.2 ± 2.9 | +6.7 | 19 of 20 | -4.9 [-9.1, -1.0] | 5 | 54.3 |
| Points alone (sign chosen) | plus | 54.4 | 54.9 | -5.1 [-12.3, 1.4] | 55.7 | 53.8 | 49.5 ± 3.1 | +4.9 | 20 of 20 | -7.4 [-15.3, 0.2] | 4 | 49.8 |
| Replaceability: fewer same-role units with a higher Value | i | 53.9 | 52.2 | -7.8 [-12.3, -4.1] | 54.7 | 53.0 | 46.5 ± 2.7 | +7.4 | 20 of 20 | -8.0 [-11.9, -4.3] | 2 | 52.3 |
| Replaceability: fewer same-army units at least as good at its main job | i | 47.7 | 47.5 | -12.5 [-16.2, -8.4] | 43.5 | 47.6 | 43.4 ± 2.2 | +4.3 | 19 of 20 | -14.2 [-18.8, -9.5] | 2 | 45.2 |
Each feature alone (family b; its sign chosen on the training armies)
Sign: + higher is better, as most training folds chose it (folds agreeing of 16). All 75 features listed, sorted by the score with ties as half a pair (flags and whole-number stats tie on most pairs, so the strict score, ties wrong, sits far below 50 for them).
| Feature | Sign (folds) | Ties as half | Scrambled, ties as half | Real − scrambled | Strict (ties wrong) | vs ledger, strict [95%] | Armies better | Gemini |
|---|---|---|---|---|---|---|---|---|
| L_kill | + (16) | 62.6 | 50.8 | +11.9 | 62.6 | +0.8 [-1.2, 3.1] | 10 | 57.4 |
| j_antiElite | + (16) | 60.7 | 48.9 | +11.7 | 60.6 | -1.2 [-6.2, 3.4] | 9 | 57.0 |
| L_soak | + (16) | 60.6 | 50.9 | +9.7 | 60.6 | -1.3 [-4.4, 2.1] | 7 | 54.4 |
| j_antiHorde | + (16) | 58.8 | 50.7 | +8.1 | 58.4 | -3.5 [-10.2, 2.8] | 8 | 54.0 |
| j_antiTank | + (16) | 58.5 | 51.2 | +7.3 | 58.4 | -3.5 [-8.0, 0.7] | 7 | 54.7 |
| j_burst | + (16) | 58.2 | 49.2 | +9.0 | 58.2 | -3.7 [-7.7, 0.0] | 6 | 57.7 |
| j_antiChar | + (16) | 57.3 | 50.8 | +6.5 | 57.3 | -4.6 [-7.3, -2.0] | 3 | 55.7 |
| j_bodies | − (16) | 56.7 | 50.6 | +6.2 | 54.6 | -7.3 [-14.7, -0.4] | 4 | 49.8 |
| pts | + (16) | 56.1 | 51.3 | +4.8 | 54.4 | -7.4 [-15.3, 0.2] | 4 | 49.8 |
| OC | + (16) | 55.7 | 50.2 | +5.5 | 38.5 | -23.3 [-30.4, -16.4] | 1 | 36.7 |
| j_actor | − (16) | 55.3 | 50.8 | +4.5 | 54.1 | -7.8 [-16.7, 0.8] | 5 | 48.3 |
| j_denial | + (16) | 55.1 | 50.9 | +4.2 | 54.3 | -7.6 [-14.8, -0.5] | 4 | 51.4 |
| debuffKind | − (16) | 54.9 | 48.6 | +6.2 | 22.9 | -38.9 [-47.2, -29.9] | 1 | 22.0 |
| epic | + (16) | 54.8 | 50.0 | +4.8 | 15.5 | -46.4 [-54.4, -37.9] | 0 | 15.1 |
| kc_elite | + (16) | 54.7 | 50.0 | +4.6 | 54.6 | -7.2 [-14.5, 0.6] | 4 | 51.7 |
| inv | + (16) | 54.2 | 50.9 | +3.3 | 31.7 | -30.2 [-38.3, -21.6] | 1 | 36.7 |
| kc_tank | + (16) | 54.1 | 50.5 | +3.6 | 54.1 | -7.8 [-14.3, -2.2] | 4 | 51.1 |
| M | + (16) | 53.9 | 51.4 | +2.5 | 40.6 | -21.3 [-27.7, -15.0] | 0 | 42.9 |
| deepStrike | + (16) | 53.9 | 49.4 | +4.5 | 22.6 | -39.2 [-47.8, -30.5] | 0 | 23.4 |
| L_hold | + (16) | 53.8 | 50.4 | +3.5 | 53.4 | -8.4 [-14.8, -1.8] | 5 | 52.2 |
| joinSquad | + (16) | 53.5 | 49.3 | +4.2 | 26.6 | -35.3 [-42.4, -28.4] | 0 | 28.4 |
| enabler | + (16) | 53.4 | 49.7 | +3.7 | 12.4 | -49.5 [-58.3, -39.9] | 1 | 13.5 |
| kc_char | − (16) | 53.4 | 50.7 | +2.7 | 53.4 | -8.5 [-17.5, -0.0] | 4 | 50.6 |
| lifted | + (16) | 53.2 | 49.9 | +3.3 | 26.4 | -35.5 [-43.3, -28.3] | 0 | 27.1 |
| j_screen | − (16) | 53.2 | 51.1 | +2.1 | 51.8 | -10.1 [-18.1, -2.3] | 4 | 48.0 |
| cpKind | + (16) | 53.2 | 50.4 | +2.8 | 13.1 | -48.7 [-54.3, -43.0] | 0 | 13.0 |
| unlock | − (16) | 53.1 | 50.2 | +2.9 | 26.6 | -35.3 [-44.7, -26.9] | 0 | 21.3 |
| infantry | + (16) | 52.7 | 49.3 | +3.4 | 24.5 | -37.4 [-43.8, -30.5] | 0 | 27.3 |
| vehicle | − (16) | 52.5 | 49.7 | +2.8 | 20.0 | -41.9 [-51.1, -32.8] | 0 | 23.9 |
| psyker | + (16) | 52.5 | 49.8 | +2.7 | 10.0 | -51.8 [-60.0, -42.5] | 1 | 9.6 |
| j_depends | + (16) | 52.4 | 49.5 | +2.9 | 44.1 | -17.8 [-26.4, -9.7] | 4 | 40.6 |
| shockKind | − (16) | 52.3 | 49.6 | +2.7 | 14.2 | -47.7 [-55.0, -40.6] | 0 | 13.5 |
| j_hide | + (16) | 52.3 | 49.8 | +2.5 | 24.1 | -37.8 [-44.7, -31.0] | 0 | 26.7 |
| fly | − (16) | 52.3 | 50.0 | +2.2 | 14.7 | -47.1 [-52.3, -40.9] | 0 | 14.5 |
| j_reach | + (16) | 52.0 | 51.3 | +0.8 | 47.2 | -14.7 [-23.9, -5.7] | 4 | 45.7 |
| lone | + (16) | 51.9 | 50.1 | +1.9 | 6.3 | -55.5 [-61.4, -49.5] | 0 | 5.3 |
| scouts | + (16) | 51.9 | 50.2 | +1.7 | 8.9 | -53.0 [-59.9, -46.0] | 0 | 8.2 |
| j_elusive | + (16) | 51.9 | 50.6 | +1.3 | 9.2 | -52.6 [-58.5, -46.5] | 0 | 8.1 |
| battleline | + (16) | 51.8 | 49.8 | +2.0 | 8.9 | -52.9 [-59.6, -46.3] | 0 | 7.8 |
| j_cp | + (16) | 51.6 | 49.8 | +1.8 | 9.2 | -52.6 [-57.9, -47.7] | 0 | 9.4 |
| aura | + (16) | 51.4 | 49.7 | +1.7 | 7.0 | -54.8 [-59.1, -50.2] | 0 | 8.1 |
| models | + (16) | 51.3 | 49.7 | +1.6 | 31.0 | -30.8 [-38.1, -23.8] | 1 | 32.9 |
| j_engineM | + (16) | 51.3 | 50.0 | +1.3 | 4.9 | -57.0 [-61.8, -51.9] | 0 | 4.7 |
| j_rival | − (16) | 51.3 | 49.9 | +1.4 | 6.1 | -55.8 [-62.2, -49.0] | 0 | 6.2 |
| joinLead | − (16) | 51.1 | 50.4 | +0.7 | 24.6 | -37.3 [-47.3, -27.8] | 0 | 21.4 |
| infiltrators | + (16) | 51.1 | 49.7 | +1.4 | 5.8 | -56.1 [-61.6, -50.4] | 0 | 7.2 |
| Ld | + (15) | 50.8 | 50.3 | +0.6 | 17.6 | -44.3 [-53.3, -35.4] | 0 | 17.8 |
| L_score | + (16) | 50.8 | 49.4 | +1.4 | 50.1 | -11.8 [-18.2, -4.9] | 3 | 53.1 |
| kc_horde | + (16) | 50.8 | 51.3 | -0.5 | 50.7 | -11.1 [-18.9, -2.7] | 4 | 48.6 |
| synapse | + (16) | 50.7 | 50.1 | +0.7 | 2.2 | -59.7 [-63.9, -55.4] | 0 | 2.1 |
| L_presence | − (15) | 50.5 | 50.5 | -0.0 | 49.7 | -12.2 [-20.8, -4.1] | 4 | 47.3 |
| L_actions | − (15) | 50.5 | 50.6 | -0.2 | 49.7 | -12.2 [-20.9, -4.1] | 4 | 47.2 |
| Sv | + (15) | 50.4 | 50.1 | +0.3 | 31.8 | -30.0 [-37.9, -22.2] | 0 | 33.6 |
| transportKw | − (16) | 50.3 | 49.6 | +0.8 | 7.7 | -54.2 [-59.7, -48.7] | 0 | 8.5 |
| leader | − (15) | 50.3 | 50.8 | -0.4 | 18.9 | -42.9 [-51.8, -34.7] | 0 | 16.2 |
| W | − (16) | 50.3 | 49.6 | +0.6 | 44.7 | -17.2 [-24.5, -10.1] | 1 | 49.6 |
| transportCap | − (16) | 50.3 | 49.5 | +0.8 | 7.8 | -54.1 [-59.6, -48.6] | 0 | 8.8 |
| j_tough | + (16) | 50.3 | 49.2 | +1.1 | 50.1 | -11.8 [-18.4, -5.5] | 4 | 46.2 |
| fightsFirst | + (15) | 50.1 | 50.4 | -0.3 | 2.8 | -59.1 [-65.6, -52.7] | 0 | 3.2 |
| j_oc | + (16) | 49.8 | 50.1 | -0.3 | 48.9 | -13.0 [-18.8, -7.2] | 3 | 50.6 |
| stealth | + (15) | 49.7 | 50.3 | -0.5 | 4.9 | -57.0 [-63.1, -51.1] | 0 | 5.6 |
| L_spawn | − (15) | 49.6 | 50.0 | -0.4 | 0.4 | -61.5 [-66.5, -56.5] | 0 | 0.5 |
| ocPer100 | + (16) | 49.5 | 50.2 | -0.7 | 48.5 | -13.4 [-19.6, -7.1] | 3 | 49.8 |
| j_lands | + (15) | 49.4 | 50.0 | -0.6 | 32.2 | -29.7 [-39.3, -20.1] | 1 | 37.5 |
| j_hold | + (16) | 49.4 | 50.1 | -0.7 | 49.0 | -12.9 [-19.3, -6.3] | 3 | 50.0 |
| meleeShare | + (15) | 49.4 | 50.8 | -1.4 | 19.9 | -41.9 [-49.8, -33.6] | 0 | 24.6 |
| j_shock | − (15) | 49.2 | 49.6 | -0.4 | 1.6 | -60.2 [-66.5, -54.0] | 0 | 2.6 |
| fnp | − (14) | 49.1 | 50.6 | -1.4 | 7.5 | -54.4 [-61.7, -47.3] | 0 | 10.7 |
| monster | − (15) | 49.0 | 50.4 | -1.4 | 5.3 | -56.6 [-60.9, -51.9] | 0 | 3.6 |
| range | − (15) | 48.6 | 50.0 | -1.4 | 38.0 | -23.9 [-32.6, -14.9] | 2 | 43.4 |
| j_recover | − (12) | 47.5 | 49.8 | -2.3 | 11.7 | -50.1 [-56.3, -44.0] | 0 | 15.4 |
| T | + (12) | 45.2 | 49.6 | -4.4 | 32.9 | -29.0 [-35.7, -22.4] | 0 | 34.4 |
| woundsPer100 | − (11) | 43.9 | 50.2 | -6.2 | 43.0 | -18.9 [-25.0, -12.8] | 0 | 47.2 |
| character | + (9) | 42.8 | 50.1 | -7.4 | 15.9 | -46.0 [-53.1, -39.0] | 0 | 17.5 |
| supportKinds | − (11) | 41.3 | 50.3 | -9.0 | 19.8 | -42.1 [-49.4, -35.3] | 0 | 21.6 |
What the fits value: coefficients
Pairwise logistic, every feature, L2
Fitted on all 16 armies at λ 300 (λ in the held-out folds: 300, 100, 30). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| epic | +0.07 | 16 | 0.06 to 0.24 |
| psyker | +0.07 | 16 | 0.06 to 0.20 |
| lone | +0.07 | 16 | 0.06 to 0.25 |
| L_soak | +0.07 | 16 | 0.05 to 0.22 |
| M | +0.06 | 16 | 0.04 to 0.27 |
| j_antiElite | +0.06 | 16 | 0.06 to 0.14 |
| j_antiTank | +0.06 | 16 | 0.04 to 0.26 |
| j_antiHorde | +0.06 | 16 | 0.05 to 0.14 |
| debuffKind | -0.06 | 16 | -0.16 to -0.05 |
| fly | -0.05 | 16 | -0.19 to -0.03 |
| deepStrike | +0.05 | 16 | 0.05 to 0.13 |
| j_engineM | +0.05 | 16 | 0.03 to 0.12 |
| L_kill | +0.04 | 15 | -0.03 to 0.06 |
| cpKind | +0.04 | 16 | 0.03 to 0.10 |
| infantry | +0.04 | 16 | 0.03 to 0.14 |
| enabler | +0.04 | 16 | 0.03 to 0.17 |
| j_hide | +0.04 | 16 | 0.03 to 0.14 |
| L_hold | +0.04 | 16 | 0.02 to 0.07 |
| joinSquad | +0.04 | 16 | 0.03 to 0.11 |
| unlock | -0.03 | 16 | -0.08 to -0.01 |
Pairwise logistic, every feature, L1
Fitted on all 16 armies at λ 10 (λ in the held-out folds: 3, 10). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| M | +0.19 | 16 | 0.17 to 0.38 |
| epic | +0.17 | 16 | 0.15 to 0.35 |
| j_antiTank | +0.16 | 16 | 0.13 to 0.46 |
| L_soak | +0.16 | 16 | 0.12 to 0.45 |
| psyker | +0.16 | 16 | 0.14 to 0.31 |
| j_antiHorde | +0.16 | 16 | 0.07 to 0.25 |
| lone | +0.15 | 16 | 0.16 to 0.38 |
| fly | -0.14 | 16 | -0.33 to -0.10 |
| debuffKind | -0.13 | 16 | -0.21 to -0.11 |
| infantry | +0.12 | 15 | 0.00 to 0.24 |
| cpKind | +0.08 | 16 | 0.01 to 0.12 |
| j_engineM | +0.08 | 16 | 0.07 to 0.14 |
| deepStrike | +0.08 | 16 | 0.02 to 0.19 |
| enabler | +0.05 | 16 | 0.01 to 0.22 |
| lifted | +0.05 | 16 | 0.03 to 0.20 |
| kc_elite | +0.05 | 14 | 0.00 to 0.20 |
| L_presence | +0.05 | 15 | 0.00 to 0.33 |
| joinSquad | +0.04 | 15 | 0.00 to 0.16 |
| j_hide | +0.04 | 15 | 0.00 to 0.29 |
| j_antiElite | +0.02 | 14 | 0.00 to 0.41 |
Pairwise logistic, the seven lines only, L2
Fitted on all 16 armies at λ 100 (λ in the held-out folds: 100, 3, 10, 300, 30, 1000). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| L_soak | +0.18 | 16 | 0.06 to 0.51 |
| L_kill | +0.16 | 16 | 0.03 to 0.18 |
| L_score | +0.09 | 16 | 0.04 to 0.37 |
| L_hold | +0.07 | 14 | -0.22 to 0.07 |
| L_presence | +0.04 | 15 | -0.02 to 0.39 |
| L_spawn | -0.04 | 15 | -0.28 to 0.00 |
Pairwise logistic, datasheet + keywords + support only (no ledger), L2
Fitted on all 16 armies at λ 10 (λ in the held-out folds: 3, 30, 10). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| M | +0.44 | 16 | 0.29 to 0.64 |
| infantry | +0.43 | 16 | 0.21 to 0.64 |
| epic | +0.34 | 16 | 0.20 to 0.41 |
| lone | +0.29 | 16 | 0.14 to 0.39 |
| fly | -0.27 | 16 | -0.36 to -0.18 |
| supportKinds | -0.26 | 16 | -0.41 to -0.09 |
| psyker | +0.21 | 16 | 0.12 to 0.27 |
| debuffKind | -0.20 | 16 | -0.30 to -0.14 |
| T | +0.18 | 16 | 0.03 to 0.34 |
| cpKind | +0.18 | 16 | 0.13 to 0.23 |
| enabler | +0.18 | 16 | 0.11 to 0.31 |
| transportKw | -0.15 | 16 | -0.33 to -0.07 |
| vehicle | -0.11 | 16 | -0.18 to -0.01 |
| W | +0.11 | 16 | 0.02 to 0.18 |
| deepStrike | +0.11 | 16 | 0.07 to 0.15 |
| synapse | +0.10 | 15 | -0.00 to 0.19 |
| stealth | -0.10 | 16 | -0.15 to -0.01 |
| pts | +0.10 | 16 | 0.02 to 0.17 |
| range | -0.10 | 16 | -0.17 to -0.05 |
| character | -0.09 | 15 | -0.16 to 0.05 |
Pairwise logistic, design/jobs.md's jobs only, L2
Fitted on all 16 armies at λ 1000 (λ in the held-out folds: 1000, 300, 30, 100). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| j_antiElite | +0.05 | 16 | 0.04 to 0.19 |
| j_antiHorde | +0.04 | 16 | 0.03 to 0.19 |
| j_antiTank | +0.04 | 16 | 0.03 to 0.23 |
| j_burst | +0.04 | 13 | -0.04 to 0.06 |
| j_antiChar | +0.03 | 13 | -0.02 to 0.05 |
| j_engineM | +0.03 | 16 | 0.01 to 0.18 |
| j_bodies | -0.02 | 16 | -0.18 to -0.02 |
| j_hide | +0.02 | 16 | 0.02 to 0.21 |
| j_oc | +0.02 | 16 | 0.01 to 0.05 |
| j_hold | +0.02 | 16 | 0.02 to 0.07 |
| j_cp | +0.02 | 16 | 0.01 to 0.15 |
| j_denial | +0.02 | 12 | -0.07 to 0.03 |
| j_elusive | +0.01 | 16 | 0.01 to 0.08 |
| j_lands | +0.01 | 15 | -0.00 to 0.02 |
| j_screen | -0.01 | 15 | -0.04 to 0.01 |
| j_shock | -0.01 | 15 | -0.05 to 0.03 |
| j_recover | +0.01 | 15 | -0.00 to 0.03 |
| j_tough | +0.00 | 9 | -0.06 to 0.02 |
| j_actor | +0.00 | 14 | -0.01 to 0.12 |
Pairwise logistic, every feature + log Value, L2
Fitted on all 16 armies at λ 100 (λ in the held-out folds: 300, 100, 30). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| lone | +0.14 | 16 | 0.05 to 0.25 |
| epic | +0.13 | 16 | 0.06 to 0.24 |
| M | +0.13 | 16 | 0.04 to 0.27 |
| psyker | +0.13 | 16 | 0.04 to 0.19 |
| j_antiTank | +0.11 | 16 | 0.04 to 0.26 |
| fly | -0.10 | 16 | -0.19 to -0.03 |
| L_soak | +0.10 | 16 | 0.04 to 0.20 |
| debuffKind | -0.10 | 16 | -0.16 to -0.05 |
| j_antiHorde | +0.09 | 16 | 0.03 to 0.14 |
| j_antiElite | +0.09 | 16 | 0.04 to 0.13 |
| deepStrike | +0.09 | 16 | 0.04 to 0.13 |
| j_engineM | +0.08 | 16 | 0.03 to 0.12 |
| infantry | +0.08 | 16 | 0.03 to 0.12 |
| j_hide | +0.07 | 16 | 0.03 to 0.11 |
| L_presence | +0.07 | 16 | 0.00 to 0.17 |
| enabler | +0.07 | 16 | 0.03 to 0.17 |
| joinSquad | +0.06 | 16 | 0.02 to 0.11 |
| cpKind | +0.06 | 16 | 0.03 to 0.08 |
| logValue | +0.06 | 16 | 0.02 to 0.08 |
| OC | +0.06 | 16 | 0.02 to 0.10 |
Ordinal logistic on letters, every feature, L2
Fitted on all 16 armies at λ 30 (λ in the held-out folds: 30, 300, 3). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| M | +0.24 | 16 | 0.05 to 0.44 |
| j_antiTank | +0.23 | 16 | 0.05 to 0.63 |
| lone | +0.21 | 16 | 0.04 to 0.42 |
| epic | +0.21 | 16 | 0.05 to 0.40 |
| psyker | +0.21 | 16 | 0.04 to 0.36 |
| fly | -0.20 | 16 | -0.37 to -0.05 |
| debuffKind | -0.17 | 16 | -0.34 to -0.05 |
| j_hide | +0.14 | 16 | 0.03 to 0.30 |
| j_engineM | +0.14 | 16 | 0.04 to 0.23 |
| L_soak | +0.14 | 16 | 0.05 to 0.54 |
| infantry | +0.14 | 16 | 0.04 to 0.26 |
| enabler | +0.13 | 16 | 0.04 to 0.32 |
| deepStrike | +0.12 | 16 | 0.03 to 0.20 |
| j_antiElite | +0.11 | 15 | -0.04 to 0.28 |
| j_antiHorde | +0.11 | 16 | 0.03 to 0.24 |
| L_presence | +0.11 | 16 | 0.02 to 0.57 |
| L_hold | +0.11 | 16 | 0.04 to 0.38 |
| kc_elite | +0.10 | 16 | 0.03 to 0.19 |
| j_antiChar | -0.10 | 11 | -0.40 to 0.02 |
| range | -0.10 | 16 | -0.15 to -0.01 |
Pairwise logistic on within-army ranks, every feature, L2
Fitted on all 16 armies at λ 3 (λ in the held-out folds: 3). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| M | +0.79 | 16 | 0.63 to 0.90 |
| psyker | +0.76 | 16 | 0.51 to 0.83 |
| L_soak | +0.72 | 16 | 0.44 to 0.79 |
| epic | +0.71 | 16 | 0.57 to 0.77 |
| fly | -0.64 | 16 | -0.81 to -0.30 |
| j_rival | -0.63 | 16 | -0.70 to -0.43 |
| debuffKind | -0.59 | 16 | -0.73 to -0.48 |
| j_hide | +0.58 | 16 | 0.49 to 0.60 |
| infantry | +0.56 | 16 | 0.36 to 0.63 |
| L_kill | +0.53 | 16 | 0.38 to 0.61 |
| OC | +0.50 | 16 | 0.44 to 0.56 |
| L_hold | +0.50 | 16 | 0.35 to 0.57 |
| deepStrike | +0.50 | 16 | 0.34 to 0.59 |
| joinSquad | +0.49 | 16 | 0.29 to 0.63 |
| j_antiTank | +0.44 | 16 | 0.27 to 0.54 |
| j_engineM | +0.42 | 16 | 0.30 to 0.46 |
| enabler | +0.39 | 16 | 0.22 to 0.50 |
| j_denial | +0.38 | 16 | 0.27 to 0.40 |
| lone | +0.37 | 16 | 0.24 to 0.50 |
| j_oc | -0.34 | 16 | -0.39 to -0.24 |
Role intercepts + Value within role (pairwise fit)
Fitted on all 16 armies at λ 100 (λ in the held-out folds: 100, 30, 10, 1). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| withinRoleValue | +0.23 | 16 | 0.15 to 0.47 |
| Hammer | +0.09 | 16 | 0.08 to 0.38 |
| Disruptors | -0.08 | 16 | -0.21 to -0.05 |
| Action units | +0.06 | 16 | 0.03 to 0.39 |
| Transports | -0.05 | 16 | -0.38 to -0.02 |
| Heavy anti-armour | +0.03 | 16 | 0.02 to 0.95 |
| Anvil | -0.03 | 16 | -0.75 to -0.01 |
| Force multipliers | -0.03 | 14 | -0.06 to 0.02 |
| Screens | -0.01 | 15 | -0.36 to 0.01 |
Archetype intercepts + Value within archetype (pairwise fit)
Fitted on all 16 armies at λ 3 (λ in the held-out folds: 1, 100, 3). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| archetype0 | -0.94 | 16 | -1.13 to -0.13 |
| archetype6 | +0.39 | 16 | 0.08 to 0.47 |
| withinArchValue | +0.36 | 16 | 0.13 to 0.47 |
| archetype1 | +0.35 | 16 | 0.03 to 0.54 |
| archetype3 | +0.35 | 16 | 0.02 to 0.55 |
| archetype2 | -0.32 | 16 | -0.54 to -0.02 |
| archetype5 | +0.14 | 16 | 0.01 to 0.20 |
| archetype4 | +0.02 | 11 | -0.10 to 0.15 |
Value + value over replacement (pairwise fit)
Fitted on all 16 armies at λ 3 (λ in the held-out folds: 10, 1, 30, 3). Coefficient per army-SD of the feature (on the score's logit scale); folds = how many of the 16 leave-one-out fits share its sign; range over the folds.
| Feature | Coefficient | Folds same sign | Range over folds |
|---|---|---|---|
| logValue | +0.58 | 16 | 0.35 to 0.64 |
| vorBest | -0.12 | 13 | -0.19 to 0.06 |
The ledger's line weights against the fitted ones
Ledger: each line's mean share of a unit's Value (Track A). Fitted: the lines-only pairwise model's coefficients as a share of their absolute sum (a different scale: per army-SD of log(1 + the line), so read the order and sign, not the size).
| Line | Ledger share of Value | Fitted share (sign) |
|---|---|---|
| kill | 85.3% | +27.7% |
| soak | 2.7% | +30.9% |
| score | 3.6% | +15.4% |
| actions | 0.9% | +0.0% |
| hold | 5.6% | +11.3% |
| spawn | 0.0% | −7.1% |
| presence | 1.9% | +7.6% |
Boosting: split-gain importance (all 16 armies)
| Feature | Share of gain |
|---|---|
| L_kill | 14.1% |
| L_soak | 10.3% |
| OC | 6.4% |
| j_antiHorde | 6.0% |
| j_oc | 4.4% |
| j_denial | 3.9% |
| infiltrators | 3.9% |
| infantry | 3.8% |
| L_hold | 3.0% |
| deepStrike | 2.9% |
| joinSquad | 2.8% |
| M | 2.6% |
| j_actor | 2.3% |
| j_rival | 1.9% |
| kc_elite | 1.9% |
By army (held out)
| Army | Lists | Pairs | a_value | b_L_kill | c_all_l2 | c_sheet_l2 | d_ord_all | h_gbm | j_all_max | x_knn | x_ens_sheet | k_share |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Tyranids | 3 | 583 | 89.5 | 82.2 | 77.0 | 74.6 | 79.6 | 72.9 | 89.5 | 72.0 | 89.0 | 95.5 |
| Space Marines | 2 | 1515 | 62.1 | 61.2 | 59.4 | 55.4 | 62.0 | 67.6 | 54.3 | 63.0 | 59.7 | 83.8 |
| Adeptus Mechanicus | 1 | 322 | 71.4 | 66.8 | 52.5 | 50.0 | 56.2 | 61.2 | 76.7 | 58.4 | 63.7 | 71.1 |
| Agents of the Imperium | 1 | 223 | 56.1 | 57.8 | 61.4 | 67.7 | 62.8 | 57.0 | 39.0 | 62.3 | 67.3 | 28.7 |
| Astra Militarum | 1 | 1319 | 67.6 | 70.1 | 65.1 | 66.1 | 65.5 | 62.0 | 58.4 | 61.3 | 69.6 | 64.4 |
| Blood Angels | 1 | 124 | 61.3 | 63.7 | 50.0 | 75.8 | 49.2 | 64.5 | 80.6 | 71.8 | 72.6 | 52.4 |
| Dark Angels | 1 | 2325 | 53.7 | 57.8 | 52.4 | 50.2 | 49.2 | 58.1 | 52.4 | 50.8 | 52.3 | 54.2 |
| Death Guard | 1 | 476 | 55.0 | 52.9 | 56.5 | 47.1 | 51.7 | 52.5 | 57.8 | 58.4 | 52.1 | 47.3 |
| Drukhari | 1 | 207 | 58.0 | 58.9 | 58.5 | 53.6 | 52.2 | 58.0 | 61.8 | 60.9 | 57.5 | 73.9 |
| Emperor's Children | 1 | 191 | 69.6 | 71.2 | 67.0 | 63.9 | 70.2 | 57.6 | 68.6 | 49.2 | 66.0 | 73.3 |
| Genestealer Cults | 2 | 112 | 55.4 | 68.8 | 66.1 | 78.6 | 65.2 | 67.0 | 56.3 | 66.1 | 67.9 | 80.4 |
| Leagues of Votann | 1 | 105 | 61.0 | 64.8 | 62.9 | 74.3 | 52.4 | 54.3 | 50.5 | 62.9 | 74.3 | 83.8 |
| Orks | 1 | 606 | 63.5 | 62.9 | 69.0 | 67.3 | 63.5 | 66.5 | 62.7 | 80.5 | 68.0 | 60.6 |
| Space Wolves | 1 | 140 | 50.7 | 52.1 | 52.9 | 65.0 | 60.0 | 37.9 | 47.9 | 38.6 | 55.7 | 87.1 |
| Thousand Sons | 1 | 309 | 38.2 | 38.5 | 61.8 | 64.7 | 65.0 | 60.5 | 40.8 | 57.0 | 52.8 | 66.0 |
| World Eaters | 3 | 219 | 76.7 | 72.6 | 70.3 | 68.9 | 66.2 | 69.4 | 81.3 | 73.1 | 79.0 | 90.9 |
Gemini (GEMINI AGGREGATE letters, reference only)
Every model above was scored on the held-out army's Gemini pairs (the Gemini column; human letters only in training, so Gemini's letters are a second, independent target). Against the ledger on those pairs (15 armies, bootstrap over armies):
| Model | Gemini pairs | vs ledger [95%] | Armies better |
|---|---|---|---|
| List share alone (DIAGNOSTIC ONLY: popularity, never in a letter) | 67.1 | +10.4 [1.2, 19.2] | 10 of 15 |
| POST HOC: datasheet-only model on within-army ranks, L2 | 66.3 | +9.6 [1.2, 18.8] | 10 of 15 |
| Pairwise logistic, datasheet + keywords + support only (no ledger), L2 | 64.0 | +7.4 [1.3, 13.3] | 11 of 15 |
| Pairwise logistic on within-army ranks, every feature, L2 | 63.2 | +6.5 [2.0, 11.3] | 10 of 15 |
| Ensemble: ledger Value rank + datasheet-only model rank | 62.9 | +6.2 [3.0, 9.8] | 11 of 15 |
| POST HOC: datasheet-only features + log Value, L2 | 62.2 | +5.5 [0.9, 10.3] | 9 of 15 |
| Pairwise logistic, every feature, L1 | 62.0 | +5.3 [-0.4, 11.2] | 9 of 15 |
| Pairwise logistic, every feature, L2 | 61.0 | +4.3 [-0.3, 9.4] | 10 of 15 |
| Pairwise logistic, every feature + log Value, L2 | 60.8 | +4.1 [-0.5, 9.1] | 10 of 15 |
| Shallow pairwise boosting (depth 2, 100 rounds) | 60.4 | +3.7 [-1.2, 8.7] | 10 of 15 |
| POST HOC: ensemble, ledger Value rank + every-feature ranks model rank | 59.6 | +2.9 [0.7, 5.3] | 9 of 15 |
| h boosting, trained on human + 0.5 × Gemini | 59.0 | +2.3 [-3.7, 7.9] | 10 of 15 |
| Ordinal logistic on letters, every feature, L2 | 58.9 | +2.2 [-3.3, 8.0] | 7 of 15 |
| 25 nearest neighbours across armies (24 features) | 58.7 | +2.0 [-2.5, 6.6] | 10 of 15 |
| Archetype intercepts + Value within archetype (pairwise fit) | 57.9 | +1.2 [-3.3, 6.5] | 9 of 15 |
| Ensemble: ledger Value rank + boosting rank | 57.5 | +0.8 [-2.3, 3.4] | 9 of 15 |
| Value + value over replacement (pairwise fit) | 57.4 | +0.7 [-0.1, 1.5] | 9 of 15 |
| Pairwise logistic, design/jobs.md's jobs only, L2 | 57.0 | +0.3 [-5.3, 5.4] | 8 of 15 |
| Leaders at their best joined form's Value | 56.9 | +0.3 [-2.2, 3.0] | 6 of 15 |
| Every unit at max(solo, best joined as leader or squad) | 56.8 | +0.1 [-3.9, 4.1] | 10 of 15 |
| Ledger Value (Track A) | 56.7 | +0.0 [0.0, 0.0] | 0 of 15 |
| Leaders at max(solo, best joined) | 56.5 | -0.2 [-3.0, 2.6] | 7 of 15 |
| Pairwise logistic, the seven lines only, L2 | 56.5 | -0.2 [-4.0, 3.3] | 7 of 15 |
| Value within archetype (T-716) | 56.4 | -0.3 [-4.3, 4.3] | 5 of 15 |
| Absolute output: Value × points | 55.6 | -1.1 [-4.8, 2.2] | 7 of 15 |
| Role intercepts + Value within role (pairwise fit) | 54.9 | -1.8 [-5.9, 2.0] | 6 of 15 |
| Value within role (Jordan's eight, by rules) | 54.3 | -2.4 [-7.2, 2.9] | 6 of 15 |
| Value over replacement: less the best other same-role unit | 53.8 | -2.9 [-5.4, -0.6] | 5 of 15 |
| Value over replacement: less the role's median | 53.6 | -3.1 [-7.4, 1.3] | 5 of 15 |
| Replaceability: fewer same-role units with a higher Value | 52.3 | -4.4 [-7.2, -2.1] | 2 of 15 |
| Points alone (sign chosen) | 49.8 | -6.9 [-14.9, 1.0] | 6 of 15 |
| Replaceability: fewer same-army units at least as good at its main job | 45.2 | -11.5 [-18.1, -4.0] | 2 of 15 |
Trained on human + 0.5 × Gemini pairs (Gemini at half a single human list's weight), scored on held-out human pairs only:
| Model | Human only | Human + 0.5 × Gemini | Difference |
|---|---|---|---|
| Pairwise logistic, every feature, L2 | 61.4 | 62.1 | +0.7 (8 armies up, 6 down) |
| Pairwise logistic, the seven lines only, L2 | 59.6 | 59.3 | -0.3 (6 armies up, 8 down) |
| Shallow pairwise boosting (depth 2, 100 rounds) | 60.4 | 61.6 | +1.2 (8 armies up, 8 down) |
The eight roles as the rules assign them (counts per army)
| Army | Anvil | Hammer | Heavy anti-armour | Action units | Screens | Force multipliers | Disruptors | Transports |
|---|---|---|---|---|---|---|---|---|
| Tyranids | 0 | 11 | 2 | 6 | 0 | 9 | 21 | 3 |
| Space Marines | 4 | 15 | 2 | 12 | 0 | 12 | 28 | 11 |
| Adeptus Mechanicus | 0 | 11 | 0 | 2 | 0 | 4 | 15 | 2 |
| Agents of the Imperium | 0 | 7 | 0 | 1 | 0 | 7 | 10 | 4 |
| Astra Militarum | 1 | 15 | 3 | 8 | 2 | 16 | 19 | 8 |
| Blood Angels | 1 | 3 | 0 | 2 | 0 | 9 | 5 | 0 |
| Dark Angels | 3 | 14 | 3 | 14 | 0 | 16 | 33 | 10 |
| Death Guard | 0 | 8 | 0 | 2 | 0 | 10 | 14 | 2 |
| Drukhari | 0 | 3 | 0 | 7 | 0 | 5 | 19 | 3 |
| Emperor's Children | 0 | 5 | 0 | 5 | 1 | 4 | 6 | 2 |
| Genestealer Cults | 1 | 19 | 1 | 11 | 2 | 16 | 29 | 11 |
| Leagues of Votann | 0 | 2 | 0 | 3 | 0 | 6 | 8 | 3 |
| Orks | 0 | 9 | 0 | 8 | 0 | 8 | 20 | 10 |
| Space Wolves | 1 | 3 | 0 | 3 | 0 | 6 | 8 | 0 |
| Thousand Sons | 0 | 9 | 0 | 3 | 0 | 11 | 10 | 2 |
| World Eaters | 0 | 4 | 0 | 4 | 0 | 5 | 15 | 2 |