T-632's next step. T-632 found that on 52 Tyranid units every model lands at 68 to 73% on units it hasn't seen, and its learning curve was still rising. So this card asks whether more examples help: put every army's graded units through the same variables, and test on an army the model never saw. Command: node tools/diagnose.js --pooled (17.6 minutes on 4 worker threads; --quick for a 3-minute smoke run). Every number below is in 2026-10-04-pooled-diagnosis-tables.md and -pooled-diagnosis.json. This is a report: no ledger number, page default, data file, setting, benchmark or frozen input changed. Our own numbers; unit names only.
The short answer
- No. Pooling 16 armies doesn't let any model beat the ledger on an army it has never seen.
- Left out one army at a time, every model lands at 58 to 63% of pairs ordered right (the 16 armies pooled).
- The ledger as it is gets 60.0%. So does its re-fitted form (59.9%).
- The best flexible model averages 2 points better per army (boosted trees, +2.2 ± 2.2, better on 10 of 16 armies). That is inside the noise.
- Scrambled letters give 48 to 51%, so the models learn something real. They learn roughly what the ledger already knows.
- The one model that seemed to win was cheating through twins. Nearest neighbours scored 63.4% pooled. That came entirely from the Space Marine family: the Dark Angels', Blood Angels' and Space Wolves' pools hold the Space Marines' own datasheets. When all four Marine armies are left out together, every flexible model falls to 50 to 54% on them. The ledger holds 57%.
- More armies barely help. The learning curve over 3, 6, 9, 12 and 15 training armies is flat within noise: the trees go 58 → 59 → 57 → 59 → 60% (army mean). T-632's hope was that several hundred graded units would carry a flexible model toward the reviewers. With 618 units over 16 armies, it doesn't.
- Pooling doesn't help Tyranids. Trained only on the other 15 armies, every flexible model does worse on Tyranids (58 to 70%) than the ledger (73%). Adding the other armies to Tyranids' own training units lifts the flexible models by 0.6 to 2.6 points. They stay 4.5 or more points under the ledger.
- Within an army, the reviewers' taste is learnable; across armies, it isn't. With units held out inside each army (the model sees the other units of the same army), the interactions model reaches 67.4% pooled, against the ledger's 60.3%. It gains most in the armies the ledger reads worst: Drukhari, Thousand Sons, Death Guard, Dark Angels and Astra Militarum. What the reviewers reward in one army doesn't travel to the next.
- What does travel, in every army:
- Kill (the ledger's own line) is the steadiest signal of all. Alone it orders 60% of pairs, with a positive sign in 11 armies and a negative sign in none.
- Kill into elite targets and monsters is next: heavy infantry, cavalry and beasts, monsters, vehicles. These are T-632's "scarce answers", the strongest of its nine candidates here.
- Being cheap to fit helps a little: the Value at the cheapest size.
- A floor on Score: units at the very bottom of their army for objective bodies lose about half a letter.
- What doesn't travel:
- Precision, psychic and dedicated transports carry no signal across armies.
- Melee Soak and actions point the wrong way in most armies, as T-632 found for Tyranids.
- Which candidate mechanic to try next: T-632's number 2, "scarce answers". Its signal holds in 7 to 10 of 16 armies, under both the trees and the neighbours. Then a capped Score (number 3) and threat on arrival (number 1), both weaker. Arrival Kill as built in step 18 (T-633) was reverted, which fits: arrival's signal here is small.
- The misses shared across armies are leaders, psykers and transports.
- Rated higher by the reviewers than any model gives: Librarians, the Thousand Sons' Sorcerer and Infernal Master, the Miasmic Malignifier, the Venom, the Stormraven, Biovores and Jakhals.
- Rated lower: fast melee units (Serberys Sulphurhounds, Tzaangor Enlightened, Maulerfiend) and Dark Angels units the reviewer put at F.
- That is the all-armies report's finding again (3 Oct): a leader's or a transport's worth is in what it does for another unit, and no solo variable holds that.
The headline table
What's scored:
- The target: each graded unit's consensus, its mean letter (S 5 … F 0) over its own army's lists.
- The lists:
- Tyranids: the five headline lists, as T-632 used them.
- Space Marines: their two fitting lists, Auspex full and Tobias. Tistaminis and the superseded Auspex summary are left out.
- Every other army: every panel it has, all of them reference lists, as tools/ledger-all.js and tools/roles.js scored them. These include the two 10th-edition Leagues of Votann lists and the retired Orks list.
- The units: 618 graded units of the 16 armies' 733 ranked units. Ungraded units still count in each army's scales.
How it's measured:
- Pairwise: of two units in the same army that the consensus orders, the share a model orders the same way. A tie in the model's scores counts as wrong. Pairs across armies are never formed, in the fit or in the score.
- Pooled: every within-army pair counted once, so Dark Angels (93 graded) and Space Marines (84) weigh most.
- Army mean: the 16 armies weighed equally.
| Model | Left out: pooled | Left out: army mean | Tyranids | Space Marines | Left out vs lists | Units held out, pooled (± SE) | Scrambled, left out (SD) |
|---|---|---|---|---|---|---|---|
| Ceiling (each list against the others; 5 armies) | – | – | 88.5% | 91.0% | – | – | – |
| (a) the ledger as it is (step 6) | 60.0% | 57.8% | 73.3% | 60.4% | 62.3% | 60.3% ± 0.2 | 49.1% (0.2) |
| (a1) the ledger's form, its 7 weights re-fitted | 59.9% | 57.8% | 73.2% | 60.4% | 62.2% | 60.1% ± 0.2 | 49.4% (0.2) |
| (b) linear on all 144 variables, ridge or lasso | 59.1% | 59.8% | 64.8% | 62.1% | 58.8% | 62.2% ± 0.5 | 50.9% (3.4) |
| (b2) linear on a short core list (25), ridge | 58.6% | 57.7% | 69.5% | 58.6% | 59.4% | 60.9% ± 0.4 | 48.1% (0.8) |
| (c) linear + every pairwise interaction (kernel) | 59.1% | 59.7% | 57.7% | 63.7% | 58.2% | 67.4% ± 0.6 | 50.1% (1.8) |
| (d) boosted trees, depth 1 to 3, monotone | 61.8% | 60.0% | 62.0% | 62.3% | 60.4% | 65.9% ± 0.7 | 49.3% (0.0) |
| (e) nearest neighbours (any army) | 63.4%\* | 56.3% | 57.9% | 69.6%\* | 61.0% | 65.4% ± 0.5 | 49.9% (2.6) |
\* Twins: see the Space Marine family below.
Notes on the table:
- (a) has nothing to fit, so its "left out" figure is the same as its in-sample one. It was tuned on Tyranids and Space Marines, so on those two it is optimistic. On the other 14 armies it is an honest held-out figure: it never saw them.
- (a1) chose the strongest pull toward step 6 in 13 of 16 folds, so it is step 6, as on Tyranids alone. Re-fitting the seven weights on 15 armies finds nothing better for the 16th.
- Ceilings (each list against the mean letter of its army's other lists):
- Tyranids 88.5%, Space Marines 91.0%, World Eaters 89.3%, Genestealer Cults 84.9%.
- Leagues of Votann 59.3%: its two 10th-edition lists barely agree with each other.
- The other 11 armies have one list each, so they have no ceiling.
Paired against the re-fitted ledger, leaving one army out (per-army gain in points; SE over the 16 armies):
| Pair | Mean gain ± SE | Armies better | Pooled gain | Scrambled spread (SD) |
|---|---|---|---|---|
| (d) trees − (a1) | +2.2 ± 2.2 | 10 of 16 | +1.9 | 0.9 |
| (b) linear − (a1) | +2.0 ± 2.7 | 9 of 16 | −0.9 | 0.3 |
| (c) interactions − (a1) | +1.9 ± 3.0 | 10 of 16 | −0.9 | 0.6 |
| (b2) core linear − (a1) | −0.1 ± 2.1 | 6 of 16 | −1.4 | 2.6 |
| (e) neighbours − (a1) | −1.5 ± 2.8 | 7 of 16 | +3.4\* | 3.1 |
None is beyond one SE of 0. The flexible models win where the ledger is worst and lose where it is good. Below are the per-army wins and losses for (d), the trees, against (a), the ledger as it is:
- Wins:
- Drukhari 66 against 46.
- Astra Militarum 75 against 66.
- Death Guard 56 against 47.
- Thousand Sons 54 against 46.
- Space Wolves 54 against 47.
- Losses:
- Tyranids 62 against 73.
- Genestealer Cults 51 against 69.
- Emperor's Children 52 against 57.
Leave one army out, per army (against the consensus)
| Army | Graded | Ceiling | (a) | (a1) | (b) | (b2) | (c) | (d) | (e) |
|---|---|---|---|---|---|---|---|---|---|
| Tyranids | 52 | 88.5% | 73.3% | 73.2% | 64.8% | 69.5% | 57.7% | 62.0% | 57.9% |
| Space Marines | 84 | 91.0% | 60.4% | 60.4% | 62.1% | 58.6% | 63.7% | 62.3% | 69.6% |
| Adeptus Mechanicus | 32 | – | 68.0% | 68.0% | 52.6% | 50.6% | 48.9% | 65.5% | 46.7% |
| Agents of the Imperium | 26 | – | 44.8% | 44.8% | 52.9% | 47.5% | 59.6% | 53.4% | 44.4% |
| Astra Militarum | 61 | – | 66.0% | 66.0% | 64.1% | 64.6% | 66.3% | 74.9% | 66.0% |
| Blood Angels | 19 | – | 54.0% | 54.0% | 71.8% | 53.2% | 67.7% | 54.0% | 43.5% |
| Dark Angels | 93 | – | 54.8% | 54.8% | 52.9% | 54.0% | 53.4% | 58.2% | 68.5% |
| Death Guard | 36 | – | 47.1% | 47.1% | 51.5% | 51.3% | 52.3% | 55.9% | 50.2% |
| Drukhari | 24 | – | 46.1% | 46.1% | 62.6% | 55.2% | 70.9% | 65.7% | 60.9% |
| Emperor's Children | 23 | – | 56.5% | 56.5% | 61.3% | 56.5% | 59.7% | 52.4% | 57.6% |
| Genestealer Cults | 24 | 84.9% | 68.8% | 68.8% | 50.8% | 59.2% | 55.0% | 51.2% | 57.5% |
| Leagues of Votann | 20 | 59.3% | 61.0% | 61.0% | 62.3% | 57.1% | 59.7% | 63.6% | 47.4% |
| Orks | 45 | – | 70.0% | 70.0% | 64.9% | 70.5% | 63.2% | 69.1% | 61.4% |
| Space Wolves | 21 | – | 46.9% | 46.9% | 60.6% | 66.9% | 56.3% | 54.4% | 58.1% |
| Thousand Sons | 28 | – | 45.6% | 45.6% | 61.2% | 53.1% | 58.6% | 54.0% | 55.3% |
| World Eaters | 30 | 89.3% | 61.5% | 61.5% | 61.3% | 54.8% | 62.0% | 63.0% | 55.1% |
The Space Marine family left out together
Why this check: the Dark Angels', Blood Angels' and Space Wolves' pools are their own datasheets plus the Space Marines' (tools/ledger-armies.js). Leaving one of the four out still trains on near-copies of most of its units. Here each of the four is scored with all four left out of the training (the main run's figure in brackets):
| Army | (a) | (b) | (b2) | (c) | (d) | (e) |
|---|---|---|---|---|---|---|
| Space Marines | 60.4% | 55.7% (62.1) | 51.9% (58.6) | 56.6% (63.7) | 56.8% (62.3) | 55.5% (69.6) |
| Blood Angels | 54.0% | 59.7% (71.8) | 58.1% (53.2) | 62.9% (67.7) | 50.0% (54.0) | 58.1% (43.5) |
| Dark Angels | 54.8% | 46.3% (52.9) | 47.1% (54.0) | 47.1% (53.4) | 50.4% (58.2) | 52.4% (68.5) |
| Space Wolves | 46.9% | 62.5% (60.6) | 68.8% (66.9) | 56.3% (56.3) | 48.1% (54.4) | 56.3% (58.1) |
| The four, pooled | 57.1% | 51.1% | 49.9% | 51.8% | 53.2% | 54.0% |
What it shows:
- The neighbours' lead on Space Marines and Dark Angels (69.6% and 68.5%) was those armies' twins being found in each other.
- With the family out, every flexible model is at or near a coin flip on Marines.
- The main leave-one-army-out table therefore flatters the flexible models on these four armies. Without them, the ledger's edge is a little larger than the table shows.
Units held out within armies (5 repeats × 5 folds, stratified by army)
All 16 armies are trained together on four-fifths of each army's units. The scores are on the pairs among the fifth left out of the same fit.
| Army | (a) | (a1) | (b) | (b2) | (c) | (d) | (e) |
|---|---|---|---|---|---|---|---|
| Tyranids | 73.2% | 72.9% | 67.8% | 68.7% | 68.0% | 66.9% | 62.9% |
| Space Marines | 61.1% | 60.9% | 61.5% | 58.3% | 69.6% | 65.9% | 70.4% |
| Adeptus Mechanicus | 70.1% | 69.8% | 52.3% | 56.3% | 52.1% | 64.7% | 55.1% |
| Agents of the Imperium | 43.3% | 43.4% | 51.1% | 51.9% | 58.9% | 44.2% | 47.5% |
| Astra Militarum | 66.5% | 66.4% | 71.5% | 74.8% | 78.9% | 77.5% | 74.9% |
| Blood Angels | 50.3% | 50.3% | 68.8% | 68.6% | 66.1% | 63.4% | 65.5% |
| Dark Angels | 54.9% | 54.8% | 59.0% | 56.7% | 66.1% | 65.2% | 67.0% |
| Death Guard | 46.8% | 46.6% | 52.8% | 52.2% | 58.0% | 57.5% | 54.5% |
| Drukhari | 49.1% | 49.1% | 63.1% | 49.7% | 70.8% | 69.2% | 58.1% |
| Emperor's Children | 53.1% | 52.3% | 60.5% | 54.4% | 66.5% | 56.1% | 58.3% |
| Genestealer Cults | 69.7% | 69.7% | 62.2% | 66.2% | 58.5% | 56.4% | 56.2% |
| Leagues of Votann | 65.4% | 65.4% | 61.3% | 56.6% | 63.6% | 57.8% | 52.0% |
| Orks | 67.9% | 67.7% | 68.0% | 69.9% | 66.1% | 69.4% | 57.6% |
| Space Wolves | 49.8% | 49.8% | 65.5% | 67.8% | 61.8% | 51.7% | 40.1% |
| Thousand Sons | 46.5% | 46.5% | 62.8% | 53.8% | 68.6% | 61.3% | 60.3% |
| World Eaters | 61.9% | 62.2% | 63.3% | 60.3% | 59.0% | 66.5% | 58.3% |
| Pooled | 60.3% | 60.1% | 62.2% | 60.9% | 67.4% | 65.9% | 65.4% |
What it shows:
- Here the flexible models do beat the ledger: (c) by 7 points pooled.
- They have seen the army, so they learn its reviewer's taste: Drukhari +22, Thousand Sons +22, Blood Angels +16, Emperor's Children +13, Dark Angels +11, Astra Militarum +12.
- Where the ledger already reads the reviewer well, they lose: Tyranids −5, Genestealer Cults −11, Adeptus Mechanicus −18.
- Set beside the left-out table, this is the main finding: the signal the flexible models find is army-specific. It doesn't carry to an army they haven't seen.
- One caution: the Marine armies' twins are in training here too (a unit's Dark Angels copy teaches its Space Marines copy), so some of the gain on those four is the same memory.
Does pooling help Tyranids?
Held-out pairwise on Tyranid units against their consensus, with the same folds in each column:
| Model | Tyranids alone (their own training folds) | From the other 15 armies only | The other 15 plus Tyranids' own training folds |
|---|---|---|---|
| (a) the ledger as it is | 73.2% | 73.3% | 73.2% |
| (a1) re-fitted | 71.5% | 73.2% | 72.9% |
| (b) linear, all variables | 67.2% | 64.8% | 67.8% |
| (b2) core linear | 66.9% | 69.5% | 68.7% |
| (c) interactions | 65.4% | 57.7% | 68.0% |
| (d) trees | 66.0% | 62.0% | 66.9% |
| (e) neighbours | 61.6% | 57.9% | 62.9% |
What it shows:
- Other armies' units add 0.6 to 2.6 points to a flexible model that also has Tyranids' own units. On their own, they teach less about Tyranids than Tyranids' own 40 training units do.
- Nothing reaches the ledger's 73%, let alone the 88.5% ceiling.
- A caveat on the first column: it is this tool's pooled pipeline restricted to Tyranids. It runs 5 repeats with an inner 3-fold, against T-632's 10 repeats with an inner 4-fold. It also takes its logs from the pooled table's skew and stops the fits after 200 steps. Its (c) lands at 65.4% where T-632's lands at 73.3%. The columns compare with each other, not with T-632's table: see doubts.
The learning curve over the number of training armies
Setup:
- Each model is trained at its most-chosen leave-one-army-out setting (no inner tuning) on m random armies other than the one scored.
- There is one random draw per left-out army per size, so 16 fits per point.
- The figures are pooled / army mean.
- The 15 column is the tuned leave-one-army-out run.
| Model | 3 | 6 | 9 | 12 | 15 |
|---|---|---|---|---|---|
| (d) trees | 59.5% / 58.1% | 61.4% / 58.8% | 55.1% / 56.5% | 65.5% / 59.4% | 61.8% / 60.0% |
| (e) neighbours | 57.1% / 54.1% | 59.8% / 55.9% | 56.5% / 55.3% | 65.5% / 58.2% | 63.4% / 56.3% |
| (b2) core linear | 52.4% / 53.3% | 59.7% / 58.3% | 56.6% / 57.3% | 58.0% / 57.1% | 58.6% / 57.7% |
| (a1) re-fitted ledger | 59.9% / 57.8% | 59.9% / 57.8% | 59.9% / 57.8% | 59.9% / 57.8% | 59.9% / 57.8% |
| (a) the ledger | 60.0% / 57.8% | (no fit) | 60.0% / 57.8% |
What it shows:
- On the army mean, the trees rise about 2 points from 3 to 15 armies. The pooled figure jumps about with which large army is drawn.
- Nothing like T-632's 6 to 8 point climb from 16 to 41 Tyranid units shows here. The units of other armies aren't "more of the same" for a new army.
The why
What the pooled models key on (permutation importance on the army left out)
How it's measured: each army is scored by the fit on the other 15. A variable, or a group of variables, is shuffled among that army's units. The drop is in pooled pairwise, in points (each army weighted by its pairs). "Armies" counts how many armies drop by more than half a point.
| Family | Trees (d): drop | armies | Neighbours (e): drop | armies |
|---|---|---|---|---|
| Ledger lines | 4.8 | 13 | 2.1 | 10 |
| Weapons | 1.7 | 8 | 3.5 | 9 |
| Kill by class | 1.4 | 10 | 0.6 | 7 |
| Datasheet | 1.3 | 10 | 2.4 | 11 |
| Rule flags | 0.7 | 13 | 5.1 | 10 |
| Interactions | 0.3 | 8 | 0.6 | 7 |
Reading the two models:
- The trees lean on the ledger's own lines, chiefly Kill and Score, and on points per model.
- The neighbours lean on flags and weapons. That is how they find twins: the same keywords and profiles.
- Trees, top variables: points per model 2.0 (11 armies), the Kill line 1.1 (13 armies), the Score line 0.6, melee damage into chaff 0.6, Kill into monsters 0.6, Lone Operative 0.5, the charge's lands 0.5, melee Kill 0.4, Kill into light vehicles 0.4, Devastating Wounds 0.4.
T-632's candidate mechanics, checked across armies
Each candidate is its group of variables (listed in tools/diagnose.js's CANDIDATES), shuffled together on the army left out. Beside that is each variable alone: the share of within-army pairs it orders right, and the armies where its sign is + or − (Spearman beyond ±0.1).
| # | Candidate | Trees: drop (armies) | Neighbours: drop (armies) | Alone, the best of its variables | Verdict |
|---|---|---|---|---|---|
| 2 | Scarce answers (Kill into monsters and vehicles, anti-tank guns) | 2.0 (7) | 2.2 (10) | Kill into light vehicles 56.3% (+8/−2); ranged anti-vehicle 55.5% (+7/−6); monsters 54.3% (+6/−3) | Backed: the strongest across armies |
| 3 | Objective bodies (Score, OC per point, Battleline) | 0.8 (8) | 0.5 (8) | bodies × OC × Move 56.7% (+9/−5); OC 56.0% (+8/−4) | Backed, weakly. The trees use it as a floor: the bottom of an army's Score loses about half a letter |
| 1 | Threat on arrival | 0.5 (7) | 0.5 (5) | arrives × melee Kill 54.3% (+9/−2); arrives 54.1% (+9/−2) | Weak. It points the right way in 9 armies but carries little. T-633's arrival Kill was reverted |
| 8 | Value at the cheapest size | −0.2 (2) | 1.0 (9) | 59.6% (+10/−1) | Mostly the same signal as Value itself. Strong alone, nothing extra once Value is in |
| 7 | Cheap reach (actions, indirect fire) | 0.2 (6) | 0.4 (10) | points at the cheapest size 55.7% (+10/−1); actions line 47.2% (+4/−7) | Not backed as actions; the actions line points down in 7 armies |
| 4 | Character sniping (Precision) | 0.0 (3) | 0.2 (6) | Kill into characters 54.6%; Precision 49.7% | Not backed across armies (a Tyranid signal) |
| 6 | Psychic | 0.1 (2) | −0.1 (4) | psyker 52.1% (+7/−3) | Not backed. Psychic weapons are among the over-rated units' traits across armies |
| 5 | Transports | −0.2 (0) | 0.2 (6) | transport 48.8% (+1/−6) | Negative, as in T-632: the ledger over-credits transports' own stats. A transport's carrying is what's missing |
| 9 | Melee Soak | 0.0 (0) | 0.1 (6) | 49.0% (+2/−8) | Negative in 8 armies. The T-632 sign holds: melee staying power is over-credited |
Ranked for the next ledger steps:
- Scarce answers (2).
- A capped Score (3).
- Arrival (1) only as part of a bigger change.
- Transports (5) as their carrying, and melee Soak (9) counted down. Both point the same way in most armies, but as corrections, not as new value.
Precision, psychic and cheap reach were Tyranid-specific.
The steadiest single signals (alone, within army)
| Variable | Pairs right | Armies + / − |
|---|---|---|
| The Kill line (c_kill) | 60.1% | 11 / 0 |
| Value at step 6 | 60.0% | 10 / 2 |
| Value at the cheapest size | 59.6% | 10 / 1 |
| Kill into heavy infantry | 58.8% | 10 / 0 |
| Kill into cavalry and beasts | 58.6% | 11 / 1 |
| √(Kill × Soak) | 58.1% | 8 / 4 |
| The Hold line | 57.1% | 8 / 3 |
What it shows:
- No single variable reaches the ledger's own Value by much.
- Many variables flip sign between armies. Ranged damage per point, for one, is + in 7 armies and − in 7.
- That flipping is why the flexible models don't travel: each army's reviewer rewards a different mix.
The trees on all 16 armies (depth 3, 100 trees; thresholds in SDs within the army)
The splits that carry the most gain:
- The Kill line (8.5% of the gain), at about the army's median. The partial dependence rises +0.3 letter there: above-median Kill is a step, not a slope.
- Points per model (7.7%), above about +0.6 SD (+0.2 letter): expensive models are rated higher than their dice alone say.
- The Score line (7.6%), at −1.2 SD: the bottom tail loses about half a letter, and beyond it Score barely matters. This is the "saturating" shape of T-632's candidate 3, seen from below.
- Melee Kill (5.6%), at −1.4 SD: having almost none costs about half a letter.
- Fly (−0.4 letter where it's rare in the army), Infantry (+0.4) and Pistols (−0.3). These keyword splits differ by army and are the likeliest to be noise.
The interactions (one split under the other):
- Kill × Score (4.2% of the gain).
- Kill × points per model (3.7%).
- Kill × melee Kill (2.0%).
- Score × Kill into chaff.
- Arrival × melee Kill × Fly.
So it is the ledger's Kill read differently by role, the same shape tools/roles.js tried with its fixed roles.
The units every flexible model misses the same way, leaving their army out
Each model's scores are mapped onto the army's letters by rank. The figure is the mean over the five flexible models minus the consensus, and all five miss in the same direction. The ledger's error is beside it.
| Rated higher by the reviewers | Army | Consensus | Models | Ledger | Rated lower by the reviewers | Army | Consensus | Models | Ledger | |
|---|---|---|---|---|---|---|---|---|---|---|
| Stormraven Gunship | Dark Angels | A | −4.0 | −4 | Centurion Assault Squad | Dark Angels | F | +3.6 | +4 | |
| Miasmic Malignifier | Death Guard | S | −4.0 | −4 | Tzaangor Enlightened | Thousand Sons | D | +3.4 | +4 | |
| Biovores | Tyranids | S | −3.9 | −0.6 | Serberys Sulphurhounds | Adeptus Mechanicus | F | +3.0 | +4 | |
| Suppressor Squad | Dark Angels | S | −3.8 | −3 | Company Heroes | Dark Angels | F | +3.0 | +3 | |
| Librarian in Phobos Armour | Dark Angels | S | −3.6 | −4 | Lazarus | Dark Angels | F | +3.0 | +3 | |
| Librarian | Dark Angels | S | −3.6 | −5 | Deathwing Terminator Squad | Dark Angels | D | +3.0 | 0 | |
| Sorcerer | Thousand Sons | S | −3.6 | −4 | Maulerfiend | Thousand Sons | C | +3.0 | +3 | |
| Jakhals | World Eaters | S | −3.5 | −2.3 | Heavy Intercessor Squad | Dark Angels | D | +2.8 | +3 | |
| Venom | Drukhari | S | −3.4 | −5 | Grey Knights Terminator Squad | Agents of the Imperium | C | +2.6 | +3 | |
| Infernal Master | Thousand Sons | S | −3.4 | −4 | Primaris Psyker | Astra Militarum | C | +2.6 | +1 |
What the misses share:
- The under-rated 30, on average, sit low in their army on Kill, Value, durability against anti-tank, Hold and points. They are cheap support pieces: leaders and psykers whose worth is in the squad they join, transports that carry, and an aura piece.
- The over-rated 30 sit high on melee damage (into vehicles and elite infantry), psychic weapons and Value. They are units whose dice look good but whom the reviewers don't take.
- Each army has its own handful, from 1/3 in Blood Angels to 19/16 in Space Marines (higher / lower). No army is free of them.
- Biovores are the one Tyranid in the top 10. Jordan's hand dial keeps the ledger close (−0.6 against −3.9 for the models).
What this means
- Data is not the bind across armies; the reviewers' differences are. T-632 read its rising Tyranid curve as "more data would help". Pooled, the extra data comes from other armies with other reviewers and other tastes, and it doesn't help a new army. The ceiling the ledger can reach on an army it was never tuned on looks like about 60% pooled: what Kill and Value give.
- The ledger is not the problem on held-out armies. Its form, re-fitted on 15 armies, comes back as step 6. Every flexible model is within noise of it, and below it once the Marine twins are removed.
- The next measurable gains are the shared misses, not a better fit:
- a leader's and a transport's worth to another unit (the joined-character and carrying measurements);
- scarce anti-elite and anti-tank answers (T-632's candidate 2);
- a Score floor rather than a slope.
Each would be one step, judged on lists and on units held out (--on).
- Within-army learning is real but is fitting a reviewer. A model trained on half an army's letters orders the other half 7 points better than the ledger in the armies the ledger reads worst. That is a per-army calibration of one reviewer's taste, not a rule of the game. It shouldn't move the ledger.
Method notes
- The variables: T-632's 144 per unit (tools/diagnose.js's DICT and rowOf, unchanged).
- Each army is at its best solo form at the step-6 baseline, with the cheapest size's columns, read through the same frozen inputs (BSData 374f505, inputs
30f0056e6103). - A chapter's datasheets and effects take its parent's as well, as tools/ledger-armies.js does.
- The Tyranid rows match T-632's table exactly (7,488 cells checked).
- Each army is at its best solo form at the step-6 baseline, with the cheapest size's columns, read through the same frozen inputs (BSData 374f505, inputs
- Per-army scales:
- Each column takes log(1 + x) where it is non-negative and skewed (skew above 1 over the pooled table).
- It is then standardised within its army (mean 0, SD 1 over all the army's ranked units, graded or not; a column constant in an army is 0), as tools/roles.js did. No letter is read in this step.
- (a1) reads its seven lines at weight 1, each army's divided by the SD of its step-6 Value.
- (e) averages its neighbours' consensus standardised within their own army, over training units only.
- The loss: pairwise logistic over within-army pairs, each pair weighted by its consensus gap. Each army's weights are scaled to total 25 × its training units, so an army counts by its units, not its pairs (n²).
- (c) is the kernel (1 + x·z/p)² with ridge.
- (d) is Newton-leaf boosting with the dictionary's monotone signs and at least 5 units a leaf. Column subsampling was dropped: on all 16 armies it changed the in-sample fit by under a point and doubled the time.
- The grids are T-632's, trimmed: lasso from 0.05, kernel down to 0.0003, k up to 30, 200 steps.
- Tuning:
- Leaving one army out: an inner cross-validation over 4 groups of the 15 training armies (pairs within the inner left-out armies only).
- Units held out: an inner 3-fold over the training units, stratified by army.
- No setting ever sees the scored army.
- Scrambled: each army's consensus and letters permuted together among its graded units, 2 permutations through the whole leave-one-army-out run.
- Runtime: 1,057 s (17.6 minutes) on 4 worker threads.
- Leave one army out: 4 minutes.
- The Marine family: 0.6 minutes.
- Scrambles: 7.6 minutes.
- Units held out, with Tyranids alone: 3.4 minutes.
- Curve and importance: 1.8 minutes.
- Tests: tests/unit/diagnose.test.js has a smoke test of the new pure functions: within-army pairs and weights, within-army scoring and folds, per-army standardisation, and rank mapping.
- Doubts:
- Units held out runs 5 repeats, not T-632's 10, to keep the full run under 20 minutes. The SEs are 0.2 to 0.7 points;
--unit-repeats 10doubles that part. - The Tyranid-only figures in this pipeline are below T-632's for the flexible models: (c) 65.4% against 73.3%, (b) 67.2% against 69.7%. The pipelines differ in folds, inner folds, repeats, log choice, step count and grids, and T-632's (c) had a scrambled spread of 6.6 points. So the difference is within what the fold choice alone can move. The comparison across columns here is like for like.
- 11 of the 16 armies have one reviewer each, so their "consensus" is one person's letters, and the ceiling is unknown there. Those armies' per-army figures move several points with a few units.
- The learning curve's points are one random draw of training armies per left-out army.
- Units held out runs 5 repeats, not T-632's 10, to keep the full run under 20 minutes. The SEs are 0.2 to 0.7 points;