A methods review for Jordan, read-only. Built on the brief, the earlier critique (not repeated here), round 9's report, design/scoring.md section 6 and design/decisions.md. Numbers are from reports/tiers/2026-10-02-tyranids-round9.json and 2026-10-02-round8-sliders.json (and only the mi list counts from docs/data/rankings.json, to rebuild the list-share pool of 1,105 units). Nothing here changes a score.
How the numbers were made. Round 9 changes only Anvil, so round-9 jobs are round 8's percentiles with round 9's Anvil (Tyranids: their own round-9 jobs). This rebuild reproduces the Tyranid ρ (0.493), the Ork ρ (−0.118) and every unit's list-share ρ (0.077; round 8 0.162) exactly. Not the Marines: 27 units whose best detachment changed differ by over 2 points, so the rebuilt Marine baseline is −0.245 (the run's −0.099); compare within that column only. SEs: about ±0.13 Tyranids, ±0.19 Marines, ±0.15 Orks, ±0.03 every unit's list share (n = 1,105), the only yardstick that separates options closer than 0.1.
1. Do the letters matter? Rank first, letters on top
Short answer: yes, score by position and add the letters afterwards as a display layer. The yardsticks already work that way, and the half that doesn't is the noisy half.
- Every ρ is computed on the continuous score: moving the S cut ±5 re-letters 401 units and leaves the Tyranid ρ at exactly 0.494. Cuts are presentation.
- The letter-agreement counts (strict, raw, within one) depend on where the cuts fall, and they lose information. On Auspex, ρ against our continuous score is 0.493; against our all-armies letters 0.456; against the proposed-cut letters 0.402. On Marines the all-armies letters happen to score +0.04 against the score's −0.10. That is the coarse ruler adding noise, not signal.
- Both yardsticks are within-army quantities. An expert tier list ranks one codex, and list share is "of the lists that could take it". So neither can validate the all-armies ruler. The one cross-army check we have, S share per army against top-quarter lists, is negative (−0.39 at the current cuts, −0.24 at the proposed ones). Nothing yet shows that our letters are comparable across armies.
- Ranking inside each army costs nothing on the panels (Auspex 0.501 within the army against 0.493) and lifts every unit's list-share ρ from 0.077 to 0.120, because it takes out army-level offsets that list share can't see.
On rank, the yardsticks gain stable, cut-free numbers with honest SEs, and lose the same-letter counts, which visitors understand; keep those as a secondary line on frozen quotas. The product loses an absolute S across armies; keep it as the default view if Jordan wants, labelled an untested claim.
Recommendation. Score = the pool percentile of the composite (0 to 100, uniform by construction). Publish the overall rank and the within-army rank. Letters = cuts on the overall score at frozen shares (S 8%, A 15%, B 27%, C 30%, D 20%, the round-9 proposal), set once, so meta drift shows as letters moving. Validate on within-army rank.
2. Are we scoring too high?
Short answer: yes. 36% of all units have a job at 90 or more and 21% score 95 or more. Almost all of it is arithmetic: the best of four uniform percentiles. Correlated jobs and the percentile transform add almost nothing, and the steps add the last bit.
Counts of units per 10-point bin, all 1,161 units:
| Bin | 0s | 10s | 20s | 30s | 40s | 50s | 60s | 70s | 80s | 90s | 100+ |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Hammer | 117 | 118 | 116 | 121 | 116 | 115 | 115 | 111 | 116 | 116 | – |
| Anvil | 116 | 117 | 117 | 111 | 110 | 106 | 105 | 131 | 125 | 123 | – |
| Runner | 128 | 104 | 117 | 113 | 117 | 117 | 116 | 118 | 114 | 117 | – |
| Banner | 660 | 49 | 26 | 42 | 43 | 27 | 31 | 51 | 116 | 116 | – |
| Best allowed job | 3 | 1 | 4 | 16 | 35 | 55 | 100 | 195 | 333 | 419 | – |
| Score (best + steps) | 3 | 1 | 4 | 16 | 35 | 54 | 99 | 191 | 313 | 401 | 44 |
| best^k, rarity-corrected (a plain re-rank gives 116 a bin) | 68 | 93 | 99 | 125 | 135 | 125 | 148 | 129 | 118 | 121 | – |
- Any allowed job ≥ 90: 419 units (36.1%). ≥ 95: 220 (18.9%). Score ≥ 95: 247 (21%), 44 of them above 100. The jobs are flat by construction: each puts 10% of units in its top bin, apart from Banner, where 648 units lend nothing and the "smooth start" puts 232 of the 513 lenders in 80 to 100.
- The order statistic explains it. If the four jobs were independent, the best of k uniform percentiles would put 387 units at 90+ and 208 at 95+ (1 − 0.9^k and 1 − 0.95^k, summed over each unit's allowed jobs). Shuffling each job's column independently, which keeps the real marginals including Banner's zeros, gives 392 and 211. The real figures (419, 220) are within 7% of that.
- Correlation is a small net push. Spearman over all units:
| Hammer | Anvil | Runner | Banner | Points | |
|---|---|---|---|---|---|
| Hammer | 1 | 0.61 | −0.34 | −0.54 | 0.58 |
| Anvil | 1 | −0.33 | −0.44 | 0.44 (round 8: 0.02) | |
| Runner | 1 | 0.17 | −0.69 | ||
| Banner | 1 | −0.27 |
Hammer and Anvil move together, which lowers the maximum. Runner and Banner run against both, which raises it. The two nearly cancel.
- Three jobs largely track cost: Hammer +0.58 (absolute survival), Runner −0.69 (the points gate), round 9's Anvil +0.44. The composite is cost-neutral (0.02) only because they cancel, which is fragile.
- What spreads units along 0 to 100. A monotone transform of the composite keeps every ρ, so the fix is free. Either re-rank the composite over the pool, or use best^k (the chance that k uniform draws all fall below your best). The re-rank is flat (116 a bin); best^k gives 68 to 148 a bin, with 121 units at 90+ and 57 at 95+. Use the plain re-rank: the per-unit k version lifts units with fewer allowed jobs (Support characters) and drops Auspex to 0.437. z-scores or log-ratios to the median belong one layer down, on the raw quantities (spec below), where they keep real gaps ("twice the median") that a percentile throws away. On the final score they would re-create a pile-up, because the composite is a maximum and its distribution is skewed.
3. Raw first, then weighted, then curved: the four-layer architecture
What it fixes. Meta knobs turn without rerunning the engine, so the Lab's sliders can cover every knob; sweeps become cheap and exact; "traceable to a named number" becomes literal (raw × weight × curve × ceiling); lenses become config presets, not code.
What it risks.
- Knob count. 39 swept constants today, plus about 11 meta weights and 4 ceilings. Every knob carries a source (measured, fitted on a named window, or Jordan's dated judgement), and none is set by maximising a panel's ρ: decisions.md's Auspex rule, extended to config.
- Circularity. Layer 2 is fitted on winning lists, and list share is the main yardstick. A time split is required: fit on one window, validate on the next.
- The "never changes" layer moves. The sweep's biggest movers are geometry (the Advance 288 letters, the objectives' distances 286, reserves 206). Those are inputs to layer 1, not meta. Put them in a versioned scenario block: changing it reruns the engine, and it is never a slider.
- Best-context selection is circular today. The best detachment, loadout and way in are chosen by the composite, so a ceiling change can change which raw row a unit uses. Layer 1 must keep a row per context, and layer 4 picks the best context after composing.
How far the tool is. Halfway. The raws exist inside tiertest.js, and rankings.json publishes the live Offense/Defense ones. But meta weights are applied before the percentile, constants are globals read mid-run, and context selection uses the composite. The round-8 sliders file is a working layer 4. The unlocking step: a raw dump plus a pure compose(raw, config) that reproduces round 9 exactly.
The spec (one page)
Layer 1: the raw table (build/tier-raw.json; one row per unit × detachment × delivery × loadout context; derived numbers only, never datasheet text; regenerated when the BSData commit, engine version or scenario hash changes)
| Column | Meaning |
|---|---|
id, army, unit, ctx{det, way, kit}, pts, models | identity and context |
allow{H,A,R,B} | eligibility flags (OC 0, joins as Support, one use, immobile) |
dmg[GEQ,MEQ,TEQ,LV,HV] | enemy points destroyed per 100 pts per turn in range |
inRange[1..5] | the delivery profile: chance of being in range each round |
htd[6 attackers] | absolute hits to destroy (heals included) |
ocTurns | OC × min(turns alive, 5), before the meta mix (per attacker, so layer 2 can mix) |
hold | position share: midfield and enemy-half objectives held |
reach[mid,far,deep] | share of the game it can reach each band |
ld, belowHalf | Leadership pass chance, share of the game below half strength |
perUnit100, footprint | units per 100 pts, footprint |
lend{leader,aura,transport,debuff,shock,cp,spawn} | points of value lent per game, by kind (CP and Battle-shock in counts, priced in layer 2) |
Layers 2 to 4: the config (data/tierconfig.json; each knob has a default, a source and a date)
| Knob | Default | Source |
|---|---|---|
scenario (gap 30", Advance 3.5", objectives 15/21/27", DS 9", reserves 18", rounds 5) | as today | 11th-edition rules + builder; versioned; layer 1 input, not a slider |
target.w[5] | 10/6/38/22/23 | fitted on winning lists' points, window W1 |
fire.w[6] | 8/20/37/7/22/6 | fitted on winning lists' weapons, window W1 |
price.cp, price.shock, price.debuff | 30 pts, 10% of median, 0.5 | Jordan, 2 Oct |
hammer.peak | 0.75 | builder (step g); applied on the log-ratios |
anvil | ocTurns per 100 pts × hold × Battle-shock | see section 5 |
runner.parts | reach, turns alive, Ld, 0.5 × (units, footprint) | builder |
curve | per raw quantity ℓ = log2(x ÷ pool median), floor −4; job raw = Σ w·ℓ; job score = pool percentile | method |
job.ceiling{H,A,R,B} | 1, 1, 1, 1 | meta lens; set from a named source, never from a panel |
compose | best allowed job × ceiling, + steps (+5 per other job ≥ 90, +2 at 80 to 89) | Jordan's rule |
final | pool percentile of the composite | method |
letters | frozen cuts at S 8 / A 15 / B 27 / C 30 / D 20% of the calibration run | Jordan to approve once |
pool | non-Legends, fieldable in 2,000 pts | method |
The curve. Log-ratio to the pool median puts every raw quantity on one scale ("one step = twice the median") and turns a weighted product into a weighted sum, so a weight means the same thing whatever the units. The job is then ranked, because the order is what the panels test and a rank survives outliers (Titans). The composite is ranked again at the end, which removes the best-of-four pile-up (section 2).
Lenses. A lens is a named partial config: target.w, fire.w, job.ceiling and price.* only, never scenario or letters. For example "Armour meta": HV 35%; "Objective missions": Anvil and Runner ceilings 1 with Hammer 0.9. The client applies it to the published layer-3 job raws (our own numbers) and shows "under the Armour lens: A", always beside the default letter, never replacing it. Because lenses are our software over our own derived numbers, they can be a supporter perk under the house rules. The default letter and the raw table stay free.
4. Weighted sum or best job
Option comparison on the rebuilt round-9 jobs (ρ; every unit's list share over 1,105 units; the Marine column is a rebuild, compare within it):
| Option | Auspex | Marines | Orks | List share, every unit | Tyranid list share | One sentence? |
|---|---|---|---|---|---|---|
| (c) Best job + steps (round 9) | 0.493 | −0.245 | −0.117 | 0.081 | 0.453 | "Best job, +5 per other top-10% job, +2 per top-20%" |
| Best job only | 0.488 | −0.122 | −0.112 | 0.084 | 0.454 | yes |
| (a) Sum, 25% each (disallowed = 0) | 0.507 | 0.008 | −0.073 | 0.149 | 0.394 | "Your four jobs averaged" |
| (a) Sum, 40/30/20/10 | 0.440 | −0.221 | 0.153 | 0.115 | 0.232 | yes |
| (a) Sum, Hammer 90 / Anvil 5 / Runner 3 / Banner 2 | 0.300 | −0.368 | 0.344 | 0.041 | 0.084 | yes |
| (b) Ceilings H100 A90 R90 B90 | 0.488 | −0.312 | 0.018 | 0.097 | 0.433 | "Best job, each marked out of what the meta pays" |
| (b) Ceiling on Anvil only, ×0.8 | 0.531 | −0.125 | −0.174 | 0.129 | 0.537 | same |
| (b) Anvil ×0.8, with steps | 0.531 | −0.164 | −0.181 | 0.124 | 0.537 | same |
| Soft maximum, τ = 10, Anvil ×0.8 | 0.515 | −0.138 | −0.145 | 0.120 | 0.540 | "Mostly your best job, a little for the others" |
| Soft maximum, τ = 40, Anvil ×0.8 | 0.547 | −0.014 | −0.052 | 0.140 | 0.504 | hard |
| max(sum's percentile, best job) | 0.471 | −0.144 | −0.128 | 0.097 | 0.450 | "Specialist or all-rounder, whichever is better" |
| 0.75 best + 0.25 mean (dropped; reference) | 0.495 | −0.079 | −0.125 | 0.095 | 0.452 | – |
Reading it.
- The panels can't choose. Every reasonable rule is within one SE on Auspex, and in a 6⁴ grid of ceilings the best for Orks (0.27) gives Auspex 0.28. Only list share can separate options.
- On list share, the sum wins (0.149 against 0.081, about 2 SE), but it buries the specialists Jordan protects: Biovores (22 of 23 lists, Auspex S) fall to the 17th percentile, against the 70th under steps. It also lifts generalists Auspex rates S (Zoanthropes 15 → 45, Hyperadapted Raveners 11 → 54). Experts and lists contain both kinds, which is why no single rule wins everywhere.
- Jordan's "Hammer 90%, Anvil 5%" fails as a weighted sum (Auspex 0.30, list share 0.04): a sum with lopsided weights turns the score into Hammer alone. As a ceiling it works. It caps how high a pure Anvil can go without stopping a 99-Hammer unit being S, and it keeps the specialist rule.
- The one knob that helps everything but Orks is the Anvil ceiling. Lowering Anvil alone from 1 to 0.6 lifts Auspex 0.49 → 0.60, Tyranid share 0.45 → 0.62 and every unit's share 0.081 → 0.130. That is monotone, not a lucky point. The same sweep on Runner or Banner makes things worse. This is a symptom: Anvil is mismeasured (section 5). The ceiling demotes the Hierophant and the Tervigon (0 lists), but it also drops Hormagaunts (Auspex S, 13 of 23 lists) from the 70th to the 33rd percentile.
- The combining rule is second order: best only, steps and soft maximum at τ ≤ 10 agree within 0.02 once Anvil is capped.
Reconciliation. Meta importance goes in as per-job ceilings in the 0.7 to 1 range, not as weights in a sum. The rule stays "best allowed job (× its ceiling), plus steps". One sentence for a visitor: "A unit's score is its best job, with each job marked out of what the current meta pays for it." Ceilings default to 1. Set Anvil's from a named source only after section 5's fix, and expect it to be close to 1 if the fix works.
5. The round-9 list-share drop
Short answer: it comes from the division. More precisely, removing one division by points multiplied Anvil by the unit's cost, which made expensive units into Anvils. The gate didn't cause it and the cuts can't have: list-share ρ uses the continuous score.
- The biggest rises are the units no list takes. Warlord Titan 19.8 → 99.7 (score 24 → 99.7, D → S); Warbringer 17.4 → 97.8; Reaver 18.4 → 95.1; Warhound 34.3 → 96.3; the Phantom and Revenant Titans; the Hierophant 36 → 93.5; the Ta'unar 30.7 → 87.3. By cost band, Anvil moved −12.6 (≤ 60 pts), −8.2 (61 to 90), −0.2 (91 to 130), +6.0 (131 to 200), +24.4 (201 to 400) and +43.3 (over 400). Scores followed: +28 over 400 pts. The 408 zero-share units gained +1.8 on average, and the 100 most-taken lost −1.5.
- The biggest falls are cheap, often-taken bodies: Custodes Prosecutors 93.7 → 57.7 (share 0.63, A → C), Tzaangor Enlightened 91.6 → 55.2 (0.45), Von Ryan's Leapers 79 → 44.
- Rises or falls? Keep round 9's Anvil but never let it rise above round 8's (the minimum of the two): list share 0.154, nearly all of round 8's 0.162. Take the maximum instead: 0.092, nearly all of round 9's 0.077. The drop is the rises. The gate can only lower Anvil, so it can't cause rises, and its sweep (one-turn gate: 0 letters; two-turn gate: 28) shows it barely touches anything. A cap on letters can't change a ρ on the score.
- Round 9's Anvil is a points measure: its ρ with points went 0.02 → 0.44, and dropping Anvil altogether (0.130) beats round 9 on list share.
Recommended form: Anvil = OC × min(turns alive, 5) ÷ points × Battle-shock × position (× revive), with no gate. It reads "objective-control turns bought per point". It divides once, as the critique wanted, but survival stops counting at the game's length, so a 3,500-pt Titan can't buy its way up. It replaces the gate: Tu below 1 already scales the value down, smoothly rather than in steps. On the Tyranid raws it is cost-neutral (ρ with points −0.00, against round 8's −0.10 and round 9's +0.42), and it ranks the gaunts, Gargoyles and Warriors first and the Hierophant nowhere. Turns alive should use the meta-mixed fire as Hammer's does, so it stays absolute (Jordan's round-5 rule).
6. What else to change first
Ranked by expected effect on both the panels and list share. None of these is in the critique or decisions.md.
- Fix Anvil (section 5). It is the single biggest lever on both yardsticks, as the Anvil-ceiling sweep shows.
- Change the list-share headline. The pooled ρ mixes 27 armies and 408 tied zero-share units. Report instead (i) the median of per-army ρ (round 8: 0.188; round 9: 0.113), (ii) the AUC for taken against never-taken (0.585 → 0.536), and (iii) ρ among taken units only. And track an overfit gauge: Tyranids' per-army ρ is 0.45, second of 27 armies, while the median is 0.11. The tuned army leads because it is tuned. When the Tyranid lead shrinks while the median rises, the method is generalising.
- Limit the pool to fieldable units. Eight units cost more than a 2,000-pt list, and both Titanicus factions are 100% S at the current cuts. They set everyone's percentiles and head the S-per-army table. Drop them before taking percentiles.
- Percentiles within weight class instead of a points gate. Runner's hand-set gate is the most sensitive constant (399 letters), and three jobs track cost. Ranking each job within the King of the Hill classes (≤ 75, 76 to 130, 131+ pts) makes cost-neutrality part of the method, not of tuned gates.
- A final rank transform plus frozen quotas (section 1). It changes no ρ, but it removes the visible pile-up and makes the S-per-army check meaningful.
- Layer-1 dump and pure compose (section 3). No yardstick moves, but every later experiment becomes a config diff, and the context-selection circularity goes.
Recommended next experiments (ranked)
- Anvil = OC × min(Tu, 5) ÷ pts × Battle-shock × position, no gate. Test: change
hold9, rerun--r9json. Worked if every unit's list share ≥ 0.15, no unit over 1,000 pts at S, Auspex ≥ 0.47, and Anvil's ρ with points within ±0.15. - Per-army median ρ, AUC and ρ among taken units as the list-share headline, with the Tyranid overfit gauge. Test: recompute over the round-9 history. Worked if the versions' order differs from the pooled ρ's (the pooled number has been misleading us); adopt either way.
- Pool limited to fieldable units. Test: a filter before percentiles. Worked if panel ρ moves under 0.01 and S-per-army ρ rises above −0.24.
- Final rank transform + frozen 8/15/27/30/20 cuts, panels on within-army rank. Test: re-letter round 9. Worked if S is 8 ± 1% and S-per-army ρ ≥ 0.
- Job percentiles within weight class, Runner's points gate removed. Test: re-rank round 9's jobs by class. Worked if Auspex ≥ 0.47 and every unit's share doesn't fall, with one hand-set constant gone.
- Anvil ceiling 0.8, only if 1 doesn't recover list share. Test: multiply in compose. Worked if every unit's share rises 0.02 or more, the per-army median rises, Auspex doesn't fall more than 0.05, and Hormagaunts stay B or better.
- Raw dump +
compose(raw, config). Test:--rawdumpand a reproduction check. Worked if 0 units differ from round 9 and a full sweep runs in under a minute. - Composite-rule bake-off (best, steps, soft maximum τ ≤ 10, ceilings) after 1 and 3. Worked only if a rule beats steps by more than 0.03 on every unit's list share while a 99/0/0/0 unit still outranks 80/80/80/80.