Tactical Reroll
⋯

The tier system's method: an independent appraisal (2026-10-02)

From reports/tiers/2026-10-02-method-appraisal.md , rendered when the site is built.

A methods review for Jordan, read-only. Built on the brief, the earlier critique (not repeated here), round 9's report, design/scoring.md section 6 and design/decisions.md. Numbers are from reports/tiers/2026-10-02-tyranids-round9.json and 2026-10-02-round8-sliders.json (and only the mi list counts from docs/data/rankings.json, to rebuild the list-share pool of 1,105 units). Nothing here changes a score.

How the numbers were made. Round 9 changes only Anvil, so round-9 jobs are round 8's percentiles with round 9's Anvil (Tyranids: their own round-9 jobs). This rebuild reproduces the Tyranid ρ (0.493), the Ork ρ (−0.118) and every unit's list-share ρ (0.077; round 8 0.162) exactly. Not the Marines: 27 units whose best detachment changed differ by over 2 points, so the rebuilt Marine baseline is −0.245 (the run's −0.099); compare within that column only. SEs: about ±0.13 Tyranids, ±0.19 Marines, ±0.15 Orks, ±0.03 every unit's list share (n = 1,105), the only yardstick that separates options closer than 0.1.


1. Do the letters matter? Rank first, letters on top

Short answer: yes, score by position and add the letters afterwards as a display layer. The yardsticks already work that way, and the half that doesn't is the noisy half.

On rank, the yardsticks gain stable, cut-free numbers with honest SEs, and lose the same-letter counts, which visitors understand; keep those as a secondary line on frozen quotas. The product loses an absolute S across armies; keep it as the default view if Jordan wants, labelled an untested claim.

Recommendation. Score = the pool percentile of the composite (0 to 100, uniform by construction). Publish the overall rank and the within-army rank. Letters = cuts on the overall score at frozen shares (S 8%, A 15%, B 27%, C 30%, D 20%, the round-9 proposal), set once, so meta drift shows as letters moving. Validate on within-army rank.

2. Are we scoring too high?

Short answer: yes. 36% of all units have a job at 90 or more and 21% score 95 or more. Almost all of it is arithmetic: the best of four uniform percentiles. Correlated jobs and the percentile transform add almost nothing, and the steps add the last bit.

Counts of units per 10-point bin, all 1,161 units:

Bin0s10s20s30s40s50s60s70s80s90s100+
Hammer117118116121116115115111116116–
Anvil116117117111110106105131125123–
Runner128104117113117117116118114117–
Banner66049264243273151116116–
Best allowed job314163555100195333419–
Score (best + steps)3141635549919131340144
best^k, rarity-corrected (a plain re-rank gives 116 a bin)689399125135125148129118121–
HammerAnvilRunnerBannerPoints
Hammer10.61−0.34−0.540.58
Anvil1−0.33−0.440.44 (round 8: 0.02)
Runner10.17−0.69
Banner1−0.27

Hammer and Anvil move together, which lowers the maximum. Runner and Banner run against both, which raises it. The two nearly cancel.

3. Raw first, then weighted, then curved: the four-layer architecture

What it fixes. Meta knobs turn without rerunning the engine, so the Lab's sliders can cover every knob; sweeps become cheap and exact; "traceable to a named number" becomes literal (raw × weight × curve × ceiling); lenses become config presets, not code.

What it risks.

How far the tool is. Halfway. The raws exist inside tiertest.js, and rankings.json publishes the live Offense/Defense ones. But meta weights are applied before the percentile, constants are globals read mid-run, and context selection uses the composite. The round-8 sliders file is a working layer 4. The unlocking step: a raw dump plus a pure compose(raw, config) that reproduces round 9 exactly.

The spec (one page)

Layer 1: the raw table (build/tier-raw.json; one row per unit × detachment × delivery × loadout context; derived numbers only, never datasheet text; regenerated when the BSData commit, engine version or scenario hash changes)

ColumnMeaning
id, army, unit, ctx{det, way, kit}, pts, modelsidentity and context
allow{H,A,R,B}eligibility flags (OC 0, joins as Support, one use, immobile)
dmg[GEQ,MEQ,TEQ,LV,HV]enemy points destroyed per 100 pts per turn in range
inRange[1..5]the delivery profile: chance of being in range each round
htd[6 attackers]absolute hits to destroy (heals included)
ocTurnsOC × min(turns alive, 5), before the meta mix (per attacker, so layer 2 can mix)
holdposition share: midfield and enemy-half objectives held
reach[mid,far,deep]share of the game it can reach each band
ld, belowHalfLeadership pass chance, share of the game below half strength
perUnit100, footprintunits per 100 pts, footprint
lend{leader,aura,transport,debuff,shock,cp,spawn}points of value lent per game, by kind (CP and Battle-shock in counts, priced in layer 2)

Layers 2 to 4: the config (data/tierconfig.json; each knob has a default, a source and a date)

KnobDefaultSource
scenario (gap 30", Advance 3.5", objectives 15/21/27", DS 9", reserves 18", rounds 5)as today11th-edition rules + builder; versioned; layer 1 input, not a slider
target.w[5]10/6/38/22/23fitted on winning lists' points, window W1
fire.w[6]8/20/37/7/22/6fitted on winning lists' weapons, window W1
price.cp, price.shock, price.debuff30 pts, 10% of median, 0.5Jordan, 2 Oct
hammer.peak0.75builder (step g); applied on the log-ratios
anvilocTurns per 100 pts × hold × Battle-shocksee section 5
runner.partsreach, turns alive, Ld, 0.5 × (units, footprint)builder
curveper raw quantity ℓ = log2(x ÷ pool median), floor −4; job raw = Σ w·ℓ; job score = pool percentilemethod
job.ceiling{H,A,R,B}1, 1, 1, 1meta lens; set from a named source, never from a panel
composebest allowed job × ceiling, + steps (+5 per other job ≥ 90, +2 at 80 to 89)Jordan's rule
finalpool percentile of the compositemethod
lettersfrozen cuts at S 8 / A 15 / B 27 / C 30 / D 20% of the calibration runJordan to approve once
poolnon-Legends, fieldable in 2,000 ptsmethod

The curve. Log-ratio to the pool median puts every raw quantity on one scale ("one step = twice the median") and turns a weighted product into a weighted sum, so a weight means the same thing whatever the units. The job is then ranked, because the order is what the panels test and a rank survives outliers (Titans). The composite is ranked again at the end, which removes the best-of-four pile-up (section 2).

Lenses. A lens is a named partial config: target.w, fire.w, job.ceiling and price.* only, never scenario or letters. For example "Armour meta": HV 35%; "Objective missions": Anvil and Runner ceilings 1 with Hammer 0.9. The client applies it to the published layer-3 job raws (our own numbers) and shows "under the Armour lens: A", always beside the default letter, never replacing it. Because lenses are our software over our own derived numbers, they can be a supporter perk under the house rules. The default letter and the raw table stay free.

4. Weighted sum or best job

Option comparison on the rebuilt round-9 jobs (ρ; every unit's list share over 1,105 units; the Marine column is a rebuild, compare within it):

OptionAuspexMarinesOrksList share, every unitTyranid list shareOne sentence?
(c) Best job + steps (round 9)0.493−0.245−0.1170.0810.453"Best job, +5 per other top-10% job, +2 per top-20%"
Best job only0.488−0.122−0.1120.0840.454yes
(a) Sum, 25% each (disallowed = 0)0.5070.008−0.0730.1490.394"Your four jobs averaged"
(a) Sum, 40/30/20/100.440−0.2210.1530.1150.232yes
(a) Sum, Hammer 90 / Anvil 5 / Runner 3 / Banner 20.300−0.3680.3440.0410.084yes
(b) Ceilings H100 A90 R90 B900.488−0.3120.0180.0970.433"Best job, each marked out of what the meta pays"
(b) Ceiling on Anvil only, ×0.80.531−0.125−0.1740.1290.537same
(b) Anvil ×0.8, with steps0.531−0.164−0.1810.1240.537same
Soft maximum, τ = 10, Anvil ×0.80.515−0.138−0.1450.1200.540"Mostly your best job, a little for the others"
Soft maximum, τ = 40, Anvil ×0.80.547−0.014−0.0520.1400.504hard
max(sum's percentile, best job)0.471−0.144−0.1280.0970.450"Specialist or all-rounder, whichever is better"
0.75 best + 0.25 mean (dropped; reference)0.495−0.079−0.1250.0950.452–

Reading it.

Reconciliation. Meta importance goes in as per-job ceilings in the 0.7 to 1 range, not as weights in a sum. The rule stays "best allowed job (× its ceiling), plus steps". One sentence for a visitor: "A unit's score is its best job, with each job marked out of what the current meta pays for it." Ceilings default to 1. Set Anvil's from a named source only after section 5's fix, and expect it to be close to 1 if the fix works.

5. The round-9 list-share drop

Short answer: it comes from the division. More precisely, removing one division by points multiplied Anvil by the unit's cost, which made expensive units into Anvils. The gate didn't cause it and the cuts can't have: list-share ρ uses the continuous score.

Recommended form: Anvil = OC × min(turns alive, 5) ÷ points × Battle-shock × position (× revive), with no gate. It reads "objective-control turns bought per point". It divides once, as the critique wanted, but survival stops counting at the game's length, so a 3,500-pt Titan can't buy its way up. It replaces the gate: Tu below 1 already scales the value down, smoothly rather than in steps. On the Tyranid raws it is cost-neutral (ρ with points −0.00, against round 8's −0.10 and round 9's +0.42), and it ranks the gaunts, Gargoyles and Warriors first and the Hierophant nowhere. Turns alive should use the meta-mixed fire as Hammer's does, so it stays absolute (Jordan's round-5 rule).

6. What else to change first

Ranked by expected effect on both the panels and list share. None of these is in the critique or decisions.md.

  1. Fix Anvil (section 5). It is the single biggest lever on both yardsticks, as the Anvil-ceiling sweep shows.
  2. Change the list-share headline. The pooled ρ mixes 27 armies and 408 tied zero-share units. Report instead (i) the median of per-army ρ (round 8: 0.188; round 9: 0.113), (ii) the AUC for taken against never-taken (0.585 → 0.536), and (iii) ρ among taken units only. And track an overfit gauge: Tyranids' per-army ρ is 0.45, second of 27 armies, while the median is 0.11. The tuned army leads because it is tuned. When the Tyranid lead shrinks while the median rises, the method is generalising.
  3. Limit the pool to fieldable units. Eight units cost more than a 2,000-pt list, and both Titanicus factions are 100% S at the current cuts. They set everyone's percentiles and head the S-per-army table. Drop them before taking percentiles.
  4. Percentiles within weight class instead of a points gate. Runner's hand-set gate is the most sensitive constant (399 letters), and three jobs track cost. Ranking each job within the King of the Hill classes (≤ 75, 76 to 130, 131+ pts) makes cost-neutrality part of the method, not of tuned gates.
  5. A final rank transform plus frozen quotas (section 1). It changes no ρ, but it removes the visible pile-up and makes the S-per-army check meaningful.
  6. Layer-1 dump and pure compose (section 3). No yardstick moves, but every later experiment becomes a config diff, and the context-selection circularity goes.
  1. Anvil = OC × min(Tu, 5) ÷ pts × Battle-shock × position, no gate. Test: change hold9, rerun --r9json. Worked if every unit's list share ≥ 0.15, no unit over 1,000 pts at S, Auspex ≥ 0.47, and Anvil's ρ with points within ±0.15.
  2. Per-army median ρ, AUC and ρ among taken units as the list-share headline, with the Tyranid overfit gauge. Test: recompute over the round-9 history. Worked if the versions' order differs from the pooled ρ's (the pooled number has been misleading us); adopt either way.
  3. Pool limited to fieldable units. Test: a filter before percentiles. Worked if panel ρ moves under 0.01 and S-per-army ρ rises above −0.24.
  4. Final rank transform + frozen 8/15/27/30/20 cuts, panels on within-army rank. Test: re-letter round 9. Worked if S is 8 ± 1% and S-per-army ρ ≥ 0.
  5. Job percentiles within weight class, Runner's points gate removed. Test: re-rank round 9's jobs by class. Worked if Auspex ≥ 0.47 and every unit's share doesn't fall, with one hand-set constant gone.
  6. Anvil ceiling 0.8, only if 1 doesn't recover list share. Test: multiply in compose. Worked if every unit's share rises 0.02 or more, the per-army median rises, Auspex doesn't fall more than 0.05, and Hormagaunts stay B or better.
  7. Raw dump + compose(raw, config). Test: --rawdump and a reproduction check. Worked if 0 units differ from round 9 and a full sweep runs in under a minute.
  8. Composite-rule bake-off (best, steps, soft maximum τ ≤ 10, ceilings) after 1 and 3. Worked only if a rule beats steps by more than 0.03 on every unit's list share while a 99/0/0/0 unit still outranks 80/80/80/80.