Tactical Reroll
⋯

Design review of the tier engine, after round 11 and plan v3 (2026-10-02)

From reports/tiers/2026-10-02-design-review.md , rendered when the site is built.

By the design reviewer session (Fable), read-only: CLAUDE.md, design/tier-plan.md (v3), design/tier-engine.md, design/matrix.md, design/decisions.md (to the 23:30 UTC entry), the round-11 report and JSON, benches 1 to 4, residuals 1 to 3, the brief's critique, the plan's critique and the method appraisal. Every number below is read from the tracked JSONs (round 9, 10 and 11, v5-every, metatrack.json, rankings.json, inclusion.json); nothing was re-run and no code changed. A report: nothing here changes a published score.

Verdict

The method is sound in its bones and is about to build the wrong thing first. The architecture (raw table, pure compose, one config with provenance, conventions frozen before fitting, locked panels, a decision log that forbids repeats) is right and should be kept. The matrix (T-568) is the right layer 1. But the two things that decide most letters today sit in layer 4 and would survive the matrix untouched: the final score is a cost ranking (its Spearman ρ with points is −0.42 and no unit over 250 points is S anywhere in the game), and the best-of-four jobs is an order statistic that hands S to the jobs most units cannot hold. Both can be measured and mended in compose in seconds, before a line of the matrix exists, and if they aren't, the matrix's first result will be judged through them.

The yardsticks are mostly honest and improving (SEs, shuffle baselines, three panels led, the overfit gauge). Two are not: "list share" is not what every document says it is (it counts lists of any placing, not winning lists), and two of the three panels are not tier lists. The things fitted rather than understood are named in decisions.md with unusual candour; the one that isn't named is the gate the queue hangs on: Jordan's "38 of 44 within one of Auspex" line, a target on the one panel the engine was tuned on.

The checkpoints and the critiques

The plan's queue (tier-plan.md §10) has nine steps; each gets one critique. Ranked by how many letters the fix would move, with the cheapest test for each.

RankCheckpoint (queue step)Critique
14, config and composeThe final score is a cost ranking in disguise
24, config and composeBest-of-four is an order statistic that favours the jobs most units can't hold
31, baselines and yardsticks"List share" counts every entrant, not winners, and rests on 1 to 23 lists an army
41, panels and the locked setOne panel is a tier list; the other two aren't; the queue's gate is set on the tuned one
52, conventionsThe "best of" conventions keep stacking, and their option premium has never been measured
63b, the allocation ruleThe mirror makes survival independent of size and of the unit's share of its army
73b, meta statesAt 112 winning lists the clustering will most likely return one state
85, the fitterThe constants with the most leverage are filed as "the table", so the fitter never sees them
96 and 7, the residual loopThe loop's human rulings are guarded; its mechanics aren't re-measured when the ground moves
108 and 9, the stop rule and betaThe stop rule lacks the two checks that have failed in every round

1. The final score is a cost ranking in disguise (checkpoint 4)

What the numbers say. Spearman ρ between the final score and points, over the 1,153 fieldable units, from the tracked JSONs:

Roundρ (score, points)
9 (composite)+0.02
10 (final score)−0.21
11 (final score)−0.42

Inside every army the lean is the same way: median −0.45 over the 30 armies with ten or more units (Chaos Knights −0.84, Death Guard −0.79, Thousand Sons −0.78, Necrons −0.54, Custodes −0.45). By points band on the all-armies cuts:

PointsUnitsSD
75 or less38354 (14%)38 (10%)
76 to 13041032 (8%)52 (13%)
131 to 2502385 (2%)80 (34%)
over 250122061 (50%)

Not one unit over 250 points is S in the whole game. Adeptus Custodes, the second most successful army in the corpus (11 top-quarter lists), has one S of 31. Auspex's own S list has six units over 165 points (both Norns, the Swarmlord, the Hive Tyrant, the Maleceptor, the Tyrannofex); the engine places them B, B, B, B, C and C.

Why. Points divide every job at least once: Hammer is per 100 points, Anvil per point, Banner per point, Runner twice (the points gate from 75 to 260 points and the cheapness part). The composite is the best of the four, so a cheap unit gets four chances to win on cost. Round 10 measured Anvil's ρ with points alone (0.19) and called the ±0.15 target "missed"; nobody measured the composite, which went from neutral to −0.42 while the Tyranid ρ climbed. The negative "S share per army against top-quarter lists" in every round since 8 is this lean seen from above: armies of expensive units get no S.

Why first. It moves two whole bands (360 units at 131 points and up) and it is where the matrix won't reach: the matrix changes Hammer's targets and turns alive but keeps "per 100 points", the gate and the cheapness part. Bench 3 already saw the edge of it: dividing by points^0.8 (C1a) lifted the Tyranid ρ 0.765 → 0.792 and moved the five big monsters toward Auspex, but it was run on Tyranids alone and set aside as "not the thread".

Cheapest test (compose only, seconds). (a) Add ρ(score, points), pooled and per army, as a standing line of every report. (b) In tierbench, make the points exponent a config knob α on every per-point job (raw ÷ points^α, α = 1 today), and drop the gate's double count (keep the gate or the cheapness part, not both). Sweep α from 0.5 to 1 and report, for each α: ρ(score, points), the three panels with SE, the S share per army against top-quarter lists, and the S count per band. Accept the α at which ρ(score, points) is inside ±0.15 pooled and within every army, if the panels hold within one SE. If no α does both, the jobs disagree on cost and the fitter should own α per job.

2. Best-of-four is an order statistic that favours the jobs most units can't hold (checkpoint 4)

The method appraisal showed that 36% of units have a job at 90 or more because the best of four percentiles is high by arithmetic. The consequence it did not draw is which units reach S. Zoanthropes at Hammer 85 (better than 85% of the game at hitting) end at final score 49.7, a C, because half the pool has some job above 85. The jobs are not commensurable: Banner has 648 units at zero (v5-every) and Runner is barred to 120 or more units by the gate, so a 90 in those jobs is reached by a smaller field than a 90 in Hammer, where nearly every unit competes.

The S lists show it. Round 11's Tyranid S and A within the army by best job: Runner 7, Anvil 4, Banner 2, Hammer 0. Auspex's S list is mostly hitters (Zoanthropes, Tyrannofex, both Norns, the Maleceptor, the Hive Tyrant). The all-armies top 8% at T-552: Runner 31, Hammer 27, Anvil 18, Banner 17.

The plan's common log-ratio curve (§3.3) is the right instrument and is "planned for the config step", unmeasured. The appraisal's rarity-corrected best (best^k) is the cheaper one.

Cheapest test (compose only). Three variants of the composite, each reported with the S job mix, the three panels and ρ(score, points): (a) today's; (b) each job's percentile taken among the units allowed that job, then the best; (c) log₂(value ÷ the pool median) per job with one scale, then the best. Expect Hammer units back in the Tyranid S and A, and the Marine ρ to move most, since the Marine panel is hitters and transports.

3. "List share" counts every entrant, not winners, and rests on 1 to 23 lists an army (checkpoint 1)

rankings.json carries mi: [taken, lists], and lists is inclusion.json's per-faction count of lists of any placing in the three-month window (tools/build_lists.js counts listsBy[faction] before the top-quarter test; inclusion.json's own note says "any placing"). So the "23 Tyranid lists" of every bench and residual report are 23 entrants, of which 7 finished in a top quarter (metatrack.json: July 2, August 4, September 1). The decision log's yardstick line ("how often top-quarter lists take it"), tiertest.js's report text ("top-quarter lists, July to September"), the plan and the explainer all say winning lists. Every list-share number in eleven rounds is popularity among entrants, not among winners.

That is not useless (entrants' choices are information) but it is a different claim, it is the second signal every round's stop rule leans on, and the base is thin: 503 entrant lists across 33 armies; 112 top-quarter lists; 24 of 31 armies have fewer than 5 top-quarter lists, median 3. The per-army median ρ (the headline since round 10) is a median of 27 correlations each on a handful of lists, which is why a data fix alone moved it 0.165 → 0.139. The AUC of 0.585 for "taken against never taken" is barely over chance and depends on armies whose "never taken" is everything but one list.

Cheapest test (a script, minutes). Rebuild the list-share lines with the top-quarter unit counts metatrack.json already holds (units per month), beside the entrants' version, for rounds V2b to 11 from their tracked JSONs. If the order of rounds changes, some switch was kept on the wrong signal. Then: rename the line "entrants' list share", report each army's n, exclude armies under 10 lists from the median, and state the pooled ρ's SE (about 0.03) as the only list-share number that can separate rounds.

4. One panel is a tier list; the other two aren't; the queue's gate is set on the tuned one (checkpoint 1)

From the panel files themselves:

Jordan's rule for restarting every-army runs ("38 of 44 within one of Auspex") is a target on that panel, and benches 2 to 4 met it by changing the yardstick to Auspex's own letter counts. The plan already says the headline must be held-out and shuffle-adjusted; the gate contradicts it.

Cheapest test (free). For Space Marines, compute the AUC of "mentioned by the reviewer against not mentioned" over all 84 Marine units: it uses the omission and costs nothing. For Orks, report lift over shuffle only and drop its ρ from the headline. For Auspex, read the transcript and date it (an hour), so its patch alignment and reasons exist. Replace the 38-of-44 gate with "lift over shuffle on every fitting panel, none worse than frozen round 11".

5. The "best of" conventions keep stacking, and their option premium has never been measured (checkpoint 2)

A unit's score is the best over its detachments, its loadout per target, its way in, its leader (round 11), its size (round 11, reversed), and each lend on the two squads it lifts most. Each maximum over noisy options rises with the number of options. Space Marines, with 16 Codex detachments and the most shared datasheets, hold 14 S of 84 with one top-quarter list. The brief's critique asked for the regression of score on option counts; decisions.md §4 still says "not yet tried", and the matrix spec adds the best legal detachment in layer 4 again.

Cheapest test (compose only). Regress the final score on log(contexts per unit) with the Hammer and Anvil raw rows as controls; then re-score with each army's most-listed detachment fixed and lends on the median reachable squad instead of the best two. Report the S count per army and the panels. If the premium is real, the mend is to subtract it (the expected maximum under shuffled options) rather than to add more conventions.

6. The mirror makes survival independent of size and of the unit's share of its army (checkpoint 3b)

In matrix.md §7.5, X is a marginal entrant: each enemy gun's share into X is proportional to X's wounds present, so the hazard per wound does not depend on how many wounds X has. Six Zoanthropes and three have the same T_alive. That dissolves the size question (consistent with Jordan's minimum-size convention) but it is the opposite error to today's: today the fire is fixed and absolute survival grows with size; under the spec the fire into X grows without limit as X grows, so a brick never exhausts the enemy's budget. It also ignores that a 320-point block is a sixth of its own side, not a marginal addition to a 2,000-point field.

Cheapest test (on paper, before building). Add to §10 a worked case: three and six Zoanthropes against one fixed enemy fire budget, T_alive under the mirror and under a budget-limited version where X displaces p_X points of its own side's field (both sides at 2,000 including X). Then make T-569 step 4's test say what doubling X's size should do to T_alive (the budget's answer), not only that copies of X elsewhere leave it unchanged. Put "X displaces its teammates or joins on top of them" to Jordan as the spec's seventh question.

7. At 112 winning lists the clustering will most likely return one state (checkpoint 3b)

Jordan chose the top quarter for the field (decisions.md, 23:30): 112 lists for July to September, 13 of them Necron. k-medoids with the spec's own stability gate (bootstrap Jaccard ≥ 0.6, clusters under 8% merged) will probably give k = 1, or states that are one army's lists. A state built from 25 lists makes every unit's Hammer_s depend on which three armies happened to place that quarter. The spec names the risk; the product ("S in armour, B in horde") depends on the states existing.

Cheapest test (needs no matrix). Run T-569 step 2 now on the corpus and report k and the Jaccard. If it fails, the generic list is one field and the skews are named synthetic fields (the generic field re-weighted by class: "horde", "armour", "elite"), which are stable, explainable and free, with clustered states added only when the corpus supports them.

8. The constants with the most leverage are filed as "the table", so the fitter never sees them (checkpoint 5)

Round 9's sweep ranked the constants by letters moved: Runner's points gate (399), the Advance distance (288), the objectives' distances (286), Strategic Reserves' arrival distance (206), Runner's half-weight parts (183). The argued-over worths (a CP, Battle-shock, debuffs, Synapse) move 0 to 3. The plan fits about 15 weights (prices, bars, steps, ceilings, gates) and classes the geometry as the scenario block, "never a slider". So the five constants that move the most letters are either fitted on nothing (the gate's 75 and 260 come from King of the Hill's weight classes) or frozen as builders' guesses.

Cheapest test (the sweep exists). Re-run --sweep on round 11, list the top ten by letters moved, and mark each as: sourced (the mission pack, a terrain layout), a weight the fitter owns, or a convention Jordan sets. None may stay "builder". The Runner gate in particular should be derived (the points of the units that do actions in the corpus's lists) or fitted, not inherited from a game mode's classes.

9. The loop's human rulings are guarded; its mechanics aren't re-measured when the ground moves (checkpoints 6 and 7)

The consensus-wrong guard is good. But each engine mechanic that comes out of a card is built to the datasheet with a unit test and switched on for every army, which is right, and then its effect on the letters is reported once, on the day. Decisions.md shows the pattern: a fix reverses two rounds later (CP 0 → 30, Feed the Swarm, bestSize) because a later change moved the ground under it. The matrix will move the ground under all of round 11's survival fixes (turnsHarmonic, stealthSurvival, the Lone Operative ×2, the solo ×0.5); the spec says which retire but not which must be re-measured.

Cheapest test. Keep the per-switch table round 11 printed ("each fix alone" and "all but one") as a standing section of every round, so a mechanic whose contribution vanishes or flips sign is seen when it happens, not two rounds later.

10. The stop rule lacks the two checks that have failed in every round (checkpoints 8 and 9)

The stop rule (held-out ρ holds, list share falls by at most one SE, the specialist rule holds) does not include the S-share-per-army sign (negative in every round since 8) or ρ(score, points). Both are cross-army checks, and the panels and list share are within-army quantities, so nothing in the rule can catch an all-armies ruler that is wrong across armies, which is exactly the view the site shows by default.

Cheapest test (a rule). Add to the stop rule: ρ(score, points) pooled and per army within ±0.15, and the S share per army against top-quarter lists not negative. Add "frozen round 11 on the untouched panels" as the floor a new round may not fall under.

What is fitted rather than understood, as the record stands

Decisions.md names most of these itself; this is the list in one place, with what each rests on.

ItemRests onStatus
Round 11's switch setthe benches' fit to Auspex (the report says so)kept, every army
A CP at 30 points (45 lands four characters, bench 2)Jordan "for now"; T-558 to price itkept
The turns cap at 5 against 6 (bench 3: chosen with the panel in view)Jordan: 5decided
Runner's gate 75 and 260King of the Hill's weight classesnever swept against a source
Battle-shock 10%, debuff share 50%, position floor 0.5, Synapse 75%builders' guesses, kept by Jordanmove almost no letters
Lends on the two best squadsone rule for every lend; "two" is a builder'sunmeasured against the median squad
The letter cuts (8/15/27/30/20%)frozen 2 Octfine as a display layer
Target weights 10/6/38/22/23 and fire weightsfitted on entrants' listsretiring with the matrix

The conventions decided at 23:30 (minimum size, cap 5, lends on the two best squads, the job names, the six matrix questions) are the right kind of decision and are correctly out of the fitter's reach. Two are worth one more look in the light of the above: "two best squads" (critique 5) and the field from the top quarter (critique 7, since 112 lists is few).

What is sound and must be kept

The next round

Round 12 should be a compose-only round, before the matrix is read by anything. Everything in it runs in tierbench in seconds on round 11's raw table, and every result is a number the matrix round will otherwise be confounded with:

  1. The list-share correction (critique 3): the top-quarter version beside the entrants' version, rounds V2b to 11 re-scored, each army's n shown.
  2. The yardstick table rewritten (critique 4): lift over shuffle per panel; Space Marines as AUC of mentioned against not; Orks demoted to a sanity list; the 38-of-44 gate retired.
  3. The cost exponent α and the gate's double count (critique 1), with ρ(score, points) per army and the S band table as standing lines.
  4. The three composites (critique 2), with the S job mix.
  5. The option-premium regression and the fixed-detachment what-if (critique 5).
  6. The sweep's top ten, each assigned an owner (critique 8).

In parallel and independent of all of it: T-569 step 2 (the clustering) on the corpus now, to learn whether meta states exist at this size; the Auspex transcript read and dated; a second real tier list for any army Jordan can find. T-568 (the matrix) can be built meanwhile, since it has its own parity tests, but its first comparison should be against a compose that already carries round 12's fixes, measured as one change.


For the Opus supervisor: text to paste

1. A card for TASKS.md (next free number is T-575 after T-574):

T-575🤖🔲Round 12: a compose-only round from the design review (reports/tiers/2026-10-02-design-review.md)Six experiments in tools/tierbench.js on round 11's raw table, no engine change, no matrix: (1) list share rebuilt from metatrack.json's top-quarter unit counts beside the entrants' count rankings.json carries (mi counts lists of any placing; inclusion.json says so), rounds V2b to 11 re-scored on both, each army's n shown, armies under 10 lists out of the median; (2) the yardstick table: lift over shuffle per panel, Space Marines as the AUC of mentioned-against-not over all 84 Marine units, Orks reported as a sanity list only, the 38-of-44 gate retired; (3) a config knob α on every per-point job (raw ÷ points^α) and the Runner gate's double count removed, α swept 0.5 to 1, with ρ(score, points) pooled and per army and the S count per points band (≤75, 76–130, 131–250, >250) as standing report lines; (4) three composites: today's, each job's percentile among the units allowed it, and log₂(value ÷ pool median) with one scale, each with the S job mix; (5) the final score regressed on log(contexts per unit) with the Hammer and Anvil raw rows as controls, then the most-listed detachment fixed per army and lends on the median reachable squad; (6) --sweep on round 11, the top ten constants by letters moved, each assigned an owner (sourced, fitted, or Jordan's convention). Report: reports/tiers/2026-10-02-tyranids-round12.md; decisions.md and scoring.md entries. Nothing to the site.

2. A line for design/decisions.md, under "Yardsticks" at the top:

Correction (2 Oct, design review): "list share" as computed in every round to date is the share of entrant lists (any placing, July to September; rankings.json mi, from inclusion.json), not of top-quarter lists. Tyranids: 23 entrant lists, 7 in a top quarter. The top-quarter version is round 12's to add beside it; until then every list-share number in this file is entrants' share.

3. Edits to design/tier-plan.md §2 (success):

  • Add: ρ(score, points), pooled and within each army, inside ±0.15; and the S share per army against top-quarter lists not negative. Both are cross-army checks the panels can't make.
  • Add: no round may fall under frozen round 11 on any panel it did not tune on.
  • Replace the 38-of-44 Auspex line wherever it is cited (T-564's brief) with lift over shuffle on every fitting panel.

4. Edits to design/matrix.md:

  • §7.5, after the hazard: "Because the share of each gun's fire into X is proportional to X's wounds present, the hazard per wound does not depend on X's size: T_alive is the same for a unit at any size, and a unit never exhausts the field's fire. This is a consequence of treating X as a marginal entrant, chosen for v1 and to be measured (step 4)."
  • §11, T-569 step 4's test, add: "doubling X's size changes T_alive by what a fixed enemy fire budget implies (the §10 worked case), not by zero."
  • §12, a seventh question: "Does X join its side on top of a 2,000-point field, or displace p_X points of it? Recommended: displace, so a 320-point block is a sixth of its own side."

5. A note for design/tier-engine.md §7 (where we are): one sentence that the design review found the final score leaning on cost (ρ −0.42 with points, no S over 250 points) and that round 12 measures it before the matrix is read.

6. A line for reports/tiers/2026-10-03-baselines.md §6 and tools/tierbaseline.js's header comment: "List share here is the share of entrant lists (any placing in the window; rankings.json mi, from inclusion.json), not of top-quarter lists; the top-quarter version is round 12's."


Addendum, 3 Oct: the review against the branch at 51b7d59d94

Three cards landed after this review was written (T-570 the baseline battery, T-572 the frozen conventions, T-574 a locked Astra Militarum panel) and the supervisor ruled, by the plan's step-1 rule, "redesign the jobs before fitting". Read with them, the review changes in four places and stands in the rest.

Everything else above stands, and the paste text now carries the next free card number (T-575) and a sixth line for the battery's report.

Jordan's decisions on the review, 3 Oct (for decisions.md, paste as one entry)

3 Oct: the design review's questions, Jordan's answers (reports/tiers/2026-10-02-design-review.md).

  • Round 12 runs now, beside T-568: the compose-only experiments (T-575) go to one sub-agent while the matrix's first rounds build; the matrix is first measured against a compose that already carries round 12's results.
  • The every-army gate is lift over shuffle on every fitting panel, none worse than its frozen predecessor; the "38 of 44 within one of Auspex" line is a bench sanity number from now on, never a gate (T-564's brief is superseded on this point).
  • The Ork panel (Sprues & Brews) is dropped: it is a pre-points codex review, not a tier list, and a random order scores 37.7 of 45 within one on it. Move reports/tiers/panel/2026-10-02-orks-spruesandbrews.json out of the panel set (a retired/ folder keeps the record); tools/tiertest.js's hand-listed panels and tools/tierbaseline.js stop reading it; the held-out pooled number is the Space Marine panel alone until a real Ork tier list exists.
  • A second Tyranid panel, fitting: reports/tiers/panel/2026-10-03-tyranids-second.json, a reviewer other than Auspex (possibly Hivemind Hobbies; Jordan to confirm the channel, title and date), about early August (before the 5 Sep update; no graded unit's points moved over 10% since), 44 units (S 6, A 11, B 16, C 9, D 2), transcribed from Jordan's summary with timestamps. Not locked (correction, 3 Oct: only the Astra Militarum panel is locked; the Votann discussion turned out to be 10th edition and is reference only, so the plan's second lock is still open). A second Hivemind Hobbies transcription already existed (T-576, 2026-10-03-tyranids-hivemind.json, 49 units, the grades Staple / Strong Choice / Detachment Dependent / Spicy Pick / Biomass); the two share 43 units and agree on 16 letters (Spearman 0.57, two-letter gaps on both Norns, the Deathleaper, the Neurolictor and Venomthropes), so they are either two different videos or one video summarised unreliably; Jordan is being asked, and neither counts as the second reviewer until he answers. First uses: the two-reviewer ceiling for Tyranids (√ of the two panels' ρ on the 37 units both grade), and frozen round 11 scored on it through tools/tierbaseline.js, which says whether round 11 learned Tyranid worth or Auspex's style. tools/tiertest.js lists its panels by hand and must add it.

The card numbers in the paste text above: the compose-only round is now T-579 (T-575 went to the second Space Marine panel).