▸ list builder / methodology

Signal · this list

How the win-rate estimate is calculated

The percentage in your Signal panel is a read of past tournament results — how often lists shaped like yours have actually won games. It is a pattern found in real data, not a prediction of your next game and not a score of how "good" your list is.

46.9%
est. win rate
faction 42.3% +4.6 · n=124 players
↑ recomputes live as you edit

Before the math — read it right

✓ What it is

The share of games won by tournament lists that resemble yours, weighted toward recent, larger events.

✕ What it isn't

A prediction of any single game, a skill rating, or proof that a unit causes wins. It's correlation in results.

→ How to use it

A directional signal. A few points of difference is noise; a large, stable gap is worth listening to.

The number, in two parts

Your faction's baseline, plus what your list changes

Faction baseline — average for every list of your faction This list — the net of what your specific choices add or subtract

Every unit, enhancement, and detachment nudges that second part up or down — which is why each row in the builder shows its own ± contribution.

How the cogitator gets there

Seven steps, in the order they're computed

01Where the data comes from

Every game is counted from each player's own perspective: a win scores 1, a draw 0.5, a loss 0. Those results come from ingested matched-play tournament games — competitive events, not casual or narrative play.

This is why a Crusade list shows a matched-play data note: the numbers are drawn from a different format than the one you're playing, so treat them as a rough guide there.

02How games are weighted

Not every game counts equally. Each result carries a weight — the product of two curves — so the estimate tracks the current meta and leans on stronger events.

Recency half-life 28 days

weight = 0.5 (age ÷ 28d) — a game loses half its say every 28 days.

Event size skill proxy

bigger events weigh more (log scale, full weight ≈ 200 players).

Because weights are uneven, sample size is reported as an effective count (n_eff) — closer to the true amount of independent evidence than a raw tally.

03The faction baseline

The baseline (42.3% above) is the weighted average win rate of every list in your faction, using the weights from step 02. It's the "starting line" your specific list is measured against — above it means your choices trend better than the faction's average build.

04Your list's model

For each faction we train one regression that learns how a list's contents relate to winning. Your list's probability is:

P(win) = σ( b0 + Σ βi · xi ) σ = logistic curve · b0 = faction intercept · βi = each feature's learned weight · xi = your list's features

The features (the x's) describe what's actually in the list:

FeatureWhat it captures
unita datasheet is present in the list
duplicatesrunning 2–3 copies (each extra copy counts for less)
squad sizea unit taken at its max model count
enhancementa specific enhancement is equipped
wargearcount of a swappable weapon option (capped)
detachmentthe detachment the list is built around
unit pairstwo units that show up together — so context matters

Illustrative — how a few features might net to the +4.6:

Illustrative figures for shape only — open your list to see its real per-unit contributions.

Thin evidence is pulled toward zero (ridge shrinkage), so one lucky unit in five lists can't swing the number. As more games arrive, the model trusts the data more.

The formula above is only the model term. On its own, adding up many individually-promising features can climb to a number no real list has ever achieved — so before anything is shown, the estimate is grounded against real lists (step 05).

05Grounded in real lists

A sum of correlations can run away: stack every unit the data likes and the raw formula will happily claim 90%+ — a number no real list has ever sustained. Three guards keep the estimate honest, and they're baked into the same artifact your browser scores with:

GuardWhat it does
real-list anchormost of the estimate comes from how actual tournament lists similar to yours performed — a whole-list comparison, not a sum of unit correlations. If nothing like your list has really played, it reads near the faction baseline instead of extrapolating.
support clampthe model term is capped at the range of the real lists it was trained on — no list can outscore the best list the data has actually seen.
calibrationthe model is tested on games it didn't train on; if it's over-confident there, all its predictions are compressed toward the faction's base rate before you see them. It can be shrunk, never amplified.

The practical upshot: an exotic pile of individually "good" units reads close to the faction baseline — not as a fantasy 92% — because nothing like it has proven itself in real games. Estimates only leave the baseline when real, similar lists earned it.

06Live recompute as you edit

The whole trained model is sent to your browser, so every edit re-scores instantly — no server round-trip, no waiting. Add a unit and the estimate moves; that movement is the ± you see on each row.

The in-browser scorer is checked against the training code by an automated parity test, so what you see matches what the model actually learned.

07Confidence & limits
low data · faction < 75 players low-confidence feature < 30 lists rare features pruned < 5 lists

When evidence is thin the panel says so, and you should weight the number accordingly. Two limits are worth keeping in mind:

Associational, not causal. Strong players tend to bring strong lists, so a unit's positive weight partly reflects who plays it, not only the unit itself.

The meta moves. The estimate follows results as they're ingested; a fresh edition or a big event can shift it before the trend is settled.

Optional · narrow the field

Filtering to the opponents you'll face

The matchups list shows your faction's win rate against each opponent faction. Flip the filter switch and tap the factions you expect to face — the estimate and every per-unit ± are recomputed against only that field. Useful when the field is known: a small event, your local meta, or a team tournament's locked factions.

This isn't just re-averaging the matchup rates. On top of the faction model (steps 03–04) a second layer is trained: for every matchup with enough games, a set of opponent-specific offsets that nudge the model's coefficients toward what actually happens in that pairing.

P(win) = σ( etag + d0 + Σ di · xi ) etag = faction model log-odds · d0 = matchup baseline shift · di = per-feature offset vs that opponent

Against a field of several factions the score is the game-weighted average of the per-opponent predictions, so a common opponent counts more than a rare one. The offsets sit on top of the global model, so turning the filter off restores the exact all-opponents number — your selection is kept (dimmed) so you can flip between "the whole meta" and "my field" freely.

Where it's honest about thin data

Matchup data is sparse — split thousands of games across ~900 faction-vs-faction cells, then by which units were present, and most cells are thin. The model degrades gracefully and says so:

SituationWhat happens
matchup under 30 gamesno offsets trained — that opponent scores through the league-wide model
unit measured vs fieldits row shows a — the ± is specific to the selected opponents
thin unit effecta muted — measured but low-confidence
narrow fielda panel banner reports what fraction is matchup-backed vs league-wide

Offsets are ridge-shrunk harder than the main effects and only shipped when they move the number meaningfully, so a unit with no real matchup signal reads at its league-wide value rather than inventing a difference. In practice only a handful of units per matchup carry a credible opponent-specific tilt — the rest, honestly, we can't yet split by opponent.

A different question · what people bring

"X% of similar lists took this"

Everything above predicts winning. This one doesn't predict anything — it counts. Next to a unit, an enhancement, a weapon or a leader you'll see how often lists like yours made the same choice, straight from parsed tournament decklists.

Common is not the same as good, and we're careful not to imply otherwise. A unit almost everyone brings often scores a win-model effect near zero — when a choice is nearly universal, there's no contrast left for the model to measure, and its value gets absorbed into the faction baseline. So adoption is shown in a deliberately plain style, with no red/green: it tells you what the field looks like, not what to do.

"Similar" means a specific set of lists

A percentage is meaningless without knowing what it's out of, so the cohort is always named on screen, with its size. We use the most specific one with enough lists behind it:

CohortWhat it holds
your exact detachmentslists running the same combination you are
your primary detachmentevery list leading with it, whatever it's paired with
your factionthe fallback — labelled faction-wide so you know

A cohort needs at least 25 lists to be used at all, and a single choice needs to appear at least 3 times before it gets a number. Below that you'll see no read rather than a percentage — "too few lists to say" and "nobody does this" are different claims and we won't blur them.

All lists, or only winning ones

By default the cohort is every tournament list in the current window — they're already a competitive population. The winners toggle narrows it to lists that won at least two thirds of at least three games. That's a much smaller sample: it takes usable detachment-level cohorts from 142 down to 46, so it will often fall back to a broader rung, and the panel says so when it does.

Where each number comes from — and its limits

SignalCounted againstWorth knowing
unitslists in the cohortplus average copies among lists that took it
enhancementsthat unit's appearances"no enhancement" is shown too — usually it's the majority
wargearthat unit's appearancesa squad mixes weapons, so these don't add to 100%
leadersattachments actually written downonly some lists record who leads what — see below

Decklists never state which leader joined which unit directly — they tag a unit's role and rely on where it sits in the list. We reconstruct the pairing from that, which works for 99.9% of tagged entries. But only about two thirds of lists include the tags at all, and it varies hugely by army: Ork lists record it 84% of the time, Imperial Knights lists 2% — they have almost nothing to attach. Where there isn't enough to read, the panel tells you the coverage instead of showing an empty box.

One more limit worth stating: a list that our parser can't break into units contributes nothing here, so cohort sizes are counted over lists we could actually read.

The one-line version

Use it as a signal, not a verdict. It points at what tournament results have favored so far — the deciding factors are still your plan, your matchup, and your dice.

Questions about a specific number? Ask the logis from your list — it can read your exact contributions and explain them in context.

Chapter two · the optimizer

The same number, searched instead of read

Chapter one explained the estimate in your Signal panel. The optimizer is that estimate run backwards: instead of scoring the list you built, it builds thousands of legal lists and keeps the ones the model scores best. Same data, same model, same caveats — it knows exactly what the signal knows, and nothing more.

✓ What it is

A fast, legality-checked search for lists that historically win — honoring your locks, your collection, and your chosen opponents.

✕ What it isn't

A guarantee, a meta oracle, or a replacement for playtesting. It inherits every blind spot the data has — and it will happily exploit them.

Why it can't cheat (much)

Honesty guards, in plain words

A search pointed at a raw statistical model finds the model's blind spots: it used to be able to stack individually-correlated units into lists scoring an absurd 90%+ that no real list has ever sustained. The score it maximizes now is chapter one's grounded estimate (step 05): anchored to how real, similar lists actually performed, capped at the best real list the data has seen, and compressed when the model is over-confident out of sample — minus a penalty for relying on low-data coefficients, and minus a penalty for leaving points unspent. Every list it emits also passes the same legality validator your own edits do — points, rule of three, warlord, enhancements, detachment rules.

Locks

⚑ keep and ⛉ freeze — your list, your rules

Every unit row carries a small lock that cycles through three states. ⚑ keep: the unit must stay, but the optimizer may retune its size, wargear and enhancement. ⛉ freeze: untouchable, exactly as configured. If your frozen units can't form a legal list together, the run fails with the reason — your locks are never "helpfully" overridden.

Targeting

Optimizing against someone in particular

The optimize tab reads whatever matchup selection is active in your Signal tab: no selection means "the whole field"; selected matchups mean the search maximizes the estimate against those opponents specifically. The same narrow field warning from chapter one applies — where a matchup has too few games, the number quietly falls back to the league-wide estimate, and the panel tells you how much of your field is actually covered.

Your collection

"Only units I own" — and what to buy next

Record the models you own on the My Units page and the optimizer can build strictly from your shelf — owned-only means owned-only, down to model counts. On My Units, a second unconstrained search runs alongside and surfaces the units the model most wants that you don't own, with the score ceiling you're leaving on the table.

The sweep

Comparing detachments honestly

"Compare detachments" optimizes the same inputs — locks, collection, targeting — under every detachment of your faction and lines the results up in one table. Detachments a frozen unit can't enter are shown as skipped with the reason, never silently dropped. It's the expensive button (one full search per detachment), so it has its own tighter rate limit.

Results

Options are diffs; applying is one undo

Results render as changes against your current list — green adds, struck-through removes, amber resizes — because what you're evaluating is the change, not a wall of names. Applying an option writes the whole change as a single revision: one ⌃Z takes it all back. Runs happen in your browser when possible (unmetered, works offline); the server steps in for collection-constrained runs and as a fallback.

Meters

What's limited, and what never is

The search itself is never metered. The ☲ analysis button spends a real language-model call, so it's limited (the meter shows what's left before you click), and the detachment sweep is rate-limited because it's many searches in one. The gallery's pre-written analyses are free to read.

The one-line version

Treat its lists as strong drafts, not verdicts. The optimizer finds what tournament data has favored — it can't see your local meta, your terrain, or your playstyle. Lock what you love, let it argue for the rest, and playtest before you trust it.