First simulation findings
The first runs identify two design priorities: the defensive recipe is too successful against the other tested recipes, and going first has a substantial advantage in the CPU mirror matches. These results concern this card pool and these related heuristic policies. They do not establish human balance.
Open the interactive report · Play the experimental decks
Runs and reproducibility
The initial screen used 20 seed pairs per matchup and CPU style: 720 games. A fresh confirmation batch used 100 pairs per matchup and style: 3,600 games. Both included the three recipes, their mirror matches, swapped starting positions, and balanced, aggressive and control priorities.
The screen completed all 720 games. The confirmation completed 3,597, with three games stopped at the 60-turn cap and no engine errors. All three capped games were Hold and Finish mirrors. Their recorded traces show late fatigue: two are the same board-standoff seed under different CPU styles, and the third has both boards empty. They remain incomplete results, not draws.
The confirmation's completed games had a median of 18 player turns and 69 engine actions; the 90th percentiles were 28 turns and 103 actions. These are whole-batch figures. Pacing by individual matchup is a useful next reporting addition; the current display does not isolate that difference. None of these counters establishes human match duration.
- Initial screen and inputs
- Fresh-seed confirmation and inputs
- Matching engine, simulator and CLI source archive
The reports embed their catalogs, seeds, CPU weights and source fingerprints. All 15 retained confirmation traces, including the capped games, reproduced their exact final state. A separate browser-generated report also matched the CLI fingerprints and replayed all 12 retained traces.
The defensive recipe leads
These are Hold and Finish's observed scores in the fresh confirmation. There were no draws or caps in these non-mirror matchups, so score equals win rate.
| Opponent | Balanced CPU | Aggressive CPU | Control CPU |
|---|---|---|---|
| Pressure | 83% | 61% | 86% |
| Support and Setup | 81.5% | 71.5% | 80.5% |
Each cell uses 100 seed pairs, or 200 games. Conservative 95% bounds have a radius of about 13.6 percentage points. The size and repetition of the advantage justify investigating the recipe; the aggressive Pressure matchup is less decisive than the others.
This is evidence against claiming that all three starter recipes are competitive. It does not tell us whether the cause is only the card values, the deck mixes, CPU weaknesses, or some combination. The original 20-card duel is kept unchanged for human comparison.
Starting order needs its own experiment
In confirmation mirror matches, the starting player won 62.5% of completed games with balanced priorities, 74% with aggressive priorities, and 65.1% with control priorities. Each style scheduled 600 mirror games; incomplete games are excluded from its score.
The next starting-order experiment should compare a small compensation for the second player against the same frozen card pool. Do not change Guard costs and the opening rule together: their effects would be harder to separate. The current starting rule still only skips the first player's initial draw.
Ability choices are being used
Ember Slinger and Standard Bearer both attacked and activated their abilities under all three CPU styles. That makes them useful candidates for human playtesting: the extra action option is being exercised, rather than ignored entirely.
Field Restorer's healing was uncommon with balanced or aggressive priorities and more frequent with control priorities. That is a question to inspect in play, not sufficient evidence to remove healing. A short-sighted CPU may undervalue preserving a useful unit.
Action counts include mirror seats and partial games and are reported separately by recipe and policy. A 50–50 split is not a design target; situational choices should follow the board.
One controlled cost experiment
A separate candidate increased Iron Warden's cost from 2 to 3, preserving its 1 power, 5 health and Guard. The baseline Warden otherwise has the same cost and power as Stone Sentry with an extra health point. Only the candidate catalog changed; the engine, deck recipes, seeds and CPU priorities stayed fixed.
| Hold and Finish opponent | Balanced before → candidate | Aggressive before → candidate | Control before → candidate |
|---|---|---|---|
| Pressure | 85% → 67.5% | 70% → 60% | 80% → 72.5% |
| Support and Setup | 82.5% → 80% | 77.5% → 70% | 82.5% → 80% |
This reuses the initial screen's 20 pairs per cell, so it is a controlled screening comparison, not an independent confirmation. Pressure versus Support is unchanged. All 720 candidate games completed without errors or caps. The change consistently reduced the defensive advantage, but did not establish a competitive Support recipe.
The candidate remains unpromoted. Its complete report and catalog are retained so the result is reviewable. The playable lab still uses the original experimental values.
Next design decisions
Prioritize the early defensive-unit rates and the support recipe's ability to convert a board into a win. Test one narrow revision at a time, then check a promising candidate on fresh seeds. Run the separate starting-order experiment before making claims about fair matchups.
Play the ability units in human games to assess the fun and clarity of their choices. Keep the pool small while these questions are open. Adding enough cards to reach 100 would not itself create more viable strategies.