# How Mirrorfall Was Tested

Mirrorfall has been put through deterministic, multi-agent table simulations built to find confusion, exploits, slow turns, runaway power, and campaign-state failures before they reach a live table.

These are not public human-table sessions. They are repeatable internal playtests with recorded dice, rulings, state changes, and fixes. That distinction matters: the evidence is substantial, but it should be described accurately.

## The Test Ledger

| Test surface | What was exercised |
| --- | --- |
| `56` archived test reports | Early campaigns, final verification, rule gates, max-level pressure, fun and tedium passes, and post-fix checks |
| `18` episodes in exact order | A seeded five-player campaign from Episode 0 through Episode 17, including carry-forward state and final Cut |
| `6` audience stress arcs | Book-club and romance, new and casual, secret drama, tactical optimization, quiet players, and comedy chaos |
| `3` randomized campaign paths | Chaotic casual play, optimizer and combat pressure, and an overloaded Pressure Keeper with rotating attendance |
| `1` beginner campaign to Degree 21 | Creation, advancement, custody, absence, re-entry, high-tier play, capstones, and the Final Proof Bundle |
| `8` max-level playbooks | Every playbook at Degree 21 and Fourth Ascent with mastery branches, capstones, economy limits, and source-boundary attacks |
| `16` linked sessions | Long-campaign durability for economy, advancement, relic load, faction clocks, re-entry, fatigue, and support caps |
| `44` hostile entries | Catalog review plus a deterministic mixed-threat encounter with separate player and hostile pressure |
| `17` factions and `12` fronts | Faction-turn audits and an eight-turn seeded campaign transcript across answer, ignore, and advance states |
| `27,839` rule and content rows | A conservative row-level inventory so unreviewed material stays visible instead of inheriting a blanket pass |

The machine-readable `PLAYTEST_EVIDENCE_LEDGER.json` inventories and hashes each unique evidence artifact, identifies whether it is a simulation, audit, protocol, or derived summary, and records exact mirror copies without counting them as additional tests. Its human-session count remains `0` until an observed human-table record exists.

## What We Tried To Break

The simulations attacked Mirrorfall from the angles that usually make a tabletop game wobble:

- first-session confusion and slow character creation;
- power gaming, stacked bonuses, and max-level capstones;
- rules-lawyer pressure and ambiguous timing;
- action-first players who fight, grab, split up, or force a scene;
- lone-wolf, main-character, PvP, refusal, and social-friction pressure;
- relic custody, strange objects, dangerous tools, offices, seats, and proof;
- missed sessions, rotating attendance, late entry, replacement, and re-entry;
- long-campaign economy, advancement, clocks, fatigue, and record keeping;
- attempts to turn Tommy, the Spire, the Codex, or source lore into player-owned authority.

## What Stayed Under Control

- Mark Pools remain capped at `6d6`.
- Cut Roll support and penalties remain capped at `+5/-5`.
- Relics, tools, charges, offices, seats, proof, and source authority stay in separate lanes.
- House Load prevents a group from carrying every useful object forever.
- Capstones still require route, proof, witness, cost, review, and a House record.
- Random tables create prompts, not offices, rank, proof, or authority.
- Success does not erase damage; the next session can still inspect what happened.

## What Changed Because Of Testing

The test record is not a pile of green checks. It changed the game:

- first-character creation became faster and gained a clear minimum before the first roll;
- `Want / Risk / Record` became the live before-roll contract;
- PvP, refusal, consent, table-captain, and adversarial-player procedures became explicit;
- reward, economy, House Load, relic custody, and prepared-object rules became tighter;
- action-first play gained clear answers for fights, grabs, split parties, and forced scenes;
- scene close, reconciliation, downtime, d100 use, and campaign carry-forward became faster;
- missed-session, late-entry, replacement, and re-entry procedures became usable at the table;
- high-level play gained complete Degree 21, Fourth Ascent, mastery, capstone, and Final Proof Bundle procedures;
- quiet-player, casual-player, and overloaded-Pressure-Keeper paths gained specific support.

## What The Game Produced

The strongest signal was not a score. It was the kind of scene that survived the tests.

Players refused a perfect-looking answer because its cost was wrong. A public excerpt helped the House while the sealed original remained dangerous. A damaged name returned with a visible scar instead of being polished into nothing. A temporary holder protected an office without owning it. A final Cut closed while the record still showed what failed.

That is Mirrorfall's promise: strange power, difficult choices, and consequences that remain visible enough to create the next session.

## The Honest Boundary

This evidence supports confidence in the tested rules and campaigns. It does not claim public human-table validation, universal balance, finished art, print proof, storefront readiness, or that no future table will find friction.

The next meaningful rules change must be tested against the surface it changes. Human-table sessions should be published as a separate evidence class when they occur, never blended into the deterministic ledger.
