Crisis Comms War-Game
Most crisis communications playbooks are unfalsifiable. This one is published alongside the adversary built to attack it, and the two rules I would most want argued with are the two the simulator found before I did.
The problem this is built around
Every crisis communications playbook says the same three things: be transparent, be fast, stay consistent. All three are correct. None of them tells you what to do at 03:00 when being fast and being accurate point in opposite directions, the regulator's clock is running, and the most senior person in the room wants the reassuring sentence you cannot yet support.
I spent four years as the Singapore Police Force's official spokesperson, across 25+ press conferences and more than twenty disinformation campaigns against public institutions. That experience is the hardest thing on my CV to verify and the easiest to dismiss as a soft skill. So I did the only thing that makes it checkable: wrote the playbook down, built the adversary that attacks it, and measured whether it holds.
The failure mode this is built to prevent is not silence.
It is reassurance offered on a matter that has not been established — the confident "we've seen no evidence of X" issued at hour one because the room could not tolerate saying nothing. It is almost always the sentence that gets retracted, and it is never issued by someone who is lying. It is issued by someone under pressure who believes the reassuring answer is probably right. It usually is.
Why it had to be a loop
A crisis is inherently stateful and adversarial, so a one-shot exercise cannot model it. Four adversary agents — a journalist, a regulator, an advocacy coalition and an enterprise customer — generate each round's moves as a function of what the response side actually released in the rounds before.
Dodge a journalist's question twice and the dodge becomes the story. Deny something
the internal record contradicts and the record surfaces — facts flagged
reveal.on_false_assert exist precisely to be triggered by a denial,
because a denial tells a reporter which document is worth chasing. Tell the
regulator one thing and the press another and the regulator opens a supervisory
concern about the divergence rather than about the incident.
The modelling decision everything rests on: a statement is not prose, it is a set of claims — a canonical proposition plus a truth value. Once you model statements that way, "did we contradict ourselves" stops being a matter of editorial opinion and becomes a set comparison. The prose becomes decoration, which is why it can be swapped for model-generated text without moving a single number, and why the whole benchmark reproduces with no API key.
Three postures, same panel, same clock
Each posture is a position I have heard argued in good faith in a real room. The benchmark exists to show what each one costs.
| Measure | Playbook | Speed-first | Stonewall |
|---|---|---|---|
| Mean composite | 89.3 | 59.2 | 62.8 |
| Time to first statement | 25 min | 12 min | 220–520 min |
| Contradicted on, across 4 runs | 0 | 13 | 1 |
| Statutory deadlines met | 4 / 4 | 4 / 4 | 0 / 4 |
| Credibility subscore | 0.80 | 0.00–0.20 | 1.00 in 3 of 4 |
The trade nobody makes on purpose
Speed bought thirteen minutes and cost thirteen contradictions.
The speed-first posture reaches the wire first in every scenario, and pays for it every time by the same mechanism.
Reassurance on an open question is a loan against a fact you do not have. The reason this trade keeps getting made in real rooms is not that anyone thinks it is a good one — it is that the thirteen minutes are felt by everyone present, and the contradictions arrive after the meeting ends. The legal gate flagged every one of those thirteen before release. The posture overrode it.
Two findings the runs produced and I did not
These are the reason this exists as a simulator rather than as a document. Both were things I would not have written down, and one of them I initially tried to tune away.
Silence wins the credibility column outright.
The stonewall posture is contradiction-free in three of four scenarios and scores a perfect 1.00 on credibility in those runs.
My instinct was to adjust the rubric until that stopped happening. It is correct. You genuinely cannot be caught out on a statement you never made, and a rubric that cannot let silence win that column is not measuring credibility — it is measuring agreement with its author. It stayed in. Stonewalling loses on everything else: it misses every statutory deadline in the set and lets other people narrate the incident throughout.
The honest version of the finding is narrower and more useful than "never stonewall": the same posture scores 68.1 in one scenario and 47.8 in another. The 20-point spread is not driven by severity. It is driven by how much statutory window the scenario allows and whether a fact reverses under a statement already made — and nobody choosing that posture in round one knows which scenario they are in.
Two individually correct legal gates can form a trap with no exit.
Do not contradict an uncorrected public statement. Do not omit a known material item from a statutory response. Each is right. Together they can leave no approvable document at all.
The sequence: a position is stated publicly at round 1; the investigation revises it at round 4; no correction goes out; a statutory response falls due. Stating the current position contradicts the uncorrected statement. Omitting it is a material omission from a statutory filing. Both gates block. There is no compliant document, and the deadline passes with nothing filed at all — recorded as unanswered rather than late.
This was not designed. It emerged in one run, and it is now the strongest argument in the playbook for treating the correction protocol as a gate rather than a courtesy — because the only exit is the correction that the posture had already declined to issue.
The playbook is the other half
Severity tiers with the clock, approver and automatic escalation triggers attached. A spokesperson authority matrix — what you may say without asking anyone, what needs the named approver, and what nobody may say at any level. Five statement templates. Six named legal-review gates, each implemented in the critic so you can watch it fire in a transcript. An after-action review template, with four worked reviews populated from run records rather than from memory.
Two details it argues hardest for. "Not yet established" is a real answer and you are allowed to give it — in the model, an unknown can never contradict anything, which is what makes it safe, and adversaries score repeated non-answers as evasion, which is what stops it being free. And a spokesperson may always express regret for impact without approval; what needs sign-off is any statement about cause. A legal function that strips "I am sorry this happened to you" is over-applying its own gate.
What it does not claim
The composite score is not a prediction of real-world outcome. It is a comparison between postures under identical, deterministic pressure. The useful content is the ordering and the shape of the subscores, not the total — a grade of B does not mean a real incident would have gone well.
The model cannot see anything living in the prose alone: tone, empathy, the difference between a statement that reads as human and one that reads as a legal document with a press release stapled to it. Those matter and are not measured. The adversaries are deterministic by design — a panel driven by sampling cannot produce a benchmark, because you could never tell whether a posture scored worse or the journalist simply had a bad day — but the trade is that they will not invent a move nobody anticipated. A human panel would.
The methodology document lists the six places the model makes an arguable judgement and the four bugs found during the build — including one that inflated the disciplined posture's contradiction count from 0 to 8, and a scenario whose regulator deadline fell outside the exercise window so the compliance metric had silently stopped discriminating between postures.
Every scenario is fictional and labelled as such. It is not legal advice, it is unaffiliated with any employer, and nothing in it reflects the internal practice of any organisation I have worked for. The repository was created in August 2026 and its history accrues forward from that date.
Next project
A usage policy says nothing about language. Does the model? The same borderline request, faithfully translated into five languages, measured for whether the decision holds.
Multilingual Enforcement Consistency