Enforcement Operations Simulator
A human-review operation is a system you can design and tune quantitatively, not a thing you staff up and hope about. Severity thresholds, routing rules, shift coverage, SLA targets, QA sampling, escalation triggers and appeal load all interact — and most of the interactions are non-obvious.
What it models
A 1,700-alert/day Trust & Safety enforcement queue, end to end: arrival bursts, classifier scores, routing thresholds, shift coverage, reviewer fatigue and anchoring, SLA escalation, QA sampling, and appeals — so that operating decisions can be argued with numbers.
Nothing in the reports is hand-written prose over hand-typed numbers. Every figure, table and sentence is interpolated from the same dataframes the simulation produces, so the analysis cannot drift away from the model. The 82 tests assert engine invariants and the direction of every headline finding.
Alongside the simulator sits the written operating model it exists to test: a severity matrix, decision tree, escalation paths, calibration cadence, an AAR template and a worked AAR.
Six findings
1. The aggregate SLA number is structurally incapable of showing you a failure
The modelled operation posts 98.6% aggregate attainment while missing its P1 target outright — 93.1% against a 95% commitment.
P2 and P3 are 88% of volume and always clear their generous targets, so they set the headline regardless of what happens in the bands that matter. Report by band or do not report.
2. A 5-point classifier precision drop costs 2.8 reviewer-hours a day — and that is the smallest of the three costs
The same regression costs 1.8 points of human decision accuracy, because reviewers are anchored by the score in front of them and a worse model produces more confidently wrong scores. Fund the 0.47 FTE and you have paid for the cheapest third of the damage.
3. The classifier threshold is not an SLA lever, and the 15-minute P0 SLA never breaks
Sweeping the bulk auto-close threshold moves utilisation across 33 points — comfortable to saturated — and shifts P0 attainment by 0.5 points. Strict priority with a shared reviewer pool insulates urgent work from bulk volume almost completely.
P0 is bound by how long the review honestly takes, not by the queue. You cannot buy P0 attainment with headcount or by shedding load, and a target set above that handle-time ceiling is one nobody can ever hit.
4. Uniform QA sampling cannot detect single-reviewer policy drift at any affordable rate; stratified sampling does it at 5%
Auditing one decision in two, uniformly, still reaches only 21.8% power to catch it within a week — while a stratified design at 5% reaches 94% with a 3-day median lag. The signal lives in about 6% of volume, and pooling dilutes it faster than it accumulates.
5. Coverage shape is a free lever, and it is being left on the table
The same 19 people moved from a 5/6/8 shift split to 2/6/11 take P1 — the binding band — from 93.1% to 94.5%. That is 1.4 of the 1.9 points the operation is short, at zero cost.
It does not close the gap alone, and the residual is a genuine headcount ask (2 reviewers, roughly $166,440/year) — but it should be sized against the re-cut roster, not the inherited one.
6. The obvious escalation rule breaks the thing it protects
Promoting aging items without a ceiling lets them reach P0, so imminent-harm reports queue behind merely-old ones. Adding a guard rank — one config line — is worth 25 points of P0 SLA at 12× burst load.
Why I built it
I ran police operations where deployment, resourcing and response-time commitments had to be defended with data rather than asserted. Enforcement queues are the same class of problem wearing different clothes: finite skilled reviewers, non-uniform severity, hard time commitments, and a strong institutional pull toward reporting the number that looks best.
Most of these findings are things an experienced operations lead half-knows. The value is in being able to put a number on them, and in the two that run against intuition — the QA sampling result and the escalation rule.
Next project
Ten APAC jurisdictions, one deadline board — and the measured finding that AI obligation extraction reliably finds the duty and misses the exemption.
APAC Regulatory Readiness