Last week I wrote about adding Stripe dispute handling -- three admin buttons a human had to click for every suspected fraud event. That's fine at a dozen charges a week. At the scale I'm targeting for launch, the latency between "Stripe fires a webhook" and "admin wakes up, opens the dashboard, clicks the button" is the difference between catching the fraud and eating a fifteen-dollar dispute fee.
This work rebuilt the fraud response path around a webhook-driven auto-response that refunds, deducts credits, freezes the user, disables their account, and emails them -- all within about a second of the webhook arriving. The admin buttons are gone. In their place: a dedicated Fraud section in the dashboard, two orchestrator functions (one async, one synchronous), a kill switch, a four-step unfreeze protocol, and a 450-line operating procedure.
The whole thing shipped in fourteen commits over about sixteen hours. Most of the build happened in the morning; the afternoon was end-to-end verification, which surfaced two real IAM bugs, one observability gap that turned into its own feature, and a handful of semantic bugs that only a human driving the dashboard could have caught.
The architecture decision
The first question was whether the webhook handler should do the refund-and-freeze work inline or enqueue to a worker. Inline is simpler: one Lambda, one code path, no queue. But inline has a hard twenty-nine-second budget before API Gateway times out and Stripe retries the webhook -- which would double-process the event. Async adds moving parts (a queue, a worker, a dead-letter queue, idempotency) but the webhook returns 200 in under a second, and you get a retry story for free when a downstream service has a transient failure.
Async won. The webhook does two things now: validate the Stripe signature, and enqueue. Everything expensive happens in a separate worker Lambda that reads from the queue and runs the orchestrator. On failure, the worker re-raises and the queue redelivers up to three times. Anything still failing lands in a dead-letter queue where a CloudWatch alarm pages on depth.
The kill switch
One SSM parameter controls whether the auto-response pipeline does anything. Both the webhook Lambda and the orchestrator check it before taking any irreversible action. When it's off, the webhook publishes a "kill switch active -- would have acted" alert and returns 200 without enqueueing.
Three scenarios drove this. A code bug starts freezing legitimate users and you need to stop the bleed. A runaway loop trips the rate alarm and you want to pause until you understand why. You're debugging and don't want production actions firing. The manual freeze flow deliberately ignores the kill switch -- admins can still freeze individual accounts when auto-action is off, which is the entire point of having both paths.
The four-step unfreeze protocol
Turning a freeze off is the most dangerous thing an admin can do. If the original fraud response was correct and an admin unfreezes without checking, you've undone your own protection. If it was wrong, an admin who rushes through without understanding why the freeze happened will probably make the same mistake next time.
The protocol enforces a pause. Four steps, each writing a timestamped row: Claim (one admin takes ownership), Verify (identity notes, minimum ten characters), Document (reason notes, minimum ten characters), Execute (re-enable the account, clear the frozen fields). The verify and document steps exist specifically to make the admin stop and think. The character minimums aren't about data quality -- they're about making "click through without reading" impossible.
Verification found real bugs
The morning session shipped the feature. The afternoon was an end-to-end verification checklist with four sections. Section A was dashboard smoke tests driven by Claude in a browser. Section B was Stripe-flow verification, reserved for a human operator because it exercises real test charges end to end. Section C was safety-mechanism tests. Section D was sign-off.
Section A found two IAM bugs that would have produced 500s for any admin opening those sub-views in production. The Dispute Review view filters disputes by non-key attributes, which means a DynamoDB scan, but the shared Lambda role policy only granted Query and single-item operations. Same bug on the Frozen Accounts view, which scans for users with a frozen timestamp present. Both were one-line permission additions. After finding the second one, I audited every scan call in the new handler modules and cross-referenced each against the role policy. Lesson filed: when a new handler module lands, grep for scan calls and check each one against permissions.
The DLQ test that didn't work as written
Section C was supposed to verify the dead-letter-queue path: enqueue a bad message, let it fail three times, watch the alarm fire. The original test plan said to use a message pointing at a Cognito user that doesn't exist, so the worker would fail on the disable step and the queue would redeliver.
Except it wouldn't. The idempotency guard writes the fraud event row on attempt one, before the Cognito step runs. Attempt two, the guard sees the row already exists and returns "skipped" -- which the worker treats as success. The message gets deleted. It never reaches the dead-letter queue.
The workaround was to force a failure before the idempotent write. DynamoDB rejects partition keys longer than 2048 bytes with a validation exception, so a synthetic message with a 2500-character event ID fails on every attempt, never writes a row, and gets all three retries plus the DLQ landing. The alarm fired on the first run. I wrote the trick into the operating procedure so the next operator running this test doesn't waste forty-five minutes on the same dead end.
Alarm emails that couldn't be acted on
Both CloudWatch alarms fired during verification and sent emails to support. The emails contained the raw AWS alarm metadata: threshold, dimensions, state history, metric namespace. Useful if you're a CloudWatch admin. Useless if you're the person paged at 2am who needs to know whether to flip the kill switch or just ride it out.
Fix was two-part. I added a runbook section to the operating procedure covering each alarm -- what it means, an immediate first-action command, an investigation flow, a decision tree, and an expected-recovery note. Then I updated the alarm descriptions in CloudFormation to embed a one-sentence summary plus a URL pointing at the runbook subsection. CloudWatch renders the alarm description verbatim in the SNS email, so the oncall now gets the summary, an immediate action, and a clickable link to the full procedure -- all in the alert itself, zero extra clicks.
Note to self: "what does the person paged at 2am need to see" is a deploy-time concern, not a follow-up task. Next alarm I add, I'll write the runbook section first.
Section B: bugs you can only catch by walking through
Section A and C were mechanical. Section B was a human walking through the actual dashboard with actual test charges and actual eyes on the output. It surfaced several bugs, and all of them were the kind that automated tests cannot easily find because they live in the gap between "code did what it was told" and "code did what a human would have asked it to do if they'd known to ask." Three are worth writing down.
The double deduction. After a successful test, I checked the user's credit balance and found it off by exactly the pack size. The transaction log showed two refund deductions one second apart for the same charge: one from the fraud worker, one from the charge.refunded webhook handler. The fraud worker creates a Stripe refund, which fires charge.refunded back at the webhook, and the webhook handler doesn't know "this refund came from me" so it deducts again. Different dedup keys on each path meant the idempotency guard never caught the collision.
Going back through the historical transaction log, the same pattern had been silently running since the original dispute work shipped weeks earlier. The bug had been quietly overcharging users the whole time, and nobody had noticed because nobody was running the math on every transaction row. Fix: unify all four refund paths on a canonical dedup key derived from the payment intent -- first path to insert wins, subsequent paths see the existing row and skip.
The inquiry that triggered an auto-freeze. One of the test cards was supposed to exercise the non-fraud dispute path: create a dispute, have the admin click "Do Not Freeze," verify nothing else happens. Instead, the user got auto-frozen. The admin email said the user was frozen via auto-fraud with reason "Fraudulent dispute."
That test card, in Stripe test mode, actually produces a fraudulent inquiry rather than a full dispute -- a warning_needs_response event where the issuer is fishing for information before deciding whether to formally dispute. No funds have been withdrawn yet. The inquiry often closes on its own without ever becoming a dispute. But the dispute-created handler was branching on the event's reason field alone, so the inquiry went straight to the auto-freeze path. By the time the inquiry closed harmlessly, we'd already refunded, deducted credits, frozen the account, and sent a fraud suspension email -- all for a warning that never became a real dispute.
The fix added a status check at the top of the handler: if the status begins with warning_, route through manual review regardless of the reason field. Inquiries are recorded for admin awareness but never auto-action.
The restore-credits checkbox that was always wrong. The unfreeze workflow had a "restore credits" checkbox at the Execute step that looked up how many credits were deducted during the freeze and granted them back. The intent was false-positive remediation. But in every real case where credits had actually been deducted, a Stripe refund or dispute chargeback had also returned money to the user -- so restoring credits on top was always double compensation. The test user's own history showed two such pairs, each worth one pack's worth of free credits. Fix: remove the option entirely. If an admin wants to grant apology credits after an unfreeze, the standalone credit adjustment form exists for exactly that and is audit-logged separately.
The thing Section B proved
Section A caught the IAM bugs because those produced visible 500s. But every bug Section B found was invisible to Section A. The double deduction was a math error you only notice if you're watching the transaction log. The inquiry misrouting looked like "user is frozen even after Do Not Freeze," not a stack trace. The restore-credits bug was a semantic mismatch between what the code was written to do and what a reader would assume it should do.
Those only surface when a human walks the flow as a real admin would and says "wait, that's not what that should say." The async architecture's complexity showed up in Section B not as broken state machines but as assumptions about "who does what when" that only held for the happy path.