Blog

All posts
John Damask · 2026-04-12
devlogsecurityarchitecture

Tests have been part of this project from early on -- each major feature shipped with its own test file, and I've been running them locally and in CI as I went. What I hadn't done was sit down and look at the suite as a whole. With launch a few days out, I wanted an honest read on it: what's covered, what isn't, and whether the tests that exist actually exercise the code they claim to.

The audit turned up three things I hadn't expected.

Fixture rot

A moto/Cognito version mismatch had broken the shared fixtures that a lot of tests depended on. Individually those tests passed when I wrote them; in the current pinned versions, they errored at setup. When I ran the full suite fresh, 426 tests collected and only 97 of them actually got as far as an assertion -- the rest failed during fixture setup before their test bodies ran. That's the kind of thing you don't notice until you run the whole suite against a clean environment, which I hadn't done in a while. Pinning the fixture versions and updating the few calls that needed new signatures cleared almost all of the collection failures.

Two real bugs in existing tests

The worst one was in the fraud-response tests. A kill-switch test was supposed to verify that Stripe refunds were not called when the kill switch was active. The assertion looked fine at a glance:

mock_refund.assert_not_called() if hasattr(mock_refund, "assert_not_called") else None

But mock_refund was the patch() context manager, not the mock object inside it. hasattr always returned False, so the assertion silently never ran. A regression in the kill switch would have slipped right past this test. The second bug was inverted boolean logic in a cache-safety test that made a cookie-detection assertion pass when Set-Cookie was present -- the opposite of its intent.

Those two tests had never failed because they weren't really testing anything. Easy to miss in a code review: the code reads left-to-right, looks like it's asserting something reasonable, and the suite is green.

Coverage gaps on the newest modules

Looking at what was covered and what wasn't, a pattern jumped out. The earliest modules -- upload, confirm, status, process -- had solid integration tests written alongside them. The modules that landed later -- the login and refresh flow, the admin 2FA path, the Stripe webhook dispute lifecycle, the gift code redemption flow, the CloudFront key rotation job, the security content screener -- had little or no coverage. Those are exactly the paths where a regression would cost money or leak data.

I'd been testing them manually during development, which is fine for getting a feature out the door but useless as a regression net six months from now. The audit made that gap visible in a way that doing feature-by-feature reviews hadn't.

What I added

Thirteen new test files, two bug fixes in existing tests, four files with strengthened assertions, and one file rewritten from source-greps (tests that just checked whether certain strings appeared in a file -- passes if the string is in a comment) to behavioral integration tests. Final count: 778 passing, 22 honestly skipped, 0 failures.

The new tests follow the pattern that was already working for the older modules -- real handler calls against moto-backed AWS, importlib.reload() to rebind module-level clients, external APIs mocked at their boundaries, everything inside the Lambda running real code.

Proving the new tests aren't garbage

Writing 681 new tests that "all pass" means nothing if the tests are trivially true. So I ran mutation tests against six of the new test files -- 22 deliberate regressions. I removed httpOnly from cookies, bypassed auth checks, disabled security screening, skipped credit deductions, changed boundary conditions by off-by-one. All 22 mutations were caught. The tests assert on real side effects -- DynamoDB state changes, S3 writes, credit balance deltas, HTTP status codes -- not on mocked return values.

Honest assessment

About 70% of the suite is integration tests against moto, 15% is static-analysis guards (CI trip-wires for cache safety and migration regressions), 10% is pure unit tests, and 0% is end-to-end tests. The integration layer is solid. The static-analysis tests are useful as guardrails but don't prove behavior -- they prove the code still looks like what I thought I wrote. The missing end-to-end layer -- an upload that goes all the way to a published page -- would require either a running server or a very elaborate moto setup with a mocked Anthropic response. That's a different kind of investment, and not one I wanted to take on the week before launch.

The real lesson from the audit wasn't "write more tests." It was "run the whole suite on a clean box periodically, or fixture rot will hide real gaps."