Blog

All posts
John Damask · 2026-04-26
devlogawsarchitecturesecurity

The previous post left an unprotected window: prod's signup API is unauthenticated, the allowlist is wide open, no CAPTCHA, no per-IP rate limit. A bot enumerating against bot$i@gmail.com can drive the SES account-level bounce rate over AWS's 5% threshold, or complaint rate past 0.1%, in an afternoon. At those numbers, AWS auto-pauses sending and verification email stops working for everyone until AWS accepts an appeal. This issue closes that window on the detection side. Prevention -- CAPTCHA, rate limits, disposable-domain denylists -- stays filed separately.

Architecture

Configuration Set on the Cognito user pool sends bounce and complaint events to an SNS topic, which fans them out to an SQS queue, which triggers a Lambda. The Lambda writes to a SuppressedEmailsTable and conditionally calls AdminDisableUser on the Cognito user. A DLQ with maxReceiveCount=3 catches poison messages. CloudWatch alarms watch bounce-rate above 2 percent, complaint-rate above 0.05 percent, and DLQ depth above zero -- all paging a new alerts topic that goes to alerts@nowigetit.us.

The 2/0.05 thresholds are deliberately well below AWS's 5/0.1 pause thresholds. By the time AWS pauses sending, we've already been paged twice and have time to dig in. The DLQ alarm is the urgent one: every DLQ message represents a bounce or complaint that failed all three retries -- meaning that address was not recorded and (for complaints) the user was not disabled.

The new alerts topic is parallel to the existing dispute-events topic, which pages support. It's the precedent for low-urgency operational alerts that shouldn't compete with fraud-dispute pages.

The complaint-disable design call

The non-obvious design decision is what to do on a complaint. A complaint means the recipient hit "Report spam" in their email client -- abuse from their perspective.

For an UNCONFIRMED Cognito user, a complaint almost certainly means a bot signed up with someone else's address, or a real human got an unexpected verification email and flagged it. Either way the right move is to disable the account and stop the pattern.

For a CONFIRMED user, the same complaint is a paying customer who got annoyed. Locking them out of their account would be a worse failure mode than the spam complaint itself.

The Lambda calls AdminGetUser, checks the user's status, and only calls AdminDisableUser if the status is UNCONFIRMED. SES's account-level suppression list takes care of "stop emailing them" automatically for both populations.

Hitting the IAM policy ceiling, twice

Test deploy Pass 1 hit Maximum policy size of 10240 bytes exceeded for role nowigetit-lambda-role. The shared Lambda role was already at around 9100 bytes of inline policy across 19 named blocks; adding two new ones -- SQS access for the new queue and DynamoDB access for the new table -- pushed it over.

I extracted both into a single managed policy attached to the role. Managed policies have a separate quota -- 6KB per document, attached separately from inline -- which gives back about 1KB of inline headroom and a fresh quota for the new grants.

Pass 1 retry hit it again, more subtly. I'd added one Cognito action (AdminGetUser) to the existing inline block so existing handlers using the disable/enable actions could share the action grant. That single new line was enough to push back over. Moved it into the managed policy alongside the new SQS and DynamoDB grants -- the new SES events Lambda is the only consumer anyway, so scoping it there is cleaner than putting it in the shared inline block.

New rule of thumb for this codebase: anything new gets a managed policy. The inline Policies: array is treated as full.

Rollback gotcha: deletion protection

Both Pass 1 failures triggered automatic stack rollbacks. The rollback couldn't delete the new SuppressedEmailsTable because of DeletionProtectionEnabled: true, and the stack got stuck in UPDATE_ROLLBACK_COMPLETE_CLEANUP_IN_PROGRESS retrying the delete on a loop.

Standard fix from the previous time this happened: aws dynamodb update-table --no-deletion-protection-enabled on the orphan, then delete-table, then CloudFormation's next retry succeeds and the stack reaches UPDATE_ROLLBACK_COMPLETE. Compliance-tier tables get deletion protection for production safety, and the rollback dance on first deploy is the standard cost of that protection.

Smoke testing with the SES simulator

AWS provides a set of simulator addresses on simulator.amazonses.com that produce specific event types without a real recipient.

For bounce: bounce-test-296@simulator.amazonses.com returned 5.1.1 Unknown command -- the simulator only treats exact local-parts as special (bounce, complaint, success, ooto, suppressionlist) plus their +suffix sub-addresses. Functionally identical for our purposes -- a generic hard bounce. SNS notification at alerts@, DDB row with source=bounce, Cognito user untouched.

For complaint: complaint+unconf-296@simulator.amazonses.com produced a real Complaint event with complaintFeedbackType=abuse. DDB row with source=complaint and cognito_user_disabled=true. Cognito user state went to UNCONFIRMED, Enabled=false. DLQ depth stayed at 0 throughout.

The CONFIRMED-no-disable branch is pinned by a unit test rather than end-to-end. Triggering it live requires admin-confirm-sign-up plus a forgot-password trigger to generate the second SES send -- more setup than the test pays back.

What's protected now

Detection: bounce and complaint events get recorded with full payload to a DynamoDB table with PITR + deletion protection on, last-event-wins by email-address PK. The alerts@nowigetit.us mailbox gets a copy of every raw event for visibility.

Reaction: hard bounce -- no Cognito side effect (SES account-suppression handles "stop emailing them"). Complaint with UNCONFIRMED user -- Cognito disabled. Complaint with CONFIRMED user -- account left alive.

Paging: bounce-rate above 2 percent and complaint-rate above 0.05 percent both page well below AWS's pause thresholds. A non-empty DLQ pages immediately because it represents events that failed all three retries.

Resilience: SQS Standard with maxReceiveCount=3 on a 60-second timeout / 120-second visibility absorbs transient throttles or DDB write hiccups; failures land in the DLQ where the depth alarm catches them within a minute.