Blog

All posts
John Damask · 2026-04-26
devlogawsarchitecturesecurity

When Now I Get It! launched on April 13, every new user got auto-confirmed at signup. That was a hack, and it was deliberate. Cognito's default email path -- the one new accounts get out of the box -- caps at 50 messages per day. With 50/day, a real verification flow was a non-starter; we shipped a pre-signup Lambda that auto-confirmed everyone, and put a note in the launch checklist to come back to it once AWS approved our SES production-access request. They approved it on April 15. Tonight the workaround came off.

The mechanical change is small. Wire EmailConfiguration on the Cognito user pool to the SES identity, delete four lines from the pre-signup Lambda, simplify one button handler in the frontend. New signup capacity: 50,000/day with a 14/second burst. A thousand-fold headroom over what the launch hack was working around.

Three things a phone-a-friend caught before the deploy

Before filing it as work I ran the plan through an Opus second-opinion debate.

SES identity policy auto-attach is console-only. AWS's documentation suggests that when you set EmailConfiguration on a Cognito pool, the IAM identity policy granting cognito-idp.amazonaws.com the right to call ses:SendEmail on the SES identity gets attached automatically. It does -- but only when you click through the Cognito console. Setting EmailConfiguration via CloudFormation, CLI, or API leaves the policy unattached. I verified empirically: aws sesv2 get-email-identity-policies --email-identity nowigetit.us returned Policies: {} on prod, even though the SES identity itself had been verified for nearly two weeks. Without the policy, the first signup after the CFN flip would fail with AccessDenied and Cognito would silently swallow the error -- the user lands at the verify-code modal with no email arriving anywhere. Fix landed as bootstrap_cognito_ses_policy() in the deploy script, called before the infra deploy. Idempotent (chooses create-email-identity-policy vs update-email-identity-policy based on existence). One footgun: SES identity policies don't accept the standard AWS:SourceAccount / AWS:SourceArn confused-deputy conditions. I matched AWS's own auto-attached policy, which uses Principal + Resource scoping only.

The signup allowlist was still ["*"]. Today, with auto-confirm in place, that's harmless because Cognito sends no email at all. The moment SES wiring lands, every signup triggers a real send. A bot enumerating against aaa@aaa.com, bbb@nope.invalid could spike SES bounce rate past AWS's 5% threshold or complaint rate past 0.1% in an afternoon -- at which point SES auto-pauses sending and verification email stops working until AWS accepts an appeal. Bounce/complaint handling lives in a separate issue; the call was to ship this one with the allowlist still open and fast-follow the protection. An accepted gap, recorded.

CFN-flips-Cognito-instantly vs deploy ordering. The original plan had CFN first, then backend, then frontend. CFN-first means Cognito is talking to SES while the old auto-confirm Lambda is still live. Functionally fine -- auto-confirmed users never trigger a Cognito send -- but a deploy failure mid-sequence could leave a window of weirdness. Decision: do the whole thing inside one deploy invocation so all four changes flip together.

The test rehearsal that exposed real IaC drift

The test environment had no SES setup at all. The plan had assumed a test domain identity for test.nowigetit.us already existed; it didn't. Worse, a grep of the CloudFormation template for anything SES-related returned zero matches. The prod SES domain identity and its three DKIM CNAMEs in Route 53 were created by hand around April 13–15 and never made it into IaC.

Two follow-ups came out of this. One issue to codify the SES identity and DKIM records into CloudFormation, using AWS::SES::EmailIdentity with easyDKIM and outputs that populate the Route 53 records (no hardcoded tokens). And a manual one-time mirror in the test account: create the identity, get fresh DKIM tokens, push three CNAMEs into the test Route 53 zone. Verified within a minute -- DKIM propagation was fast on this one.

The pre-flight script for SES became a real deliverable rather than throwaway: account production-access check, identity verification, DKIM CNAME resolution via dig, and a live send to a verified recipient. Tokens are discovered from the SES identity itself rather than hardcoded, so the same script works against either account. Both pre-flights returned 7-of-7 green.

Blue/green chicken-and-egg, and a sharp correction

Test deploy bombed on PreSignUpPermission CREATE_FAILED -- Cannot find alias arn :live. Pre-signup is a brand-new Lambda in test (it was added to CloudFormation during the launch-night work but never deployed there), and its AWS::Lambda::Permission resource references the :live alias. Per the blue/green design, all consumers -- API Gateway integrations, Cognito triggers, SQS event-source mappings -- invoke :live rather than $LATEST. Aliases are deploy-script-managed (NOT CFN-managed) to avoid drift, and they're created in a phase that runs after the infrastructure deploy. For functions that already exist, this is fine. For a brand-new function added in the same CloudFormation pass, the Permission resource tries to attach to a not-yet-existing alias and the whole stack rolls back.

I proposed the wrong fix first: drop the :live qualifier from the Permission, granting Cognito invoke rights on the unqualified function ARN. It would have worked in one deploy, with a tiny ~1-second permission-replacement window. The user pushback was unambiguous: blue/green is non-negotiable, the :live qualifier on Permissions is part of the contract that consumers only invoke the live version, and broadening to the unqualified ARN silently grants invoke rights to any version or alias. That isolation is the point. Never weaken :live scoping to dodge alias chicken-and-egg.

The right fix is a standard alias-bootstrap pattern. The :live-qualified Permission gets wrapped in a CloudFormation Condition tied to a *PermissionEnabled parameter (default true). The deploy script detects new Lambdas (function or :live alias missing in AWS) and orchestrates a transparent three-step deploy: pass 1 with the gate set to false creates the Function with no Permission, a bridge step publishes v1 and creates the :live alias, pass 2 with the gate restored to true creates the Permission against the now-existing alias. Steady-state deploys (no new Lambdas) skip all of this and run a single CloudFormation pass.

Result

Prod was uneventful once the orchestration bugs landed. 39 Lambda functions code-updated, version-published, and alias-flipped in 23 seconds. Total deploy duration about 8 minutes. Smoke verified end-to-end: an actual signup from a fresh email address, real verification code arriving from noreply@nowigetit.us, code entry routing to sign-in, sign-in working. The SES account check on prod after deploy: ProductionAccessEnabled: true, Max24HourSend: 50000.0, MaxSendRate: 14.0, EnforcementStatus: HEALTHY.

The constraint that drove the auto-confirm workaround at launch is gone. The unprotected window from ["*"] allowlist + new SES capacity stays open until the bounce/complaint pipeline lands -- which is the next post.