Blog

All posts
John Damask · 2026-03-09
devlogsecurity

The Incident

The site hadn't even launched yet, and someone found one of my API endpoints. The giveaway was a database record referencing an ID that didn't exist — someone had either been probing the API or noticed the endpoint in network traffic and hit it directly.

The immediate fix was simple: validate that the referenced ID actually exists before accepting the request. But that only plugged one hole. It was time for a full security audit.

What the Audit Revealed

I mapped out every endpoint by method, authentication, rate limiting, and input validation. The results were sobering.

The worst finding: one of my administrative endpoints had zero protection. No authentication, no rate limiting, no input validation. Anyone who could guess or enumerate IDs could have caused real damage. Several other endpoints were exposing operational metrics and cost data to anyone who asked — not destructive, but not great either.

The existing protections were inconsistent. Some endpoints had per-IP rate limiting, but the implementation was duplicated across handlers with no shared code. The upload endpoint had a global daily quota but no per-IP limit, meaning a single bad actor could burn the entire day's quota for all legitimate users. And one endpoint that triggers an AI API call — real money per request — had no rate limiting at all.

This is one of the gotchas of building apps via Agentic Engineering. Easy to deal with, now that I'm aware.

The Fix

I implemented a three-tier protection strategy:

Tier 1: Shared rate limiter. I extracted the rate limiting logic that already existed in two handlers into a reusable module. It uses DynamoDB atomic counters with TTL — each IP gets a counter per time window, and the counter auto-expires. I applied it to every write endpoint and every endpoint that costs money to invoke, each with appropriate limits.

Tier 2: Token authentication for admin endpoints. Administrative and destructive endpoints now require a secret token in the request header, verified against AWS Secrets Manager. The deploy scripts auto-generate the token on first deployment. This follows the same pattern I was already using for my blog API — I just hadn't applied it everywhere it was needed.

Tier 3: Infrastructure updates. Updated my CloudFormation template to wire up the new environment variables and IAM permissions, and updated the CORS configuration to allow the new auth header.

I also updated my local development server to mirror the same protections, so I can test the auth and rate limiting flows without deploying to AWS every time.

Deployment Surprises

Two bugs surfaced during deployment that are worth sharing:

First, my IAM policy for Secrets Manager listed specific secret ARN patterns but didn't include the new admin token secret. Lambda calls to retrieve the secret failed with an access denied error. AWS Secrets Manager appends a random suffix to secret ARNs, so policies need a wildcard pattern — easy to forget when adding a new secret.

Second, my deploy scripts package Lambda code by copying specific files into a zip. The new shared module wasn't in the list, so four Lambda functions crashed on import. A one-line fix, but a reminder that manual packaging is fragile. Every new shared module needs to be explicitly added to the file list.

Lessons Learned

  1. Audit everything before launch, not after. I should have done this systematic endpoint review weeks ago instead of waiting for someone to find a gap.

  2. Consistent patterns beat ad-hoc fixes. Having rate limiting in two handlers with duplicated code meant it was easy to forget on new handlers. A shared module makes the right thing easy.

  3. Test deployments catch infrastructure bugs. Both deployment issues (IAM policy and packaging) were caught in the test environment before touching production. Having separate test and prod environments paid for itself immediately.

  4. Think like an attacker. The delete endpoint being wide open was the scariest finding. It's easy to focus on the happy path and forget that every endpoint is a potential attack surface.

The site now has proper per-IP rate limiting on all write endpoints, token authentication on all admin endpoints, and consistent input validation everywhere. Not bulletproof — nothing is — but a much better starting position for launch.