Now I Get It! started as a fully public site. Upload a PDF, get a URL, share it. The gallery showed everyone's work. That was fine for a demo; it was not fine for a product people would actually pay for. People want a place where their own papers live and have the option to share if they want.
User accounts (issue #24) shipped earlier in the week. Each user got their own gallery page, and new uploads now live in a per-user S3 path: pages/users/{user_id}/{job_id}.html. The S3 bucket policy blocks public access to that prefix entirely. But when I first wired this up, private pages were served via presigned S3 URLs -- AWS credentials in query parameters, expiring after an hour, unbookmarkable, bypassing CloudFront. That was ok but I kept at it to see if there was a better design.
The design I almost built
My first plan was Lambda@Edge plus CloudFront KeyValueStore. The idea: a Lambda@Edge function on the pages/users/* behavior would check KVS for "is this job publicly shared?" and either bypass auth or validate signed cookies. No files would ever be copied between S3 paths; visibility was just a KVS flag.
But that wasn't right - Lambda@Edge runs on every request, including cache hits. A private page has exactly one viewer -- the owner -- so CDN caching provides zero benefit. All the edge compute would do is add 5-50 ms of latency on warm hits and 100-500 ms on cold starts, for nothing. It's also operationally painful: no environment variables (config has to be baked into the code or pulled from SSM at cold start), logs scatter to whichever region the function happened to execute in, deploys take 5-15 minutes to propagate, and you can't delete a replicated function until all its replicas are cleaned up -- which takes hours.
The KVS plan also introduced a sync problem I didn't need: visibility state would live in DynamoDB (the shared_via_link flag on the job record) and in KVS. A failed KVS write after a successful DDB update would leave the two out of sync, and the only way to detect the drift would be a reconciliation job.
The thing I was trying to avoid -- copying the HTML file to pages/{job_id}.html when a user toggled "share" -- is three API calls on a rare user action. Toggling visibility is not a hot path. Copying a 50-500 KB HTML file takes about 100 ms.
What shipped
CloudFront has native support for signed cookies via TrustedKeyGroups on a cache behavior. I restructured the distribution with two origins: the existing S3 website endpoint keeps serving public content, and a new S3 REST API origin with Origin Access Control serves the pages/users/* path. The pages/users/* behavior has TrustedKeyGroups set, so CloudFront validates signed cookies natively. No custom edge code.
On login and token refresh, the auth session Lambda reads an RSA private key from Secrets Manager, builds a CloudFront custom policy scoped to the user's path (/pages/users/{uid}/*), signs it with RSA-SHA1, and returns the three cookie values -- CloudFront-Policy, CloudFront-Signature, CloudFront-Key-Pair-Id -- in the API response body. The frontend sets them via document.cookie. They persist for 30 days.
Public sharing still copies the file to pages/{job_id}.html, served by the public pages/* behavior with no auth. Unsharing deletes the copy. Simple.
Rotating the key
The signing key has to rotate. I wrote a single Lambda that handles three event types: CloudFormation Custom Resource (bootstrap during stack creation), EventBridge schedule (rotate every 30 days), and manual invocation (if something goes wrong). The Custom Resource creates the initial RSA 2048-bit key pair and returns the PublicKey ID so CloudFormation can populate the KeyGroup. The private key goes to Secrets Manager; the key pair ID goes to SSM Parameter Store. After bootstrap, the rotation Lambda pulls the KeyGroup ID from SSM (it can't get it from an env var because that would create a circular dependency with the Lambda's own CloudFormation role).
Takedowns get simpler
The old DMCA takedown flow was built before private-page access existed. It would back up the original HTML, overwrite it with a takedown notice, and delete any public copy. Even the page's owner saw the takedown notice when they opened their own link -- which made sense when there was no such thing as an owner.
With signed-cookie access, the takedown flow can now just remove the public copy and set a takedown_locked flag on the job. The owner's private page stays intact. Restore just unsets the flag -- no S3 copy/restore needed, because the content was never removed. The share toggle checks takedown_locked and returns 403 if set. The my-gallery page shows a banner ("Shared page removed per copyright request") with a link to the complaint status page, and disables the share toggles.
Legacy pages from before user accounts have no user_id, so they retain the old overwrite behavior -- there's no private path to preserve.
An audit trail
I added a file_history field to the job record: an append-only list of every file lifecycle event (created, made_public, made_private, takedown_public_removed, takedown_restored, cancel_backup, cancel_overwrite). Each entry has a timestamp, the action, the S3 key involved, the actor, and an optional note. One DynamoDB list_append per event. Useful for debugging, for copyright disputes, and for account-cancellation sweeps.
Three bugs from verification
The first: browsers rejected the CloudFront cookies when I initially set them via Set-Cookie response headers. The auth API lives on execute-api.amazonaws.com; the pages live on nowigetit.us. A response from one domain can't set cookies for a different domain. The fix was to stop setting the cookies server-side at all -- return the three values in the JSON body, and let the frontend set them via document.cookie, which is same-origin from the browser's perspective because the script is running on nowigetit.us.
The second: the S3 bucket policy used StringNotEquals with arn:aws:cloudfront::account:distribution/* to exempt the OAC from the deny rule. StringNotEquals treats * as a literal character, not a wildcard. The condition never matched; S3 denied every CloudFront OAC request. Switched to the specific distribution ARN.
The third: cookie clearing didn't work on some browsers. document.cookie = "name=;Max-Age=0;Path=/;Secure;SameSite=None" silently failed. Adding spaces after semicolons fixed it: "name=; Max-Age=0; Path=/; Secure; SameSite=None". Twenty years of the web and there are still corners like this.
PR #164. 29 commits, 39 files, about 5,800 lines. Private-by-default is now the whole product's posture, and the public gallery is an opt-in, not an opt-out.