Blog

All posts
John Damask · 2026-04-12
devlogperformanceawsarchitecture

The public gallery on Now I Get It! queries DynamoDB for all completed papers that are marked public. When I first built it, every paper was public by default, so the query was basically "give me everything with status = complete" -- fast, simple, and correct. Then I added user accounts with private-by-default papers, and the query became "give me everything with status = complete, then throw away the private ones." That's still fast at 300 papers. It would not be fast at 30,000.

The problem is the FilterExpression. DynamoDB charges you for every row it reads, not every row it returns. A filter that matches 5% of the partition means you're paying for 20x the reads you actually need. As private papers become the dominant share of the table -- which they will, once real users are uploading -- the gallery query burns more RCUs per page load while returning fewer results. The latency degrades silently until one day it starts timing out.

Sparse GSI

DynamoDB has a feature called a sparse GSI. If you define a Global Secondary Index on an attribute that only exists on some items, the index only contains those items. Everything else is invisible to queries against it.

I added a PublicGalleryIndex GSI keyed on public_gallery_pk and publish_ts. Whenever a paper becomes publicly visible -- either via the "make public" toggle or the "share via link" option -- the writer Lambda sets public_gallery_pk = "public" on that item. When a paper goes private, the writer removes the attribute. The index tracks only the public set, always, automatically. Querying it returns exactly the papers the gallery should show, with zero wasted reads.

Three separate writer paths had to be updated: the visibility toggle, the account-cancel loop (which unpublishes all of a user's papers), and the admin takedown flow. Each one got a REMOVE public_gallery_pk, publish_ts added to its existing DynamoDB update expression. The visibility toggle also needed a race guard -- an attribute_not_exists condition on the SET path -- so that a backfill script and a live user toggle can't clobber each other's timestamps. Seven extra lines of code for an edge case that probably fires once a year, but the alternative was silently corrupting sort order in the gallery.

Caching the public reads

The second lever was CloudFront edge caching. The gallery and library endpoints were hitting Lambda on every single request, even though the data only changes when someone publishes or unpublishes a paper. A 60-second cache at the edge would eliminate virtually all origin traffic for steady-state browsing.

The original plan was to flip the entire api/* CloudFront behavior from CachingDisabled to a custom cache policy. A reviewer caught the risk: if any authenticated endpoint ever accidentally emitted a Cache-Control: public header, CloudFront would cache that response and serve it to every subsequent user. One bug, one cached response, and you're leaking user A's profile data to user B.

The fix was to stop thinking about it as "cache everything and hope the Lambdas don't mess up" and start thinking about it as "only these specific paths are allowed to be cached, and everything else is architecturally unreachable." I added two exact-match CloudFront behaviors -- one for api/gallery, one for api/library -- placed before the api/* fallback. Each has its own origin request policy that strips Authorization headers and cookies before the request reaches the Lambda, so even if the Lambda code tried to read caller identity, it would see nothing. A response headers policy strips Set-Cookie on the way back out. The api/* fallback stays CachingDisabled, completely untouched.

What actually improved

After deploying everything to the test environment, the gallery still took about two seconds on first load. I measured the full request chain:

The second time I loaded the page, the API call dropped to 100ms because CloudFront served it from cache. That's a real 13x improvement on the API leg. But I was optimizing step 3 when the user experience was bottlenecked on the combination of steps 1 through 3. The sparse GSI and the edge cache are doing exactly what they're supposed to do -- the query is scale-proof and repeat visits are fast. But at 354 items in the test database, the old query and the new query perform identically. The payoff is insurance against growth, not a visible improvement today.

The honest takeaway: I built the right infrastructure for a product that's about to scale, but if I wanted the page to feel faster right now, I'd need to attack the JavaScript waterfall -- inline the API call so it starts before the scripts finish loading, or server-side render the first page of cards into the HTML. Different problem, different issue.

I also added skeleton loading cards to the gallery, landing page, and profile -- shimmer animations that render immediately while the API call is in flight. Doesn't change the actual load time, but the page isn't blank while you wait.