Blog

All posts
John Damask · 2026-05-16
devlogawsapiarchitecture

I added Featured Collections to the homepage. These collections are curated from the best papers on particular topics, translated for particular audiences.

Currently, only admins can create collections and populate them with papers and since the app only supported processing one paper at a time, seeding the collections was a grind. It's also not as cost-efficient as it could be.

Anthropic's API supports batch processing, which is exactly what I needed: submit several papers at once, get a batch ID back, results land within 24 hours (most within an hour), and the bill is half.

What I built

Three new Lambdas form a fan-out → batch-submit → poll pipeline:

Upload Lambda. POST /api/admin/bulk-translate/upload takes a list of files, mints one DynamoDB row per file in the existing jobs table with status="awaiting_upload", mints one metadata row in a new BulkBatchesTable, and returns N pre-signed S3 PUT URLs. The browser uploads each PDF directly to S3 under a new admin-bulk/ prefix.

Submit Lambda. POST /api/admin/bulk-translate/submit does the heavy lifting: for each uploaded PDF it runs Anthropic's Haiku-based content screen, records token counts and paper classification on the job row, builds an anthropic.types.messages.batch_create_params.Request[] payload, and fires client.messages.batches.create(requests=[...]). Then it flips the batch metadata row and all surviving job rows to processing, conditional on their being in the expected starting state.

Poller Lambda. EventBridge fires it every five minutes. It queries a sparse (status, created_at) GSI for live batches, calls batches.retrieve, and when one reports processing_status="ended" it streams batches.results and routes each per-request outcome to a shared write_job_complete helper that lands the same row shape the single-PDF path writes.

That shared helper was the most valuable side effect of the whole arc. The single-PDF Process Lambda already had eighty lines of "build the complete-state job row" logic touching about thirty fields. The poller needed to write the same row. Field-parity drift between two writers is exactly the kind of bug that's invisible at runtime -- a missing column on bulk-processed papers means the dashboard doesn't show it, no error, no alarm. Factoring write_job_complete(table, job_id, *, parsed, mode, llm_config, screen_model, ...) while both call sites were fresh let me add a field-parity test that fakes both paths and asserts the row shapes are identical.

The regression review that earned its keep

Halfway through the implementation, after fixing two bugs the first end-to-end test surfaced (papers landing in the wrong gallery because I'd resolved the admin's user ID from the wrong table; bulk jobs being filtered too aggressively out of the operations dashboard), Claude told me, "I have a small paranoia that the regression risk is contained, but you should watch CloudWatch logs and alarms for the next hour to be safe." I said that was ridiculous, "Paranoia? This is methodical. Your suggestion to watch CloudWatch logs and alarms is not methodical."

Claude admitted as much - watching logs is what you do after you've confirmed a diff doesn't break anything -- not as a substitute for confirming it. I have it run the phone-a-friend skill to discuss how to handle this and asked, "could the changes in this PR have broken the single-PDF customer flow that has nothing to do with bulk-translate?"

The friend surfaced one thing Claude hadn't checked. The publisher module that uploads generated HTML to S3 had this at the top:

PUBLISH_BUCKET = os.environ.get("PUBLISH_BUCKET", "<staging-bucket-name>")

The default falls back to the staging bucket -- the one with a one-day S3 lifecycle policy. If any deployed Lambda ever imported this module without PUBLISH_BUCKET set in its environment, the publish would succeed (200 from put_object), the page would serve correctly for a day, and then vanish. Silent customer data loss with no alarm.

In this PR, nothing Claude added imported that module, so the bug wasn't active. I asked, "Under what conditions could this realistically happen?"

The honest answer is that today's deploy was fine, but the realistic failure mode is "next person adds a new publisher consumer Lambda, forgets the env var, and a customer's just-generated paper 404s the next day." Not the worst-case silent-deletion-after-three-days Claude initially framed it as. Still serious enough to merit prevention, just less dramatic than Claude made it sound.

The fix landed in the same PR. Both publisher modules switched from os.environ.get(..., default) to os.environ[...] -- fail-fast on Lambda init if the env var is missing. Plus two new tests: an AST walk that fails the build if anyone re-introduces a bucket-name default pointing at a real bucket, and a CFN parser test that grep-checks every Lambda that imports a publisher and asserts the matching AWS::Lambda::Function resource in CloudFormation has PUBLISH_BUCKET in its environment.

The 504 that wasn't a failure

The feature shipped to prod and I tested it for real: twelve seminal ML papers, the actual curation pass I built the feature for. The dashboard returned Submit failed: HTTP 504.

CloudWatch told a different story. The submit Lambda had completed cleanly with a duration of 113 seconds, the Anthropic batch had been created, and the poller was already pulling results while the client saw a Gateway Timeout.

API Gateway's HTTP API has a hard 29-second integration timeout. The client got 504 at twenty-nine seconds while the Lambda kept running for another minute and a half. The submit endpoint runs Haiku content-screening on each PDF sequentially before calling batches.create: a small paper screens in three seconds, a 75-page paper truncated to 25 pages screens in twenty-five. Twelve papers stacked end to end blew past the API Gateway window.

The fix is architectural and it's exactly the same shape as the rest of the bulk path: the submit endpoint should become a thin gate that flips state and async-invokes a worker Lambda, returns 202 immediately, and the UI polls a new status endpoint. The single-PDF flow already uses this Confirm → Process pattern for the same reason; applying the same instinct to bulk submit would have caught this at planning time. Filed as a separate issue, scheduled for the next session.

Claude's Lessons (that it will forget)

Regression checks aren't log-watching. "I'll keep an eye on it" is what you say after you've confirmed correctness, not instead of confirming it. When someone asks "are you sure?" about a non-trivial change, the right answer is a structured review by a fresh pair of eyes. The phone-a-friend regression pass cost twenty minutes and surfaced a latent bug that would have eventually caused a real customer incident.

Severity framing matters. Describing a footgun as "silent customer data loss" when the realistic failure is "loud 404 within hours" is technically accurate but practically misleading. Write the worst realistic case, not the worst conceivable case, and label which is which.

API Gateway's 29-second ceiling drives architecture, not just rate-limit behavior. Anywhere I'm migrating a synchronous primitive to an async-batch one, I need to audit every per-file step against the time budget, not just the headline API call. The plan doc for this work even mentioned "instead of one SQS message per job, submit all PDFs in one shot" -- and that's true at the batches.create granularity, but the per-file screening fan-out is still per-file work. The right time to find that out is the planning step, not the first real production workload.

The PR merged. Prod is running batches. The next workload pass will use the async submit refactor, and I'm now ready to put some collections of great science together.