Landing Page Overhaul
The landing page needed to do a better job of explaining what the app actually does. Tackled this as a five-part epic, working in a git worktree on a 8-improve-landing-page branch to keep main stable.
The layout was tightened up -- top-anchored instead of vertically centered, compact horizontal drop zone (icon left, text right), tighter margins throughout so everything from title to upload button fits on a 1080p screen without scrolling. The old single "description" line was split into a proper tagline ("Scientific articles, explained." at 1.3rem) and a description sentence below it (at 1rem) explaining the upload-to-interactive-page workflow. A subtle "Works best with files under 10 MB" hint was added inside the drop zone in muted text -- soft guidance, not a hard limit.
The upload zone now grays out with pointer-events: none during processing so users can't accidentally trigger a second upload. A "Recreate" button appears next to the Copy button when processing completes, letting users re-run the same PDF through Claude to get a different explanation. This required a new backend endpoint (DELETE /api/job/{job_id}) to clean up the old DynamoDB record before re-uploading, preventing duplicate entries in the library carousel. The delete endpoint shipped as a new Lambda (lambda_delete.py) with its own API Gateway route, and the IAM policy got dynamodb:DeleteItem added.
Used an index2.html staging approach for testing -- uploaded the work-in-progress to S3 as index2.html alongside the production index.html, so changes could be previewed at nowigetit.us/index2.html without affecting live users. Once approved, deployed normally and cleaned up the test file.
After merging, a few more polish passes landed directly on main: the title row was visually centered by adding right padding equal to the logo width (the logo was making the row appear shifted left relative to the centered subtitle), the footer got copyright and author link, and the carousel gained continuous auto-scrolling.
Auto-scrolling Carousel
The library carousel was static -- users had to click arrows to browse. Converted it to a continuously auto-scrolling marquee using requestAnimationFrame at 0.25px/frame for a smooth, slow drift. The trick to seamless looping: clone all carousel cards and append them, then when scrollLeft passes the original set's width, jump it back by that amount. The clones ensure the visible cards are identical at the jump point, making the reset invisible.
The scroll-snap CSS had to be disabled since it conflicts with sub-pixel continuous scrolling. Both arrow buttons now always show (no hiding at "ends" since there are no ends). User interactions pause the auto-scroll: arrow clicks pause for 5 seconds, hovering over the carousel pauses immediately, and mouse-leave resumes after 1 second. All pause/resume logic shares a single timer to avoid race conditions.
Article Field Metadata
The library carousel cards showed paper titles and authors, but nothing about what kind of paper it was. Extended the metadata extraction to include the paper's academic field and subfields. The Claude system prompt now requests <field> and <subfields> elements inside the <metadata> XML block. The _parse_metadata() function in generator.py extracts both, and they flow through DynamoDB, the status API, the library API, and local dev -- same pipeline as title/authors/date.
This was a clean vertical slice: prompt change, parser update, storage, and all API surfaces in one commit.
User-Friendly Error Messages
Discovered while testing with a large PDF (agents-of-chaos.pdf) that exceeds Claude's 185K input token limit. The user saw "Processing failed." with no explanation -- the actual error was buried in CloudWatch logs. Not a great experience when the fix is simply "try a shorter document."
Wrapped the client.messages.stream() call in generator.py with a try/except for anthropic.APIStatusError. The handler inspects the error body and maps known failures to plain-English ValueError messages: "prompt is too long" becomes "This PDF is too large for our AI to process. Try a shorter document." Rate limits, auth errors, and server errors each get their own message. An unknown API error includes the original message as a fallback.
The key design decision was using ValueError as the boundary between "user-facing" and "unexpected" errors. Both lambda_process.py and main.py now catch ValueError separately from Exception -- the former stores str(e) (the friendly message), the latter keeps the generic "Processing failed." No frontend changes needed since it already displays whatever error string the status API returns.
Prompt Injection Defense
NowIGetIt sends user-uploaded PDFs to Claude and publishes the resulting HTML to the web. That's a prompt injection vector -- a malicious PDF could contain instructions that Claude follows, injecting scripts or harmful content into the published page. This needed defense-in-depth rather than a single guardrail.
Four layers went in. First, the system prompt was hardened to explicitly tell Claude that PDF content is untrusted data and to restrict the output to educational HTML -- no external links, no scripts, no iframes. Second, generator.py gained suspicious pattern detection that logs warnings when the generated HTML contains <script>, <iframe>, event handlers, or data URIs. Third, a Haiku-based pre-screening module (content_screen.py) classifies uploaded PDFs before they reach the expensive Opus processing -- if the PDF looks like it's trying to manipulate the AI, it's rejected with an explanation before spending any Opus tokens. Fourth, CloudFront security headers (CSP, HSTS, X-Frame-Options, X-Content-Type-Options) provide browser-level backstops.
The CSP header caused an immediate production bug: connect-src 'none' blocked all fetch() calls from the frontend, including the carousel's library endpoint. Fixed it the same evening by changing connect-src to allow the API Gateway origin while still blocking arbitrary external connections that injected scripts might attempt.
Also created a test harness (tests/make_injection_pdf.py) that generates PDFs with embedded prompt injection attempts -- useful for verifying all four defense layers work together.
Gallery Page
The biggest feature since launch. The carousel on the home page shows the most recent papers, but with 17+ generated explanations there was no way to browse them all. The gallery page at /gallery.html is a full-page grid view with thumbnails, field-based filtering, and infinite scroll.
The Architecture
The gallery required three new Lambda functions working together:
Screenshot Lambda (container image): A Playwright/Chromium headless browser running in Lambda that navigates to each published explanation page, takes a screenshot, converts it to WebP (via Pillow), and uploads the thumbnail to S3. This is the first container-based Lambda in the project -- zip-packaged Lambdas can't include a browser binary. The process Lambda now fires off an async screenshot invocation after successfully publishing each new page.
Gallery Lambda (zip): Queries a new DynamoDB Global Secondary Index (StatusDateIndex, partition key status, sort key created_date) to efficiently fetch completed items in reverse chronological order. Supports cursor-based pagination (base64url-encoded ExclusiveStartKey) and an optional field query parameter for filtering by scientific discipline. Returns items with pre-computed thumbnail URLs.
Backfill script (scripts/backfill_thumbnails.py): A one-off script that scans DynamoDB for all completed items, checks S3 for existing thumbnails, and invokes the Screenshot Lambda for any that are missing. Rate-limited to 5 concurrent invocations.
The Frontend
The gallery page (gallery.html) is a standalone HTML file matching the index page's dark theme -- DM Serif Display headings, amber accents, glass-card effects, grain overlay. It renders a responsive grid (3 columns on desktop, 2 tablet, 1 mobile) with infinite scroll powered by IntersectionObserver. A field filter dropdown lets users browse by scientific discipline. Thumbnails have a graceful fallback: if the WebP fails to load, an onerror handler replaces it with a gradient colored by the paper's field.
Deployment Gauntlet
Getting the Screenshot Lambda to actually work in production required solving five separate problems, each discovered only after deploying:
ECR chicken-and-egg: CloudFormation's Screenshot Lambda resource references an ECR image URI, but the ECR repository and image don't exist until the stack creates them. The main stack failed on first deploy. Solution: split the ECR repository into a separate CloudFormation template (
aws/ecr.yaml) that deploys first, then the deploy script builds and pushes the Docker image, then the main stack deploys referencing the existing image.ARM Mac building for x86_64 Lambda: Docker on Apple Silicon builds ARM images by default. Lambda runs x86_64. Added
--platform linux/amd64to the Docker build command.Multi-arch manifest rejection: Docker buildx on ARM Macs creates a multi-architecture manifest list even when building for a single platform. Lambda's container runtime rejected it with "image manifest, config or layer media type not supported." The fix:
--provenance=falseforces a single-platform image manifest.Playwright browser path: The Dockerfile's
playwright install chromiumruns as root during build, installing to root's home directory. Lambda runs assbx_user1051, which can't access root's home. SetPLAYWRIGHT_BROWSERS_PATH=/opt/browsersin the Dockerfile so the browser installs to a globally-readable path.Chromium sandbox crash: Even with the browser found, Chromium crashed immediately with "Target page, context or browser has been closed." Lambda's sandboxed execution environment doesn't support the Linux user namespace sandbox that Chromium enables by default. Fixed with five launch flags (
--no-sandbox,--disable-setuid-sandbox,--disable-dev-shm-usage,--disable-gpu,--single-process) and bumped Lambda memory from 1024MB to 2048MB since headless Chromium needs the headroom.
Also hit an S3 static hosting gotcha: the gallery link was /gallery but S3 serves files by key name -- there's no URL rewriting. Changed to /gallery.html.
After all five fixes, the backfill ran clean: 17 thumbnails generated, zero failures, all WebP files sitting in S3 at 20-46KB each. The gallery page is live at https://nowigetit.us/gallery.html.
Two Regressions from the Security Hardening
The prompt injection defense commit broke two things that weren't caught until after the gallery page work was deployed.
Upload flow broken by CSP. The CloudFront Content-Security-Policy header included connect-src https://*.execute-api.us-east-1.amazonaws.com to allow frontend API calls, but the pre-signed upload flow PUTs files directly to the S3 bucket -- a different domain entirely. The browser blocked the S3 PUT with a CSP violation, surfacing as a cryptic "Failed to fetch" error. Added the S3 domains to connect-src. This was the second CSP-related bug from the same commit -- the first (connect-src 'none' blocking all API calls) was caught earlier, but this one slipped through because the upload flow involves a cross-origin request to S3 that wasn't covered by testing the carousel endpoint alone.
HTML extraction broken by overzealous input validation. The same security commit tightened the raw HTML detection in generator.py from "<html" in response_text (substring search) to response_text.startswith("<html") (must be at position 0). The intent was to prevent Claude from sneaking content before the HTML. The problem: Claude's response always starts with a <metadata> XML block (title, authors, field) before the HTML, by design of the system prompt. With startswith, the metadata block meant the HTML detection never matched, causing every upload to fail with "Claude did not return valid HTML." The fix was to strip the <metadata> block before running all HTML extraction checks -- code fences, truncated fences, and raw HTML detection all now operate on the cleaned response.
Both bugs shared a common cause: the security hardening was tested against injection scenarios but not against the normal happy path. A lesson in regression testing -- security changes need to be validated against the primary user flow, not just the threat model.