CSP Strikes Again
The Content Security Policy continued to be a recurring source of production bugs. Someone processed the Bitcoin whitepaper and the generated page tried to load Chart.js from cdn.jsdelivr.net, which the CSP connect-src directive blocked. Source maps from the CDN were also blocked. On top of that, the Google Analytics tag added a few days earlier was being blocked too since https://www.googletagmanager.com and https://www.google-analytics.com weren't in the CSP allowlist.
This was the third CSP fix since the prompt injection defense work -- first connect-src 'none' blocking all API calls, then the S3 pre-signed upload domain, and now CDN resources and analytics. Each time, the CSP was correct in the security sense (blocking resources not explicitly allowed), but too restrictive for what the generated pages actually need. The fix added cdn.jsdelivr.net to script-src and connect-src, and the Google Analytics domains to script-src, connect-src, and img-src. The img-src allowance was needed because GA uses pixel tracking.
Stats API Endpoints
Added two operational API endpoints for monitoring the system: one returns the number of articles processed today along with cumulative token usage and cost, and the other returns a paginated list of all articles with their metadata. Both hit a DynamoDB GSI for efficient queries by date.
These are operator-facing -- for keeping tabs on usage and cost without having to go into the AWS console. The list endpoint supports cursor-based pagination and optional date filtering, same pattern as the gallery API.
Also wrote a proper README for the project with architecture overview, API documentation for all endpoints, and setup instructions for both local development and AWS deployment.
Thumbnail Reliability Overhaul
The Problem
The gallery had been accumulating blank thumbnail cards. Out of 248 completed articles, many were missing their .thumb.webp files in S3 entirely. The screenshot Lambda -- a Docker container running Playwright with headless Chromium -- was failing silently during the normal processing flow because it's invoked asynchronously (fire-and-forget) from the process Lambda after HTML publishing.
Diagnosis
Running the existing backfill_thumbnails.py script revealed the first problem: it fired off 82 async Lambda invocations, but only ~18 actually executed. No throttling, no errors in CloudWatch -- the invocations just vanished. Switching to synchronous (RequestResponse) invocations revealed the real issues.
The first ~14 invocations would succeed, then Chromium would start crashing with "Target page, context or browser has been closed". After more invocations, the errors escalated to "ENOSPC: no space left on device, mkdtemp '/tmp/playwright_chromiumdev_profile-...'". The Lambda's 512 MB /tmp was filling up with Playwright's browser profile directories and artifact caches, which persist across invocations when Lambda reuses the same container.
The Fix (and the Fix for the Fix)
The first attempt added pkill -f chrome-headless-shell to clean up orphaned Chromium processes before each invocation. This immediately blew up every single invocation with FileNotFoundError: No such file or directory: 'pkill' -- the Amazon Linux-based Lambda container image doesn't ship with procps. The pkill error happened in the finally block after the screenshot attempt, so it masked whatever the actual screenshot result was.
The working fix replaced the shell-based process killing with a /proc filesystem scan: iterate /proc entries, read each process's /proc/{pid}/cmdline, and os.kill(pid, SIGKILL) anything matching chrome-headless-shell. Exception handling covers the race conditions (ProcessLookupError if the process exits between the scan and the kill, PermissionError for system processes, OSError for vanishing /proc entries). This cleanup runs before every screenshot attempt.
The handler was also restructured with a retry loop (MAX_RETRIES=2), an extracted _take_screenshot() function with proper try/finally for browser.close(), and a switch from wait_until="networkidle" to wait_until="domcontentloaded" plus a 2-second fixed delay. The networkidle strategy waits for all network activity to stop, which was causing timeouts on pages with long-running requests or analytics pings. domcontentloaded fires as soon as the DOM is parsed, and the 2-second wait gives JavaScript time to render interactive elements.
On the infrastructure side, the CloudFormation template got EphemeralStorage: Size: 1024 on the screenshot Lambda -- doubling /tmp from 512 MB to 1 GB as a safety buffer alongside the cleanup logic.
Backfill Results
The backfill required three rounds of synchronous invocations. Round one knocked out 18 of 53 missing thumbnails before Chromium started crashing (pre-fix container). Round two after the cleanup fix got 16 more, but the remaining 12 failed consistently -- the pkill bug was masking their results. Round three after fixing the pkill issue went 12 for 12. Final audit: 248 completed articles, 248 thumbnails present, zero missing.