PDF Metadata and the Library Carousel
Two features landed in quick succession, and the second immediately exposed bugs in the first.
Started with PDF metadata extraction. The system prompt already tells Claude to generate the HTML page, so it was straightforward to also ask it to emit a <metadata> XML block with the paper's title, authors, and publication date. The _parse_metadata() function in generator.py strips it from the HTML before publishing. The process Lambda stores paper_title, paper_authors, and paper_date in DynamoDB alongside the existing job fields.
Then came the carousel -- a horizontally-scrolling row of cards on the home page showing previously generated explanations. This required a new GET /api/library endpoint backed by lambda_library.py, which scans DynamoDB for completed jobs and returns them sorted by date. The frontend renders cards with paper titles, authors, and dates, and auto-refreshes the carousel after each new upload completes. The carousel is hidden entirely when the library is empty.
Building the carousel uncovered a TTL problem: completed DynamoDB items had a 24-hour TTL, which meant successfully processed papers disappeared from the library the next day. Fixed by stripping the TTL attribute (REMOVE #ttl) from the process Lambda's success-path update expression. Errored and in-progress items still clean themselves up.
Once both features were deployed and the carousel appeared on the live site, three bugs became immediately obvious: cards showed filenames instead of paper titles, authors weren't rendered at all, and the processing date was displayed instead of the paper's publication date. Investigation revealed the metadata extraction pipeline was actually working correctly -- the bugs were all in the frontend rendering code. The trickier issue was that the four papers already in DynamoDB had been processed before the metadata PR landed, so they had no paper_title/paper_authors/paper_date fields at all. Backfilled them by scraping titles, authors, and dates from the generated HTML pages already sitting in S3 (since the source PDFs are deleted after processing).
Branding and UI Polish
A batch of small visual improvements. Added the NowIGetIt logo to the title row and as a favicon. Replaced the old tagline with a single description line ("Scientific articles, explained.") between the carousel and the upload zone. Updated footer branding. All purely cosmetic, no backend changes.