Three Presentation Modes
The biggest feature addition since the initial launch: users can now choose how their paper gets explained. The idea came from experimenting with different prompt styles. The original single prompt tried to be everything to everyone -- accessible but not dumbed down, detailed but not overwhelming. The results were inconsistent. Some papers came out great for a general audience, others felt like they were written for someone who already understood the material.
The solution was to split into three distinct modes, each with a carefully crafted prompt:
- Native -- for a technical audience. Preserves the paper's own terminology, reproduces equations with KaTeX/MathJax, creates sortable tables and interactive data visualizations. The goal isn't to simplify but to make the paper a dramatically better reading experience than the PDF. Follows the paper's own organizational structure rather than imposing a narrative arc.
- Non-technical (default) -- for intelligent adults with no domain expertise. Leads with analogies before introducing terminology, contextualizes every number ("what counts as good?"), uses progressive disclosure with expandable "Go deeper" sections. Structured as a narrative arc: hook, foundation, core idea reveal, evidence, implications, takeaways.
- Kid-friendly -- for ages 8-13. Think Kurzgesagt, not textbook. Animated bar races instead of static tables, "click to reveal" comparisons, draggable sliders, concrete comparisons ("That's like stacking 42 million pizzas!"), mini-quizzes between sections, and a "Big Words I Learned" glossary at the end styled like collectible trading cards.
Architecture
The prompts are stored in a new DynamoDB table (${StackName}-prompts) rather than being hardcoded. This means prompts can be edited in the DynamoDB console without redeploying -- useful for iterating on prompt quality. The deploy scripts seed the table using attribute_not_exists conditional puts, so manual edits survive redeployments. File-based defaults in backend/prompts/*.txt serve as both the seed source and a fallback for local development (where there's no DynamoDB prompts table).
The security-critical parts -- the preamble that tells Claude to treat the PDF as untrusted data, the output restrictions blocking external resource loading, and the metadata extraction format -- are hardcoded as SHARED_PREFIX and SHARED_SUFFIX in generator.py. These bookend whatever mode-specific prompt body is loaded. This separation means you can freely edit the creative instructions in DynamoDB without any risk of accidentally weakening the security guardrails.
The mode parameter threads through the entire pipeline: the frontend sends it in the upload JSON body, lambda_upload validates it (returning 400 for invalid values) and stores it in the job record, lambda_confirm reads it from the job and forwards it to lambda_process, which loads the corresponding prompt from DynamoDB (with module-level caching for warm Lambda reuse) and passes it to generate_html(). The completed job record stores presentation_mode for reference.
Frontend
The mode selector is a segmented pill toggle between the description text and the drop zone. Each option has a small ? icon with a hover tooltip: "For a technical audience" / "For an intelligent audience" / "For curious little ones." Non-technical is selected by default. One small bug caught during testing: the tooltips were being clipped by overflow: hidden on the selector container (needed originally for border-radius). Fixed by removing the overflow and adding explicit border-radius on the first and last child options instead.
Deployment
Deployed first to the test environment, verified the mode selector rendered correctly and tooltips worked. Confirmed prompt seeding idempotency by re-running the seeding commands -- all three modes correctly reported "Exists (not overwritten)" with timestamps unchanged from the original seed. Then deployed to production.
Accurate Token Usage and Cost Tracking
The app had been underreporting costs since the prompt injection defense work back on February 25th. Every PDF upload makes two Claude API calls -- a Haiku screening call (content safety + NCSES field classification) and an Opus generation call (the actual interactive HTML page) -- but only the Opus call's tokens and cost were being tracked. The Haiku call was invisible: its tokens were extracted by the SDK but thrown away, never recorded anywhere. For a typical paper, the screening call processes nearly the same number of input tokens as generation (the full PDF), so the reported cost was missing a meaningful chunk.
The Change
Eight files across the full stack, touching every layer from the Anthropic API response up through the frontend display.
The foundational change was in content_screen.py: screen_pdf() now returns screen_input_tokens and screen_output_tokens alongside the existing safety/classification result. The data was already there in response.usage -- it just wasn't being captured. Error and fallback paths return zeros so the pipeline never breaks on missing data.
The ScreenConfig SSM parameter got an upgrade from a plain model name string to a JSON object matching the existing LlmConfig pattern: {"model": "claude-haiku-4-5", "input_price_per_mtok": 1.00, "output_price_per_mtok": 5.00}. The code handles backward compatibility -- if it reads a plain string from SSM (old format), it wraps it in a config dict with default Haiku pricing. This also standardized model naming: switched from date-stamped model IDs (claude-haiku-4-5-20251001) to class-level IDs (claude-haiku-4-5) everywhere, since the date-stamped variants deprecate faster.
The confirm Lambda now computes the screening cost and forwards it (along with model name and token counts) to the process Lambda via the async invoke payload. The process Lambda takes both sets of numbers and writes a comprehensive breakdown to DynamoDB: per-model fields (screen_model, screen_input_tokens, screen_output_tokens, screen_cost, gen_model, gen_input_tokens, gen_output_tokens, gen_cost), aggregate totals (total_input_tokens, total_output_tokens, total_cost), and a models field containing a JSON list with both calls' details for programmatic access.
Verification
Deployed to test and processed a biosensing paper in kid-friendly mode. The DynamoDB record showed the full picture: Haiku screening consumed 48,408 input tokens and 158 output tokens ($0.0492), Opus generation used 48,769 input and 22,015 output ($0.7942), for a grand total of 97,177 input tokens, 22,173 output tokens, and $0.8434 total cost. The frontend displayed the totals correctly (97,177 / 22,173 / $0.84). The screening call accounted for ~5.8% of total cost -- not huge for a single paper, but it adds up across hundreds of uploads.
The old input_tokens, output_tokens, input_tokens_cost, and output_tokens_cost DDB fields are no longer written. Existing records from before this change still display fine since all the API read paths use .get() with defaults -- they just show partial data (generation only) which is accurate for those older jobs.
Waitlist with Email Confirmation
Added a waitlist signup flow for when the daily quota is exhausted. Users enter their email, the backend validates it, stores a pending record in DynamoDB, and fires off a confirmation email via Postmark. Clicking the confirmation link hits a dedicated Lambda that marks the record as confirmed. This was the first external service integration beyond AWS and Anthropic -- Postmark was chosen for its simplicity and deliverability reputation.