A Bigger Window
Anthropic made the 1M context window generally available for Claude Opus 4.6 at standard pricing -- no beta header, no tiered rates above 200K tokens, just the same flat rate across the full million. For Now I Get It!, this was a big deal. The old 200K limit meant papers over about 100-150 pages would fail. The new limit handles up to 600 pages and roughly 750,000 words. Most scientific papers are well under that, but some fields -- law reviews, comprehensive meta-analyses, monograph-length reports -- push into territory that was previously off-limits.
The Opus call itself didn't need a model ID change. claude-opus-4-6 already supports 1M natively. I also bumped the max output tokens from 64K to 128K (Opus 4.6's new ceiling), making it configurable through SSM so it can be adjusted without redeploying code.
The Haiku Problem
Here's the catch: before Opus ever sees a paper, the PDF goes through a screening step powered by Claude Haiku 4.5. Haiku checks for prompt injection attempts and classifies the paper into the NCSES taxonomy (the standard classification system for scientific disciplines). Haiku is fast and cheap -- exactly what you want for a security gate. But Haiku is still stuck at 200K context with a 100-page limit. Upload a 300-page document and the screening step would choke before Opus got its chance.
Truncate for Screening, Full Document for Generation
The solution was straightforward: truncate large PDFs before sending them to Haiku, but always send the full document to Opus.
When a PDF is uploaded, the Confirm Lambda now downloads it into memory, checks the page count with pypdf, and if it exceeds a certain number of pages, creates a smaller version with many randomly pages chosen at random from the PDF. That truncated PDF gets passed to Haiku for screening via a pre-signed URL (keeping the existing URL-based flow), and the temp object gets cleaned up afterward.
The IAM Gotcha
First deploy to the test environment revealed a familiar problem: IAM permissions. The Lambda role restricts S3 access to keys under a specific prefix, and my initial code put truncated files under a different prefix that fell outside the allowed path. The Haiku call silently fell back to the original URL (by design -- truncation failures shouldn't block processing), but then the full 46-page PDF exceeded Haiku's 200K token limit and the screening failed.
The fix was trivial -- keep the temp key under the same prefix as the original file. But it's the same class of bug I hit during the takedown system work a few days ago: S3 operations fail silently or with confusing errors when IAM policies don't match the key patterns your code generates. Worth checking IAM before your first deploy, not after.
Results
Tested with LeCun's 1998 paper on gradient-based learning (46 pages) -- a classic that previously would have been at the upper edge of what the system could handle. The screening logs confirmed truncation: "PDF has 46 pages, truncating for screening." Haiku used 196K input tokens for the truncated version. Opus processed the full document at 374K input tokens -- nearly double the old 200K ceiling. The generated interactive page came back in about four and a half minutes.
The FAQ on the site previously warned about a 10MB file size limit. That's been updated to reflect that we can now handle papers up to several hundred pages. For most users, "upload whatever paper you want" is now the practical reality.