Standardized NCSES Taxonomy
The field classification system had a design flaw from the start. When metadata extraction was added (Feb 23), Claude Opus was asked to pick a "primary field" and "subfields" with only loose examples ("e.g. Biology, Physics, Computer Science"). This produced inconsistent categories -- one paper might get "Biology" while a similar one got "Biological Sciences" or "Life Sciences." The gallery's field filter and grouping were only as good as Claude's ad-hoc choices, which meant browsing by field was unreliable.
The fix was to replace the open-ended classification with the NSF's NCSES Taxonomy of Disciplines -- 15 broad fields and ~55 subfields used by the National Center for Science and Engineering Statistics for their Survey of Earned Doctorates. This gives the system a constrained vocabulary that's both well-known in academia and granular enough to be useful.
Moving Classification from Opus to Haiku
The more interesting architectural change was where classification happens. Previously, Opus did everything: read the PDF, classify the field, and generate the interactive HTML page -- all in one expensive call. But the Haiku content screening call (added for prompt injection defense) already reads the full PDF to check for adversarial content. That's wasted context if Haiku isn't also classifying.
The new flow combines screening and classification into a single Haiku call. The SCREENING_PROMPT now has two jobs: security screening (same as before) and NCSES field classification. The CLASSIFY_TOOL schema expanded from {safe, reason} to {safe, reason, field, subfields}. One PDF read, two results. Haiku costs roughly 1/60th of Opus per token, so classification is essentially free now.
This also freed Opus from the classification burden. The <field> and <subfields> XML instructions and parsing were removed entirely from generator.py, giving Opus more context window for what it's actually good at -- generating the interactive HTML explanation.
Wiring the Pipeline
The classification result needed to flow through the entire pipeline. In local dev (main.py), the screening result's field and subfields are stored in the in-memory job dict immediately after screening, then preserved through to the completion state. In AWS, lambda_confirm.py passes paper_field and paper_subfields from the screening result into the process Lambda's event payload, and lambda_process.py reads them from the event instead of from generate_html() metadata (which no longer provides them). Six files changed across the pipeline.
Backfill and Gallery
The backfill script (scripts/backfill_fields.py) was updated with the same NCSES taxonomy in its Haiku prompt, plus a --reclassify-all flag that scans every completed record rather than just uncategorized ones. Since the backfill can't send PDFs (they're deleted after processing), it continues to classify from the published HTML page, but with the NCSES taxonomy constraining the output.
A --dry-run --reclassify-all pass showed clean results across all 21 records -- consistent field names, sensible subfield assignments. The live run migrated all 21 records to NCSES-compliant classifications.
The gallery page gained sub-field grouping within each field section. Cards are now organized in a two-level hierarchy: field heading (uppercase, muted), then subfield heading (smaller, left-bordered) with a grid of cards underneath. Cards without subfield data still render under their field heading without a subfield divider. The existing text search and field filter dropdown continue to work unchanged.
FAQ Page
Added /faq.html with 12 Q&As covering what the site does, cost, security, file size limits, and daily upload caps. Same glass-card design as the other pages. Footer across all three pages now shows a centered FAQ link between the copyright and johndamask.com.