Here's what my CloudWatch logs looked like when a user hit an Anthropic API error:
Processing error for abc123: Our AI service is temporarily busy.
That's it. Was it a rate limit? A server outage? An overloaded cluster? A bad request? No idea. The error handling in Now I Get It! was catching API errors, converting them to friendly messages for users, and then logging the friendly message. All the operational details -- status code, error type, request ID, retry-after headers -- were extracted from the response and immediately discarded.
This is the kind of bug that doesn't feel like a bug. Users see a polite error message. The code doesn't crash. Everything "works." But the first time you need to diagnose why errors spiked at 2 AM, you realize your logs are useless.
The Fix
The solution was straightforward: log before you translate. A structured [ANTHROPIC_ERROR] log line now fires before the error gets converted to a user-facing message:
[ANTHROPIC_ERROR] status=429 type=rate_limit_error job_id=abc123 request_id=req_xyz retry_after=30 message=Rate limit exceeded
Every field is filterable in CloudWatch. Want to see all throttling events? Filter on status=429. Want to correlate an error to a specific Anthropic request? Search by request_id. Want to know if the API is telling you to back off for 30 seconds or 5 minutes? Check retry_after.
The error handling now uses the Anthropic SDK's specific exception classes -- RateLimitError, AuthenticationError, and so on -- instead of checking raw status codes. Each error gets categorized (throttle, auth_error, overloaded, server_error, api_error, connection_error, timeout), and that category gets stored in DynamoDB alongside the job record. Over time, this builds a queryable history of failure patterns.
The Gotcha
The first deploy to test immediately crashed. The Anthropic Python SDK defines an OverloadedError class for HTTP 529 responses, but it lives in anthropic._exceptions -- a private module. Trying to catch anthropic.OverloadedError raises AttributeError because it's not exported as a public attribute. The fix was simple: handle 529 in the general APIStatusError block by checking the status code directly. A good reminder to actually test against the deployed SDK version rather than just reading source code.
Why This Matters
Error categorization sounds like a small operational improvement, but it changes what questions you can answer. Before this change, investigating API failures meant reading CloudWatch logs line by line and guessing. Now I can query DynamoDB for all jobs where error_type = "throttle" and see exactly when and how often rate limiting is happening. That's the difference between reacting to user complaints and proactively managing API capacity.