Patterns — Handling Errors
Contents
- Error taxonomy skeleton (Python)
- Boundary handler → problem+json
- Retry with exponential backoff + full jitter
- Circuit breaker
- Fail-fast startup
- Gotchas
Error taxonomy skeleton (Python)
python
class AppError(Exception):
status = 500
title = "Internal error"
def __init__(self, detail="", **ctx):
super().__init__(detail)
self.detail, self.ctx = detail, ctx
class ValidationError(AppError): status, title = 400, "Invalid request"
class AuthError(AppError): status, title = 401, "Authentication required"
class Forbidden(AppError): status, title = 403, "Not allowed"
class NotFound(AppError): status, title = 404, "Resource not found"
class Conflict(AppError): status, title = 409, "Conflict"
class RateLimited(AppError): status, title = 429, "Too many requests"Boundary handler → problem+json
One handler at the HTTP edge; nothing below it formats responses.
python
@app.exception_handler(AppError)
def app_error(request, exc):
# Exactly one log line per failure, with context, at the boundary.
logger.warning("request_failed", status=exc.status,
error=type(exc).__name__, **exc.ctx)
return JSONResponse(status_code=exc.status, media_type="application/problem+json",
content={"type": f"https://api.example.com/errors/{type(exc).__name__}",
"title": exc.title, "status": exc.status,
"detail": exc.detail, "instance": str(request.url.path)})
@app.exception_handler(Exception)
def unexpected(request, exc):
logger.exception("unhandled_error") # full trace to logs only
return JSONResponse(status_code=500, media_type="application/problem+json",
content={"title": "Internal error", "status": 500,
"detail": "Unexpected error. Retry or contact support."})Retry with exponential backoff + full jitter
python
import random, time
# base 0.2s, factor 2, cap 30s, 5 attempts: worst-case wait ~ <60s total.
def retry(fn, retryable=(TimeoutError, ConnectionError),
attempts=5, base=0.2, cap=30.0):
for n in range(attempts):
try:
return fn()
except retryable:
if n == attempts - 1:
raise
time.sleep(random.uniform(0, min(cap, base * 2 ** n))) # full jitterHonor server hints: if the response carries Retry-After: N, sleep N seconds instead of the computed backoff.
Circuit breaker
python
class Breaker:
# 5 consecutive failures opens; probe after 30s (half-open).
def __init__(self, threshold=5, reset_after=30.0):
self.fail, self.threshold, self.reset_after = 0, threshold, reset_after
self.opened_at = None
def call(self, fn):
if self.opened_at is not None:
if time.monotonic() - self.opened_at < self.reset_after:
raise DependencyDown("circuit open")
self.opened_at = None # half-open: allow one probe
try:
out = fn()
except Exception:
self.fail += 1
if self.fail >= self.threshold:
self.opened_at = time.monotonic()
raise
self.fail = 0
return outFail-fast startup
python
def main():
cfg = load_config() # raises with the missing key named
db.ping(cfg.database_url) # unreachable DB -> crash now, not at first request
run_pending_migration_check(cfg)
serve(cfg)Gotchas
except Exception: passhides bugs for months — if a failure is truly ignorable, log it at DEBUG with a reason string.- Re-raising with
raise NewError(...) from epreserves the chain; bareraise NewError(...)destroys the original traceback. - Retrying a POST without an idempotency key can double-charge/double-create — timeouts are unknown outcome, not failure.
- Breakers per dependency, not global — one dead cache must not open the DB breaker.
- problem+json
typeURLs should be stable identifiers; they don't have to resolve, but never reuse one for a different meaning. - 500 bodies must be static — rendering them from the exception risks leaking internals and can itself fail.