SPB Git

spb/ultra-sharp-agent-skills Public

Ultra-Sharp Agent Skills — a research-first skill-authoring system + 72 production-ready skills for AI agents.

Python 100%

# Patterns — Caching

# Contents

  • Cache-aside with fallback
  • TTL jitter
  • Single-flight rebuild
  • Versioned keys
  • Purge-on-write
  • Negative caching
  • Metrics
  • Gotchas

# Cache-aside with fallback

python
import json, random

CACHE_TIMEOUT_S = 0.05  # cache slower than 50 ms is worse than origin
TTL_S = 300             # staleness tolerance for this data class

def get_product(pid: int):
    key = f"product:v2:{pid}"
    try:
        raw = redis.get(key)                     # timeout=CACHE_TIMEOUT_S on the client
        if raw is not None:
            return json.loads(raw)
    except RedisError:
        pass                                     # cache down → serve from origin
    value = db.fetch_product(pid)
    try:
        redis.setex(key, jittered(TTL_S), json.dumps(value))
    except RedisError:
        pass                                     # failing to cache is not an error
    return value

# TTL jitter

python
def jittered(ttl: int) -> int:
    return int(ttl * random.uniform(0.9, 1.1))  # ±10% desynchronizes expiry

# Single-flight rebuild

python
LOCK_TTL_S = 10  # > rebuild p99 so a crashed builder's lock self-clears

def get_report(rid: str):
    key = f"report:v1:{rid}"
    raw = redis.get(key)
    if raw is not None:
        return json.loads(raw)
    if redis.set(f"lock:{key}", "1", nx=True, ex=LOCK_TTL_S):
        value = build_report(rid)                # only this process rebuilds
        redis.setex(key, jittered(600), json.dumps(value))
        redis.delete(f"lock:{key}")
        return value
    time.sleep(0.1)                              # others: brief wait then retry once
    raw = redis.get(key)
    return json.loads(raw) if raw else build_report(rid)  # last resort: origin

# Versioned keys

python
# Invalidate a whole family by bumping the version constant in code —
# no scan-and-delete, old entries simply age out via TTL.
PRODUCT_CACHE_V = 3
key = f"product:v{PRODUCT_CACHE_V}:{pid}:{locale}"

# Purge-on-write

python
def update_product(pid: int, fields: dict):
    db.update_product(pid, fields)
    try:
        redis.delete(f"product:v2:{pid}")   # delete, don't rewrite: the next
    except RedisError:                      # read repopulates from fresh origin
        log.warning("purge failed for %s — TTL is the backstop", pid)

Delete (not set) after write: writing the new value here races concurrent readers repopulating from a stale read replica.

# Negative caching

python
NEG_TTL_S = 30  # short: absorbs miss storms without delaying creates for long

if value is None:
    redis.setex(key, NEG_TTL_S, "__miss__")
# On create: redis.delete(key) so the sentinel dies immediately.

# Metrics

python
metrics.incr(f"cache.{family}.{'hit' if raw is not None else 'miss'}")

Alert when hit rate for a family drops below its target (~80%) — usually a key-schema change or an invalidation bug, not traffic.

# Gotchas

  • Caching the serialized response of another cache-user stacks TTLs; staleness = sum of layers, not max.
  • KEYS pattern* for invalidation blocks Redis; that need is the signal to switch to versioned keys.
  • Thundering read-repair after purge-on-write: hot keys need single-flight even with purge-on-write.
  • Objects that serialize differently across app versions poison shared caches during deploys — version the key (rule 5) on format changes.
  • Read replicas + purge-on-write: purge, then repopulate-on-read may read a stale replica and resurrect old data; short TTL bounds the damage.
  • In-process caches in autoscaled fleets are N independent staleness bubbles; keep them ≤5 s or accept per-instance divergence.