Guides/Caching & request shaping
Caching & request shaping for COGS
Pricing fights start after launch. The fastest margin wins usually combine hit-rate instrumentation, cache keys that respect privacy, and smaller prompts before swapping model tiers — see also pricing & unit economics and model selection.
Stack mental model
Tactics & failure modes
| Tactic | Best COGS / latency wins | What breaks trust |
|---|---|---|
| Semantic cache | Support bots, repeated policy questions, stable corpora | Wrong answer persistence — pair TTL with eval alerts |
| Prompt / prefix cache | Shared system prompts + large doc prefixes across tenants (where allowed) | Privacy boundaries — never cross-tenant bleed keys |
| Deduped fan-out | Burst traffic to same answer — collapse in-flight requests | Thundering herd on cold miss — add jitter and backoff |
| Batch & offline queues | Summaries, indexing, low-interactive workloads | User expectation mismatch if UI promises realtime |
Latency vs spend trajectory
Relative TTFT & token spend index by path (illustrative)
Indices are directional — measure your own P95 with vendor dashboards and tracing.