Guides/Pricing & unit economics
AI pricing & unit economics
AI features fail commercially when teams surprise finance after launch. Ground conversations in token skew (output ≫ input), tool-line items, and how caching changes marginal cost — then pair with packaging experiments instead of hoping gross margin holds.
Cost drivers
Most bills enumerate input vs output tokens separately — agents and verbose completions concentrate spend on output. Hosted retrieval, web search, or bespoke tools often bill per call — treat them as first-class COGS lines, not “misc infra.”
Prompt caching lowers marginal input cost when stable prefixes repeat (system prompts, large RAG dumps). Model its adoption honestly — not every workload repeats enough to benefit.
Illustrative COGS stack
Percentages shift by product shape — use your vendor invoices to replace fiction with quarterly refresh charts for exec reviews.
Where spend often concentrates (one generic copilot-style turn)
Output-heavy agents skew output further; cached prefixes shrink input.
Margin sensitivity (schematic)
Small UX choices that lengthen outputs or invoke tools move margin faster than headline subscription price debates — scenario-plan before committing roadmap.
Margin pressure vs length & tools (schematic index)
Output priced per token — small UX changes that lengthen answers hit margin nonlinearly.
PM checklist
- Publish an internal “cost per successful task” definition aligned with analytics — not raw tokens alone.
- Tie roadmap bets (longer answers, bigger context, more tools) to margin bands before engineering commits.
- Instrument downgrade paths — cheaper models or shorter defaults when sessions exceed COGS targets.