Sub-penny pricing tier optimized for high-volume multimodal batch classification, summarization, and lightweight agent orchestrations.
Architectural Recommendation
- Swap Tier-2 summarization and embeddings guardrails from legacy 1.5 Pro to Flash-8B.
- Activate context caching for recurring static prompt templates (>32k tokens).
- Estimated monthly inference cost delta: -58% with zero regression on JSON parsing.