Caching lets a model store portions of a conversation's context so they don't have to be reprocessed on every request. In agentic coding — where the same large context (system prompt, file contents, tool definitions) is re-sent across many turns — this matters a lot: the cached portion is reused instead of reprocessed, and cached tokens are typically billed at ~10% of the normal input price (source).
Goal: keep the cache warm so each turn re-reads context at ~10% price instead of rebuilding it at full input price.
📚 Source: Optimizing your AI usage to maximize efficiency and reduce cost → "Preserve the cache"
Under UBB you pay for three token buckets (input, cached, output). Cached input is the cheap one — roughly 10% of the input rate. A long agent session re-sends a big, mostly-unchanged context every turn. If that context stays cached, you pay the cheap rate to re-read it. If the cache is invalidated, the entire context is re-sent and billed as fresh, full-price input tokens (source).
So cache discipline is a silent multiplier on every long session: get it right and the running context is cheap; break it and you re-pay for everything.
Three common actions force a full, full-price context rebuild (source):
| Action | Why it breaks the cache |
|---|---|
| Switching models mid-session | One model can't reuse another model's cache — the next request rebuilds from scratch. |
| Coming back to an old session | Caches expire after inactivity — 24h for OpenAI models, ~1h for most others. A cold session rebuilds. |
| Changing reasoning effort, context size, or the enabled tools / MCP servers mid-session | Each of these changes the cached prefix, so it can't be reused. |
- Pick a model and stick with it for the session. If you need to escalate to a stronger model, do it at a natural boundary (new session), not mid-task.
- Configure before you start. Set reasoning effort, context size, and the tool/MCP set up front, then leave them unchanged for the session.
- Use
Auto— it protects your cache. Copilot auto model selection only changes models at natural cache boundaries (a new session, or after/compact), never mid-task (source). - Returning to old work?
/compactfirst. When a session has gone cold, run/compact(Copilot CLI) so what gets rebuilt is a short summary rather than the full history. See Prune Chat History. - Run cheap subagents in their own session. A subagent has its own scoped context and doesn't inherit the main conversation, so assigning it a lighter model doesn't churn the main agent's cache the way a mid-session switch would. See Scoped Agents & Handoffs.
- 🚩 Toggling MCP servers / tools on and off mid-conversation to "try something."
- 🚩 Bumping reasoning effort up for one hard step, then back down — each flip rebuilds the cache.
- 🚩 Leaving a session for lunch and resuming hours later on a non-OpenAI model.
- 🚩 Manually hopping between models mid-task instead of letting
Autoroute at cache boundaries.
- Understand AI Credits & Per-Token Pricing
- Prune Chat History
- Choose the Right Model
- Manage Context Attachments