OpenAI details GPT-6 prompt caching costs and explicit controls
OpenAI's GPT-6 caching announcement highlights explicit breakpoints and diagnostics. Companion API documentation lists cache writes at 1.25 times uncached input pricing for GPT-5.6 and later, with discounted reads and controls to avoid caching changing content.
SessionWatcher editorial · Published · Updated · Source announcement: 2026-09-22
Related guide: How to check your Codex usage
Cache writes and reads have different rates
OpenAI's GPT-6 prompt caching announcement highlights higher cache hit rates, diagnostics and explicit breakpoints. Its companion API documentation explains the billing distinction: for GPT-5.6 and later, cache writes cost 1.25 times the standard uncached input-token rate.
Subsequent cache reads cost 0.1 times that rate on most of these models, or 0.05 times on GPT-6.1 Sol. Cache-write pricing is not an additional fee on top of ordinary input pricing: input tokens use the uncached-input, cached-input or cache-write rate.
Sources: Better prompt caching for GPT-6 · Prompt caching | OpenAI API
Explicit breakpoints control what gets written
GPT-5.6 and later support both implicit and explicit caching. In explicit-only mode, content after the last selected breakpoint uses the uncached input-token rate without a cache-write charge. This lets developers avoid writing changing content that is unlikely to be reused.
If no explicit breakpoints are placed in explicit mode, the request neither uses prompt caching nor creates cache writes. OpenAI also points developers to its Prompt Caching Dashboard for cache read hit rates and its Prompt Cache Diagnostics tool for investigating misses.
Sources: Prompt caching | OpenAI API
AI assisted reporting, checked against the linked official sources. Source pages checked 2026-10-02. Editorial process and corrections.