OpenAI improves GPT-6 prompt caching, cutting agent costs further
OpenAI raises GPT-6 cache hit rates by default and ships diagnostics; agent inference costs may fall, though all figures are vendor-reported.
Original event 2026-09-22
In a product note dated September 22, OpenAI said the GPT-6 family delivers higher prompt cache hit rates by default, with cache discounts for eligible shared prefixes reused within a 30-minute window and discounts of up to 90% on cached input tokens.
New tooling includes a caching dashboard, a cache-miss diagnostics tool, and explicit cache breakpoints. Developers can now change reasoning effort between responses without breaking cache, and prewarm known context to cut first-token latency.
Customer figures in the post are self-reported: GitHub Copilot says the share of prompt tokens needing fresh processing fell by more than 50% over recent months, Manus says its cache hit rate rose from roughly 85% to above 90%, and Strawberry Browser reports a 36% inference cost reduction. These are vendor and customer claims, not independently verified.