Next upAI x Bio Pitch Contest
News

OpenAI details GPT-6 prompt-caching controls and diagnostics

OpenAI's current documentation lays out prompt-cache diagnostics and controls for persistent agents, while the stated performance and savings benefits remain vendor claims.

D
Sep 27, 2026 · 2 min read

OpenAI’s current developer documentation lays out a Prompt Caching Dashboard, request-level miss diagnostics and controls over reusable prompt prefixes. The toolkit is aimed at multi-turn applications, including persistent agents, that repeatedly carry forward instructions, tool definitions and conversation history. OpenAI describes prompt caching as a way to reduce input cost and the time spent processing input, but the opened documentation does not include an independent benchmark for those benefits.

The prompt-caching documentation says the dashboard can monitor cache-read hit rates across an application. Separately, prompt cache diagnostics in the Responses API can compare a current request with a recent completed response and identify the first classified incompatibility involving the model, service tier, tools, settings or prompt input. Developers must use usage.input_tokens_details.cached_tokens to measure actual reuse.

The diagnostic comparison is best effort. Supplying prompt_cache_options.comparison_response_id requests a comparison but does not load the earlier conversation or change caching behavior. Diagnostics do not block or fail the request and do not change model generation.

Explicit breakpoints and prewarming are documented for GPT-5.6 and later supported models, not as GPT-6-only features. In explicit mode, developers mark supported content blocks as cache-write boundaries. Without an explicit breakpoint, the request does not use prompt caching or create cache writes; content after the final breakpoint remains uncached. A request can create up to four cache writes.

Prewarming processes shared instructions, tool definitions or reference material without generating output, preparing the same prefix for a later request. OpenAI says tokens written during prewarming are billed at the standard cache-write rate. The documentation also advises developers to preserve tool definitions, schemas and ordering, use allowed_tools or tool_choice instead of removing definitions, and append new tools or instructions so earlier context remains reusable.

For supported GPT-6 models, a configuration_update item can change reasoning effort between responses while request-level reasoning.effort remains unchanged. OpenAI’s model guidance says this preserves the earlier prompt prefix for potential cache reuse as an agent’s configuration changes.

OpenAI advertises cached-input discounts of up to 90%. For GPT-5.6 and later, its documentation prices cache reads at 0.1 times the uncached input-token rate and cache writes at 1.25 times that rate. Because uncached tokens and writes are billed separately, the read discount is not a promise of 90% savings across an entire request or application. The documentation also says eligible prompts for these models require at least 1,024 visible input tokens and remain eligible for reuse for at least 30 minutes after the latest write or reuse.

More news