Tech
GPT-6 adds caching controls to cut prompt latency and costs
GPT-6 improves prompt caching with higher hit rates, diagnostics, breakpoints and controls designed to reduce latency and costs.
What happened
GPT-6 introduces prompt-caching improvements, including higher cache hit rates, new diagnostic tools, explicit breakpoints and additional controls. The update is designed to reduce latency and costs when processing prompts, though the source provides no performance figures or release date.
Why it matters
For businesses using GPT-6, better cache performance could mean faster prompt processing and lower costs. The source supports those intended benefits but does not quantify savings, latency reductions or the workloads most affected.
Source: OpenAI News