01 What happened
OpenAI has launched an improved prompt caching system for its GPT-6 model family. The update aims to increase cache hit rates by default and reduce computational overhead for recurring context.
02 Key details
- The system offers discounts of up to 90% on cached input tokens for shared prefixes reused within a 30-minute window.
- Developers can now adjust reasoning effort between responses for GPT-6 models without invalidating the existing cache.
- A new prewarming feature prepares context in advance to minimize request latency.
03 Why it matters
This update reduces operational costs and latency for persistent AI agents. The ability to modify reasoning effort while maintaining cache efficiency allows for more granular control over model behavior.
04 Who it matters to
AI developers, AI agent developers, corporate AI platform developers and AI integration specialists.
Original sourceOpenAI