01 What happened

OpenAI has launched an improved prompt caching system for its GPT-6 model family. The update aims to increase cache hit rates by default and reduce computational overhead for recurring context.

02 Key details

  • The system offers discounts of up to 90% on cached input tokens for shared prefixes reused within a 30-minute window.
  • Developers can now adjust reasoning effort between responses for GPT-6 models without invalidating the existing cache.
  • A new prewarming feature prepares context in advance to minimize request latency.

03 Why it matters

This update reduces operational costs and latency for persistent AI agents. The ability to modify reasoning effort while maintaining cache efficiency allows for more granular control over model behavior.

04 Who it matters to

AI developers, AI agent developers, corporate AI platform developers and AI integration specialists.

Original sourceOpenAI