01 What happened

OpenAI’s documentation describes Ultrafast as the fastest service tier in the OpenAI API. It is currently available to all API users for GPT-6 Astra, subject to low rate limits. GPT-5.6 Sol has preview access.

02 Key details

  • Default GPT-6 Astra limits are 500,000 tokens per minute for API usage tiers 1–3, 1,000,000 for tier 4, and 5,000,000 for tier 5.
  • In each response.create event, set model to gpt-6-astra and service_tier to ultrafast. OpenAI strongly recommends WebSockets for applications making frequent, rapid tool calls: without a persistent connection, network overhead can reduce latency gains.
  • Ultrafast also supports HTTP requests through the SDK. OpenAI recommends using it when speed justifies the higher cost.
  • Ultrafast supports US data residency and global processing, but not EU or other non-US regional processing endpoints.
  • Organizations that work with an OpenAI account team can contact the team to request higher rate limits or preview access for GPT-5.6 Sol.

03 Why it matters

Developers can select OpenAI’s fastest API service tier for GPT-6 Astra while weighing its higher cost, low rate limits, and regional processing restrictions.

04 Who it matters to

Software developers.

Original sourceOpenAI Developers