01 What happened
OpenAI’s documentation describes Ultrafast as the fastest service tier in the OpenAI API. It is currently available to all API users for GPT-6 Astra, subject to low rate limits. GPT-5.6 Sol has preview access.
02 Key details
- Default GPT-6 Astra limits are 500,000 tokens per minute for API usage tiers 1–3, 1,000,000 for tier 4, and 5,000,000 for tier 5.
- In each response.create event, set model to gpt-6-astra and service_tier to ultrafast. OpenAI strongly recommends WebSockets for applications making frequent, rapid tool calls: without a persistent connection, network overhead can reduce latency gains.
- Ultrafast also supports HTTP requests through the SDK. OpenAI recommends using it when speed justifies the higher cost.
- Ultrafast supports US data residency and global processing, but not EU or other non-US regional processing endpoints.
- Organizations that work with an OpenAI account team can contact the team to request higher rate limits or preview access for GPT-5.6 Sol.
03 Why it matters
Developers can select OpenAI’s fastest API service tier for GPT-6 Astra while weighing its higher cost, low rate limits, and regional processing restrictions.
04 Who it matters to
Software developers.
Original sourceOpenAI Developers