All AI updates

OpenAI adds Ultrafast mode for GPT-6 Astra in the Responses API

OpenAI says API customers can use GPT-6 Astra with the Responses API's Ultrafast mode to reduce the time between generated output tokens. Availability is subject to rate limits, uses global processing, and does not support EU or other regional inference residency.

SessionWatcher editorial · Published · Updated · Source announcement: 2026-09-29

What Ultrafast mode changes

OpenAI has added Ultrafast mode for GPT-6 Astra in the Responses API. The changelog describes it as a way to reduce the time between generated output tokens. It does not provide a specific latency reduction or guarantee a response-time target.

To use it, make a Responses API request with the GPT-6 Astra model and set the service tier to "ultrafast". The announcement is for API customers; it does not establish availability in ChatGPT or Codex.

Sources: Openai release notes: 2026-09-29

Availability and residency limits

Access is subject to rate limits. Requests use global processing, and OpenAI says EU and other regional inference residency are not supported. Teams with regional residency requirements should account for that limitation before choosing this mode.

The changelog points to separate Ultrafast pricing information but does not state prices in this entry. Check the applicable pricing details before estimating costs; this announcement alone does not establish a price.

Sources: Openai release notes: 2026-09-29

AI assisted reporting, checked against the linked official sources. Source pages checked 2026-09-30. Editorial process and corrections.