OpenAI makes GPT-6 Astra Ultrafast available to API users
OpenAI has opened GPT-6 Astra's Ultrafast tier to all API users, with low default rate limits and higher-cost positioning.
OpenAI has made its Ultrafast API service tier broadly available for GPT-6 Astra, giving all API users access at low default rate limits. OpenAI calls it the company’s fastest API service tier and positions it for workloads where lower response time justifies a higher cost.
The default limits are 500,000 tokens per minute for API usage tiers 1 through 3, 1 million for tier 4 and 5 million for tier 5. Organizations that work with an OpenAI account team can ask that team for higher limits.
Developers enable Ultrafast on each Responses API request by selecting the gpt-6-astra model and setting service_tier to ultrafast. Standard HTTP requests are supported. For agentic applications that make frequent tool calls, however, OpenAI strongly recommends a persistent WebSocket connection because network overhead can erode the latency benefit.
NVIDIA says GPT-6 Astra Ultrafast runs on its Blackwell GPUs and claims the tier can generate tokens up to eight times faster than Astra Standard mode. Neither source discloses benchmark methodology, workload mix, sample size or latency distributions, so that figure has not been independently verified.
OpenAI says the tier supports U.S. data residency and global processing, but not European Union or other non-U.S. regional processing endpoints. NVIDIA separately says Ultrafast is available to eligible ChatGPT Work and Codex users, although neither source defines the eligibility conditions. GPT-6 Astra is also generally available through Microsoft Foundry.
More news

Databricks adds branch-based restores to Lakebase Postgres

AWS adds hierarchy filtering to Amazon Quick Sight dashboards

ServiceNow CoreAI introduces AutoSynthData for enterprise-agent training
