GPT-6 Astra Ultrafast Delivers Up to 8x Faster Token Generation on NVIDIA Blackwell GPUs
GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is now available through the OpenAI API to eligible ChatGPT Work and Codex users.
Powered by OpenAI model-inference optimizations that leverage the NVIDIA Blackwell architecture, Astra Ultrafast provides token generation at speeds of up to 8x faster than Astra Standard mode.
For developers, faster token generation can accelerate editing, testing, and debugging cycles for coding agents. It also reduces the time needed to generate responses between tool calls, helping interactive applications respond more quickly.
Faster responses are especially valuable when model outputs are repeated throughout an agent workflow. Agents write code, use tools, review results, and decide what to do next. Astra Ultrafast is designed to improve these time-sensitive loops, with NVIDIA AI infrastructure helping OpenAI deliver more useful model outputs when developers need them.
“NVIDIA’s significant investments in tools and documentation have enabled us to create very good models for programming Blackwell and Rubin GPUs,” said Philippe Tillet, Inference Lead at OpenAI. “Astra allows us to turn that knowledge into high-performance kernels, making NVIDIA hardware attractive across the spectrum of latency, throughput, and cost. Astra Ultrafast means faster model response as agents write code, use tools, and process complex tasks.”
Continuous performance improvements for AI inference
Performance improvements do not stop after a model is deployed. OpenAI uses its proprietary models to improve inference software running on NVIDIA GPUs and takes advantage of the platform’s programmability to test and implement additional optimizations.
This continuous work can help models respond faster and improve the productivity of deployed infrastructure over time.
“We used internal models to optimize inference on NVIDIA GPUs, and NVIDIA programmability enabled Astra Ultrafast to be faster,” said Uday Ruddarraju, Chief Technology Officer of Computing at OpenAI.
The programmable NVIDIA platform allows developers and researchers to reuse infrastructure across training, inference, and reinforcement learning as their models evolve. This flexibility enables teams to reuse computing resources as demand changes, improving utilization and helping avoid overprovisioning for individual workloads.
How to access GPT-6 Astra Ultrafast
Developers can use GPT-6 Astra Ultrafast through the API starting today. See the Ultrafast mode guide for access, pricing, and implementation details.
Source: blogs.nvidia.com


