OpenAI Launches GPT-6 Astra Ultrafast with 8x Speed Boost on Nvidia Blackwell GPUs

OpenAI has released GPT-6 Astra Ultrafast, a high-performance model variant capable of generating tokens up to eight times faster than the standard Astra mode. The model, which runs exclusively on Nvidia Blackwell hardware, is now accessible through the OpenAI API and is rolling out to eligible ChatGPT Work and Codex enterprise users.
Accelerated generation speeds can minimize the time coding agents spend during edit-test-debug cycles and decrease response wait times between tool calls. Ultrafast delivers token generation speeds up to 8 times faster in Codex at 300 tokens per second, and up to 6 times faster in the API.
OpenAI chief technology officer of compute Uday Ruddarraju stated that internal models were leveraged to optimize inference on NVIDIA hardware. Ultrafast is currently accessible within the API, ChatGPT Work, and Codex on Pro 500 and Enterprise plans.
According to company statements, the eightfold performance leap is driven by deep inference optimizations developed by OpenAI that exploit the native programmability and architectural advantages of Nvidia hardware. The acceleration is designed to impact demanding enterprise workflows, shortening critical development loops from code generation to interactive debugging tools.
Optimizing Iterative Workflows and Agentic Execution
The performance benefits of GPT-6 Astra Ultrafast are most pronounced during iterative workflows where autonomous software agents repeatedly write code, invoke software tools, and evaluate execution results. By cutting response times within these time-sensitive loops, the accelerated model provides timely outputs when developers require them most.
Philippe Tillet, inference lead at OpenAI, stated that Nvidia's investment in tooling and documentation has enabled the company to make its models exceptionally good at programming Blackwell and Rubin GPUs. This close hardware-software co-optimization is central to the release.
Philippe Tillet noted that extensive investments in documentation and tooling by NVIDIA enabled OpenAI models to become highly proficient at programming Blackwell and Rubin GPUs. Astra translates this capability into high-performance kernels that make hardware utilization effective across throughput, latency, and cost metrics.
Continuous Software Refinement and Resource Flexibility
OpenAI's deployment strategy extends beyond the initial release. The company is actively utilizing its own generative models to refine and optimize the underlying inference software running on Nvidia silicon, ensuring continuous performance improvements and maximizing throughput over time.
Uday Ruddarraju, chief technology officer of compute at OpenAI, noted that internal models were deployed to streamline inference execution. This programmable architecture provides essential operational flexibility, enabling developers and research teams to reuse core infrastructure dynamically across training, inference, and reinforcement learning workloads as model demands evolve.
By facilitating flexible resource allocation, the platform helps engineering teams improve hardware utilization and avoid overprovisioning expensive compute infrastructure for individual tasks. While pricing details for the Ultrafast API tier remain undisclosed, the dramatic token generation speedup promises substantial cost-per-token efficiencies for high-volume enterprise deployments.
Sources & Citations
- Quantum Zeitgeist report
- NVIDIA Blog primary recordPrimary / official
- OpenAIPrimary / official
