
GPT-5.6 Sol Receives Extremely Fast Ultrafast Mode
OpenAI has unveiled the Ultrafast mode for the GPT-5.6 Sol model, generating up to 750 tokens per second. This technological breakthrough promises a fourteen-fold speed increase and primarily targets developers and critical enterprise applications requiring instantaneous response times.
AI Speed Revolution: OpenAI Introduces Ultrafast Mode for GPT-5.6 Sol
OpenAI has officially introduced a new performance mode called Ultrafast. This mode is designed to dramatically accelerate work with its flagship model, the GPT-5.6 Sol version. According to official company data, this mode can generate up to 750 output tokens per second, representing up to a fourteen-fold speed increase compared to standard processing.
At this moment, this is not a feature available to regular users of the free version of ChatGPT, but a technology aimed primarily at developers and the corporate sector.
Performance Without Compromise: 750 Tokens Per Second
The Ultrafast mode, which OpenAI launched on August 13 (initially for its API), aims to eliminate the need to choose between performance and speed. In the past, developers were often forced to choose smaller or specialized models due to high latency. GPT-5.6 Sol in Ultrafast mode is intended to offer the full capabilities of the flagship model while maintaining extremely short response times.
For comparison: in late July, OpenAI launched Fast mode, which achieves roughly 2.5 times the standard speed. Ultrafast is therefore a much more aggressive shift, focused on tasks where every millisecond is critical.
Key Areas of Use for the New Mode
According to OpenAI, extreme generation speed is particularly crucial in the following industries:
- Technical incident resolution: Immediate response to system errors.
- Financial market analysis: Real-time detection of suspicious transactions.
- Customer Support: Faster and smoother interactions with clients.
- Voice assistants and e-commerce: More natural conversations without delays.
The company also highlights the benefits for researchers. Higher processing speeds shorten the cycle between launching an experiment, evaluating results, and subsequent iteration.

Partnership with Cerebras: Infrastructure for Extreme AI
Behind this stunning acceleration is a close collaboration between OpenAI and Cerebras. This company provides specialized infrastructure for so-called low-latency inference. It is thanks to this partnership that GPT-5.6 Sol can run at speeds that were previously the domain of only significantly smaller and less complex models.
OpenAI is already using Ultrafast mode internally to resolve its own technical incidents and as part of internal research.
Availability and Testing
Currently, Ultrafast is only available in a limited preview for a select group of customers. Publicly announced partners testing the technology in the fields of programming, finance, and commerce include Jane Street, Podium, Basis, and Rogo.
OpenAI plans to gradually expand access to this feature depending on available capacity. For regular ChatGPT users, this is not yet a new button in the interface, but a strategic deployment within enterprise and developer ecosystems.