OpenAI is previewing Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing. The company says the tier can generate as many as about 750 output tokens per second, aiming at work where every second counts—incident response, live support, financial checks, and real-time commerce. Ultrafast is powered with Cerebras hardware and is launching first in the OpenAI API, with early access via OpenAI’s Ultrafast form. OpenAI frames the move as a new speed class for frontier intelligence: instead of forcing teams to pick a smaller model when they need low latency, Ultrafast tries to keep flagship-level quality while shrinking wait time. Example use cases include reading logs and recent code changes during an outage, spotting suspicious transactions while markets move, and answering shoppers before a cart is abandoned. For builders, the practical takeaway is simple: latency is becoming a product feature, not only a cost line. Teams that already built on GPT-5.6 Sol can treat Ultrafast as a deployment choice for hot paths, while keeping Standard for batch or lower-priority jobs.
Source: https://openai.com/index/previewing-ultrafast/