AI News
AI News AgentProductOpenAI3 min read

OpenAI tests GPT-5.6 Sol at up to 14 times faster

OpenAI is testing Ultrafast, a new API tier that runs GPT-5.6 Sol up to 14 times faster than Standard processing. Powered by Cerebras, it can generate up to 750 tokens per second, although access is still limited to an initial group of customers.

OpenAI is testing Ultrafast, a new service tier that lets you use GPT-5.6 Sol up to 14 times faster than Standard processing. Access starts in the OpenAI API and is currently limited to an initial group of customers.

The speed comes from a collaboration with Cerebras. Under the announced terms, the system can generate up to 750 tokens per second. Tokens are the small units of text a model produces, such as words, parts of words or punctuation marks.

More speed without choosing a smaller model

Until now, when a company needed near-real-time responses, it typically had to use a smaller or specialized model. Ultrafast aims to reduce that trade-off: keep the capabilities of an advanced model while cutting wait times.

That matters especially when the response needs to arrive while a situation is still changing. For example, during a service outage, a tool could read technical logs, review recent code changes and summarize engineers' reports to help locate the cause and prepare a fix.

OpenAI also points to other possible uses:

  • Incident response: analyze logs, traces and conversations to decide what to check next during an outage.
  • Financial research and security: review market signals, assess transactions and detect suspicious activity as the data changes.
  • Customer support and voice: resolve complex issues in real time without putting the conversation on hold to search for information across multiple systems.
  • E-commerce: answer product questions, check inventory, personalize recommendations and resolve payment issues before the customer abandons the purchase.
  • Research and experimentation: turn processes that previously required waiting all night into interactive cycles of testing, analysis and adjustment during the workday.

What changes in practice

The difference is not just that a response appears sooner. It can also change how products and daily work are designed.

A research team could run an experiment, review the results, modify the hypothesis and run another test without losing momentum. A support team could consult multiple sources while the customer is still speaking. And an engineer could move faster from detecting an alert to testing a possible fix.

OpenAI says its own teams are already testing Ultrafast for incident response and research. In those processes, the model helps read and organize information and suggest next steps, but engineers remain responsible for evaluating decisions and deploying changes.

Ultrafast generates up to 750 tokens per second.

It is still a preview

The service is not generally available. OpenAI is evaluating it with companies in software development, commerce, financial research, customer support and other interactive applications. The company will use those results to decide where it adds the most value and how to expand access as capacity increases.

For you, the immediate effect is limited if you do not develop products with the OpenAI API. But the direction is clear: AI is starting to compete not only on what it can answer, but on whether it can do so at the speed of a conversation, a purchase or an operational emergency. The next point to watch is how much that speed will cost and when it is worth paying for.