OpenAI previewed Ultrafast mode in the API powered by Cerebras wafer-scale hardware. Accelerating GPT-5.6 Sol generation by up to 14x (peaking at 750 tokens/s), it targets low-latency coding loops, dynamic UI rendering, real-time voice, and automated security incident response.

Key Takeaways

  • Cerebras wafer-scale hardware accelerates GPT-5.6 Sol up to 750 tokens/sec (14x speedup).
  • Eliminates latency bottlenecks during multi-step reasoning diff and UI synthesis.
  • Currently in private API preview with planned expansion as capacity scales.
ADSponsored