OpenAI launches Ultrafast tier, speed figures are its own claim
High speed no longer requires a smaller model, but the 300 tokens-per-second figure is OpenAI's own; check real latency and bills yourself.
OpenAI introduced a new Ultrafast service tier at its developer day event on September 30, saying Codex can reach up to 300 tokens per second.
According to IT之家's report, the tier offers up to 8x faster generation in Codex and up to 6x via API. It is live now, with Pro 500 subscribers and enterprise users able to try it in ChatGPT Work and Codex.
It is not cheap: gpt-6-astra output costs $300 per million tokens for short context and $450 for long context, while cached input drops to $6. OpenAI also previewed a GPT-6.1 Sol Ultrafast to follow.
The 300 tokens-per-second figure and the speed multiples are OpenAI's own claims with no independent testing attached; the article's reference to "GPT-5.6 Sol capability" also conflicts with contemporaneous coverage of GPT-6.1 Sol, so verify before committing a workload.