Google and OpenAI both pushed the same message today: AI is now fast enough to feel impressive for those who use AI agents, each company announcing ultra fast models.
The two launches are built differently, and only one is actually in your hands today.
Myriad: When will OpenAI release GPT-6? Click to make your prediction.
Google shipped Gemini 3.7 Flash, its latest model tuned for software coding and autonomous business workflows. OpenAI opened a limited preview of GPT-5.6 Sol Ultrafast, a new service tier that runs its most capable model at up to 750 output tokens per second.
Gemini 3.7 Flash is a general-availability model. It takes up to a million input tokens (roughly 750,000 words) and returns 64,000, handling text, images, video, audio and PDFs, and it can call tools and control a computer. Google is pitching it as the cheap brain for autonomous systems that plan tasks and finish multi-step jobs with less human help.
The model is not sacrificing quality for speed. It is both more capable and more efficient, being able to complete our test coding task in 2 minutes and 13 seconds whereas the latest Flash model took more than 5 minutes. The quality gap between the two is also noticeable.
OpenAI's Ultrafast isn't a new model. It's GPT-5.6 Sol—the same model OpenAI used an AI red team to harden against prompt-injection attacks before launch—on a faster track, powered by chipmaker Cerebras. It is around 14 times faster than GPT-5.6 Sol's own standard speed.
Cerebras' wafer-scale chips generate up to 750 tokens a second, about 560 words, fast enough that a voice agent can think mid-call.
The numbers that matter
Google's own benchmark sheet puts Gemini 3.7 Flash ahead of Claude Sonnet 5, GPT-5.6 Terra, and others on 11 of 18 tested categories, including a top Code Arena web-dev score of 1,588 Elo and 30.4% on AutomationBench for enterprise workflows. That’s all based on Google's methodology, so treat the lead as the company's claim.
As any Gemini Flash, the model is also cheap. At 75 cents per million input tokens and $3.75 per million output tokens through year-end, it's half of Gemini 3.6 Flash's original rate. That intro price expires December 31, then doubles to $1.50 and $7.50, which is still cheap for a Google model.
OpenAI hasn't published head-to-head scores for Ultrafast beyond customer quotes. Jane Street AI engineer John Crepezzi said in OpenAI’s announcement that Cerebras’ speed "enables different ways of using the models." Podium product lead Courtland Lykins called it "invaluable in our voice stack," saying the speed "completely changes the call experience." The tier is invite-only for now.
The speed push lands as the labs pivot from "who's smartest" to "who's fast enough for agents." Google's timing is pointed. Its flagship Gemini 3.5 Pro is still missing, with no release date given, three weeks after 3.6 Flash and days after a DeepMind leadership reshuffle that moved Demis Hassabis aside for deputy Koray Kavukcuoglu. OpenAI, meanwhile, is renting Cerebras' speed rather than waiting on its own stack.
Gemini 3.7 Flash is live now in more than 160 countries; GPT-5.6 Sol Ultrafast is still invite-only.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。