AI
Cerebras and Gimlet Cloud Are Chasing 3,000 Tokens a Second, and Speed Turns AI Into a Collaborator
The next leap in AI will come from speed, and speed comes from faster models. Gimlet Labs and Cerebras announced a collaboration on Monday to build an ultrafast inference cloud targeting 3,000 tokens per second at production scale, with the first Cerebras powered datacenter expected online later this year. The pitch is simple. When AI answers in real time, it stops feeling like a tool and starts feeling like a collaborator.
The technical idea is disaggregation. Gimlet Cloud pairs the Cerebras Wafer Scale Engine with high throughput GPUs, then routes each phase of inference to the silicon built for it. Cerebras generates tokens fast; GPUs move volume at scale. Splitting the work this way is the engineering mechanism behind the 3,000 token target, and it reflects a shift the whole industry is making, from training compute to inference latency as the thing that decides product quality.
The companies build on work already underway. Gimlet and Cerebras have worked on joint customer engagements since last year, and an integrated system is already serving tokens in private deployments. Gimlet is also a launch partner for Cerebras's latest CS-4 systems, giving cloud customers a direct path to the newest hardware. Cerebras chief technology officer Sean Lie framed it as datacenter economics, where the fastest tokens plus the highest throughput GPUs gives every customer better value per token. Gimlet chief executive Zain Asgar put it more plainly. Inference speed decides how productive AI can be.
Why does speed matter this much? Because the workloads are changing. Agents, voice assistants, and video systems rise or fall on latency. A voice agent that pauses two seconds between sentences breaks the spell. A coding agent that streams at human reading speed keeps you in flow. At 3,000 tokens a second, a model can produce roughly a page of structured output in under a second, the threshold where real time multi agent orchestration becomes genuinely practical. Fast tokens are more valuable tokens, as the announcement puts it, because they unlock products that slow tokens cannot.
This also redraws the competitive map. The inference market has been a GPU monoculture. A purpose built cloud that mixes wafer scale engines with GPUs, tuned per phase, gives developers a new lever. Pick your speed, pick your throughput, pay for the blend. Cerebras, which trades on Nasdaq under the ticker CBRS, has spent years arguing that when AI is fast, it changes the world. Monday's announcement is that argument moving from the chip lab into production cloud infrastructure.
For you, the takeaway is about what AI will feel like a year from now. The models are already smart enough for most jobs. Presence is the next frontier, the sense of an intelligence keeping up with you in real time. Infrastructure like this is what closes that gap. The next great AI product will know more, and it will keep up.
Quick answers
What is this story about?
The next leap in AI will come from speed, and speed comes from faster models. Gimlet Labs and Cerebras announced a collaboration on Monday to build an ultrafast inference cloud targeting 3,000 tokens per second at production scale, with the first Cerebras powered datacenter expected online later this year. The pitch is simple. When AI answers in real time, it stops feeling like a tool and starts feeling like a collaborator.
Why does this story matter?
For you, the takeaway is about what AI will feel like a year from now. The models are already smart enough for most jobs. Presence is the next frontier, the sense of an intelligence keeping up with you in real time. Infrastructure like this is what closes that gap. The next great AI product will know more, and it will keep up.
Sources
New to crypto? Read the crypto glossary, browse frequent questions, read our story, or explore the story archive.