Crypt0's NewsCrypt0's News

AI

Cerebras and Gimlet Labs Chase 3,000 Tokens per Second for AI Cloud

The race for faster AI responses just got a big new bet. Cerebras Systems announced Monday it will supply AI chips and hardware drawing roughly 100 megawatts of power to cloud computing startup Gimlet Labs, with deliveries of its CS4 systems spread over one to two years. Gimlet plans to make the capacity available in its cloud in 2027, targeting inference speeds of up to 3,000 tokens per second.

Inference is the moment a model generates its answer, and speed there decides whether an AI product feels instant or sluggish. Gimlet CEO Zain Asgar told Reuters that fast inference unlocks higher value workloads across cybersecurity, voice applications and financial analysis. The partnership pairs Cerebras wafer scale compute with the Gimlet Cloud in a purpose built disaggregated inference cloud, where each phase of inference runs on the silicon best suited to it.

Gimlet will run frontier models, the most demanding AI systems, on the new hardware. Cerebras CEO Andrew Feldman described the agreement as confirmation of how easily the systems deploy alongside hardware from other companies. Gimlet will maintain and operate the systems once delivered, and the companies declined to disclose the financial terms of the deal.

The deal sits inside a wider wave of investment in inference. Training giant models put Nvidia GPUs at the center of the world, and now the market is turning to the daily work of serving answers to millions of users. Earlier this year Cerebras signed a chip supply deal with OpenAI, and last year Nvidia signed a licensing deal with Groq for that company's inference chips and hardware. Speed of response has become the next frontier of competition.

In its own announcement, Gimlet said the first Cerebras powered Gimlet Cloud datacenter is expected to come online later this year. The design integrates the Cerebras Wafer Scale Engine with GPUs into one solution, orchestrating model execution across the whole chip mix. For agentic and realtime applications, where an assistant must plan, act and respond in quick succession, that orchestration is what turns raw compute into a product that feels alive.

What it means for developers is that speed is becoming a competitive feature of AI clouds. As more inference capacity comes online, expect AI assistants that respond without pause, voice agents that hold natural conversations and coding tools that keep up with the flow of thought. The winners of the next wave of AI products will be built on infrastructure that makes the model feel instant, and this deal adds 100 megawatts to that future.

Quick answers

What is this story about?

The race for faster AI responses just got a big new bet. Cerebras Systems announced Monday it will supply AI chips and hardware drawing roughly 100 megawatts of power to cloud computing startup Gimlet Labs, with deliveries of its CS4 systems spread over one to two years. Gimlet plans to make the capacity available in its cloud in 2027, targeting inference speeds of up to 3,000 tokens per second.

Why does this story matter?

What it means for developers is that speed is becoming a competitive feature of AI clouds. As more inference capacity comes online, expect AI assistants that respond without pause, voice agents that hold natural conversations and coding tools that keep up with the flow of thought. The winners of the next wave of AI products will be built on infrastructure that makes the model feel instant, and this deal adds 100 megawatts to that future.

Sources

New to crypto? Read the crypto glossary, browse frequent questions, read our story, or explore the story archive.

← Back to Crypt0's News