Free to read. Sign up to save tools and get alerts when they change. Plus 900+ more AI tool profiles.

Sign up free
5 min read·Updated August 19, 2026

Cerebras Inference is an AI inference platform powered by the world's largest chip — the Wafer-Scale Engine — delivering the fastest token generation speeds for large language models, with partnerships from OpenAI, AMD, AWS, and Meta. The CS-4 rack system launched August 18, 2026, built on an overclocked WSE-3 Turbo rather than new silicon. Cerebras completed the largest US tech IPO of 2026 on May 14, 2026, and trades on Nasdaq as CBRS.

Share

Listen to this overview

Free preview · first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

Learning Objectives

  • Understand what Cerebras Inference is and how wafer-scale chip technology works
  • Compare Cerebras speed benchmarks to GPU-based and competing inference platforms
  • Evaluate Cerebras pricing tiers and the significance of its OpenAI and AWS partnerships

What Is Cerebras Inference?

Cerebras Inference is an AI inference platform built on the Wafer-Scale Engine 3 (WSE-3) — the largest chip ever made. While traditional GPUs use chips the size of a postage stamp, the WSE-3 is the size of an entire silicon wafer: 4 trillion transistors, 900,000 AI-optimized cores, and 44 gigabytes of on-chip SRAM memory.

The result is inference speed that consistently outperforms every alternative. On large models (70 billion+ parameters), Cerebras delivers 2 to 6 times faster token generation than both Groq's LPU and NVIDIA's Blackwell GPUs.

💡Key Concept

Wafer-Scale Engine (WSE): Instead of cutting a silicon wafer into hundreds of individual chips, Cerebras uses the entire wafer as a single processor. This eliminates the communication bottlenecks between separate chips, allowing data to flow across 900,000 cores without leaving the chip.

The CS-4 — More Speed From the Same Silicon

On August 18, 2026, at its Supernova conference, Cerebras launched the CS-4, the rack system that succeeds the CS-3. Cerebras claims more than 1,000 tokens per second on models above 10 trillion parameters, up to 30 times the speed of production graphics-processor systems, and up to 10 times more throughput per watt than the CS-3. Each rack holds three wafers, wafer-to-wafer latency drops to 2 microseconds, and first shipments begin this quarter, with general availability later in the third quarter.

The detail that changes how you read those numbers is what the CS-4 is not. The WSE-3 Turbo is not new silicon — it is the same 900,000-core, 5-nanometer wafer as the two-year-old WSE-3, with the same 44 gigabytes of on-chip memory, clocked from 1.4 gigahertz to 2.8 gigahertz. The gains come from the clock, from three wafers per rack instead of one, and from a modular redesign with separate power and compute shelves that Cerebras says cuts component count by half.

⚠️Warning

An overclock is a real gain, but it borrows from the future. Doubling the clock roughly doubles throughput today, and it is a legitimate engineering result — the power and cooling to sustain it did not exist two years ago. But it does not raise the on-wafer memory ceiling, which is the architecture's actual constraint: 44 gigabytes of SRAM is why frontier models must be split across multiple systems in the first place. It also spends headroom only once. Cerebras projects doubling throughput annually through 2029, and clock speed cannot deliver that again — the next step needs new silicon. Weigh the roadmap accordingly.

Major Partnerships (2025-2026)

Cerebras has secured partnerships with several of the biggest names in AI:

  • OpenAI (January 2026): A multi-year deal to deploy 750 megawatts of Cerebras wafer-scale systems for OpenAI inference — described as the largest high-speed AI inference deployment in the world, rolling out 2026-2028. This went live in August 2026 as OpenAI's Ultrafast service tier, serving GPT-5.6 Sol from Cerebras hardware instead of graphics processors at up to 750 output tokens per second
  • AMD (August 2026): A disaggregated inference design that splits the two phases of serving between vendors — AMD graphics processors handle prefill, where the model reads the prompt, and Cerebras handles decode, where it generates tokens. Cerebras projects roughly 10 times the speed of graphics processors alone and 5 times the throughput of Cerebras alone. Notable as much for the shape as the numbers: rivals cooperating because each phase rewards different hardware
  • AWS (March 2026): WSE chips on Amazon Bedrock, combining AWS Trainium for prefill with Cerebras for decode — the same split the AMD deal generalizes
  • Arista Networks (August 2026): Scale-out networking for CS-4 deployments
  • Meta (April 2025): Powers the Llama API with up to 18 times faster inference than GPU-based solutions

Speed Benchmarks

Cerebras consistently leads inference speed benchmarks, especially on larger models:

ModelCerebras SpeedGroq SpeedSpeedup
GPT-OSS-120B~3,000 tokens/sec~493 tokens/sec~6x faster
Llama 3.1 8B~1,800 tokens/sec~1,345 tokens/sec~1.3x faster
Llama 3.1 70B~450 tokens/sec~275 tokens/sec~1.6x faster
Qwen3 480B Coder~2,000 tokens/secNot availableLargest model hosted
Llama 4 Maverick~2,500+ tokens/secAvailable~2.5x faster than NVIDIA flagship

📝Note

Cerebras's advantage grows with model size. On smaller models (8 billion parameters), the gap narrows. On frontier models (100 billion+), Cerebras pulls significantly ahead because the entire model fits in on-chip SRAM, avoiding the memory bottleneck that slows down GPU-based systems.

Supported Models

As of March 2026, Cerebras hosts a focused selection of major open-source models:

ToolBest For
Llama 3.3 70BGeneral-purpose workhorse model
GPT-OSS-120BOpenAI's open-source reasoning model
Qwen3 480B CoderLargest hosted model; code-focused
Qwen3 235B InstructLarge multilingual instruction model
DeepSeek R1 Distill 70BReasoning-optimized model
Llama 4 MaverickLatest Llama generation; mixture-of-experts
Llama 3.1 8BFast and cheap for simple tasks

Pricing

Free$0
  • 1 million tokens per day
  • 8,192 context length
DeveloperPay-per-token
  • Higher limits
  • Production use
Code Pro$50/month
  • Coding-focused with discounted per-token rates
Code Max$200/month
  • High-volume coding workloads
EnterpriseCustom
  • Dedicated capacity
  • Fine-tuned models
  • SLAs

Per-token pricing (approximate):

ModelInput (per 1 million tokens)Output (per 1 million tokens)
Llama 3.1 8B$0.10$0.10
Llama 3.1 70B$0.60$0.60
Llama 3.1 405B$6.00$12.00

The free tier offering 1 million tokens per day is one of the most generous in the industry — enough for meaningful experimentation without a credit card.

WSE-3 vs. NVIDIA Blackwell

SpecCerebras WSE-3NVIDIA B200
Transistors4 trillion208 billion
Cores900,000 AI cores18,432 CUDA + 576 Tensor
On-chip memory44 GB SRAM192 GB HBM3e
AI compute125 petaFLOPS~4.5 petaFLOPS
Best forInference (speed leader)Training + inference (flexibility)

⚠️Warning

Raw specs do not tell the full story. NVIDIA's ecosystem (CUDA, cuDNN, TensorRT) supports virtually any model and workload. Cerebras excels at inference speed but has a narrower model catalog and does not support custom fine-tuning through its API yet.

💡Key Concept

Cerebras's role in the three-way inference shift. Stratechery's Ben Thompson argued in his May 11, 2026 piece "The Inference Shift" that AI compute is bifurcating into three workload categories that need fundamentally different hardware: training (GPUs win on bandwidth + ecosystem), answer inference (where token speed for human-facing chat matters most), and agentic inference (where humans aren't in the loop and memory capacity + cost-per-token matter more than raw speed). Cerebras's WSE-3 is positioned by Thompson as the canonical "answer inference" play — 21 petabytes per second of on-chip SRAM bandwidth versus 3.35 terabytes per second of HBM on the NVIDIA H100. When the next response in a conversation is what the user is waiting on, that bandwidth gap becomes practical latency. For agentic workloads where the model is making many tool calls without a human watching, the framework predicts cost-optimized memory-heavy hardware — possibly using slower, cheaper DRAM — will out-economize either GPUs or Cerebras-class speed silicon.

Company Details

DetailInfo
Founded2016
CEOAndrew Feldman
HeadquartersSunnyvale, California
Employees~750-800
Latest Funding$1 billion Series H (February 2026)
Public listingNasdaq: CBRS — closed its May 14, 2026 debut at $311 for a $66 billion market cap; see a live quote for the current figure
Total Raised~$2.9 billion private + $5.5 billion IPO proceeds
Key InvestorsTiger Global (lead); Benchmark; Fidelity; AMD; Coatue
IPODebuted May 14, 2026 — 28 million shares priced at $185 (above $115 to $160 range), $5.5 billion raised, stock more than doubled on debut to close at $311 for $66 billion market cap
Notable CustomersOpenAI; AWS; Meta; Group 42; Saudi MBZUAI; Mistral; Perplexity; Mayo Clinic; US Department of Energy
Websitecerebras.ai

Strengths

  • Fastest inference on large models — 2 to 6 times faster than Groq and NVIDIA on 70 billion+ parameter models
  • Wafer-scale architecture — 4 trillion transistors on a single chip eliminates inter-chip communication bottlenecks
  • Major partnerships — OpenAI (750 megawatt deployment), AWS (Bedrock integration), Meta (Llama API)
  • Generous free tier — 1 million tokens per day at no cost
  • Frontier model support — the hosted cloud runs models up to 480 billion parameters (Qwen3 480B Coder); Cerebras claims the CS-4 handles models above 10 trillion parameters at more than 1,000 tokens per second

Limitations and Considerations

  • Narrower model catalog — fewer models than Together AI or GPU cloud providers; focused on major open-source models
  • No custom fine-tuning via API — you cannot upload or fine-tune your own models (unlike Together AI or AWS SageMaker)
  • Inference only — Cerebras Cloud does not offer model training (though on-premise systems support training)
  • On-wafer memory is the real ceiling — 44 gigabytes of SRAM per wafer is unchanged in the CS-4, so frontier models still span multiple systems; the speed gain came from clock and rack density, not from relieving that constraint
  • AWS integration not yet live — Bedrock availability announced for H2 2026 but not generally available yet
  • Ecosystem maturity — NVIDIA's CUDA ecosystem is vastly more developed; Cerebras is still building out developer tooling

IPO and Public-Market Debut

Cerebras Systems completed the largest US tech IPO of 2026, pricing 28 million shares at $185 — above the $115 to $160 range — and raising roughly $5.5 billion in proceeds. The stock more than doubled on debut Thursday, opening at $385 and closing at $311 for a $66 billion market cap. The S-1 named OpenAI, Group 42, Saudi Arabia's MBZUAI (Mohamed bin Zayed University of Artificial Intelligence), and Amazon Web Services as top customers, and Cerebras swung to profitability on $510 million of 2025 revenue. OpenAI is one of Cerebras's largest customers under a multi-year contract worth more than $10 billion and holds a $1 billion secured loan plus warrants for over 33 million shares — making OpenAI a meaningful post-listing shareholder.

The successful debut validates the thesis that frontier labs will pay a premium for Nvidia alternatives when inference economics work, and gives Cerebras a public-market currency to chase Nvidia's data-center share more aggressively. It also positions Cerebras as the first AI-specialized silicon company to reach a frontier-scale public listing and serves as a practical test of Nvidia's GPU pricing power and frontier-lab compute lock-in. Ben Thompson's "Inference Shift" framework — see the InfoBox above — helps frame the answer-inference workload bet behind investor demand: when the next response in a conversation is what a user is waiting on, the on-chip SRAM bandwidth gap between WSE-3 and HBM-based GPUs becomes practical latency.

Key Takeaways

  • Cerebras Inference delivers the fastest AI inference available, powered by the WSE-3 — the world's largest chip with 4 trillion transistors
  • Speed advantage grows with model size: 6 times faster than Groq on 120 billion parameter models, with the hosted cloud supporting models up to 480 billion parameters
  • The CS-4, launched August 18, 2026, gets its gains from overclocking the existing wafer from 1.4 to 2.8 gigahertz and packing three per rack — not from new silicon, which means the 44-gigabyte on-wafer memory ceiling is unchanged and the same trick cannot be repeated next generation
  • AMD and Cerebras now split inference by phase, with AMD graphics processors doing prefill and Cerebras doing decode — a sign that serving is specializing into stages that reward different hardware
  • Major 2026 partnerships with OpenAI (750 megawatt deployment) and AWS (Bedrock integration) validate the technology at massive scale
  • Free tier offers 1 million tokens per day — ideal for experimentation; the company completed the largest US tech IPO of 2026 on May 14, pricing 28 million shares at $185 (above the $115 to $160 range) for $5.5 billion raised, with the stock more than doubling on debut to close at $311 for a $66 billion market cap and OpenAI a meaningful post-listing shareholder

Keep track of the tools you’re evaluating

  • The AI Hub on a phone: a 12-day AI Skill Streak and an expanded Content updates alert listing the saved items that changed.
  • Recommended for you on a phone: nine personalised suggestions labelled Trending in AI news, On your saved list, and Popular.
  • My AI Tools on a phone: saved tools including GitHub Copilot and OpenAI Codex, each with an Updated badge.

Swipe for Recommended for you and My AI Tools

Your AI Hub — sample data.

📰Cerebras Inference in the News

Showing the 3 stories where Cerebras Inference is tagged in Top AI Stories.

Other tools in Inference & Model Serving

Show 5 more →
🧭Recommended for you

Optional detours — these connect to what you just read, and your next lesson will be waiting.