Cerebras's technological trade-offs and use cases
"Whole-wafer-as-chip brings super-fast memory access, but at much higher costs and only if everything fits in on-chip memory."
Directionally right, but "only if everything fits" is too absolute, and the cost point needs a better metric.
- Fast memory: correct. WSE-3 puts 44 GB of SRAM on one wafer at 21 PB/s (Introl). A single LP30 chip, by comparison, has 150 TB/s (StorageReview). This is why single-stream decode speed is so high.
- Cost: right, if measured per GB of capacity. SRAM costs far more per bit than HBM. 44 GB per wafer is small next to the hundreds of GB of HBM on one current GPU package, so a large model needs many wafers. A 2T model takes ~23 (Futurum), each a ~23 kW, 15U, water-cooled system (Introl). Gross margin was about 39% in 2025, down from ~42% in 2024 (my calc from S-1/A income statement figures). That is far below GPU-vendor margins and suggests the hardware is not yet earning a pricing premium over its cost. The right investor metric is cost per token at a given speed. Cerebras can win on that at low batch sizes and lose at high batch throughput.
- "Only if everything fits": too strong.
- Inference: models larger than 44 GB are split across multiple wafers (Futurum). The model only has to fit in the aggregate SRAM of the cluster, which costs money but is not a hard wall.
- Training: MemoryX external memory (up to 1.5 PB per system) and SwarmX interconnect stream weights in, so the model does not need to fit on-chip (Introl).
- The real constraint is the KV cache. Long contexts and many simultaneous users enlarge the KV cache, which competes with weights for SRAM. That cuts concurrent users per system and raises cost per token. NVIDIA's LPX design keeps the KV cache in GPU HBM to avoid this (Q1). Real-world deployments reflect the limit: OpenAI's first Cerebras product, Codex-Spark, is a smaller, speed-optimized model (Cerebras). (Inference: that model-size choice fits the memory constraint, but OpenAI has not said that was the reason.)
Better phrasing: "Wafer-scale gives unmatched memory bandwidth. The trade-off is expensive, limited memory capacity, so economics are best for small-to-mid models, moderate context and latency-sensitive traffic. They worsen as model size, context length and concurrency grow."
"It is still useful for Training and Answer inference, but less so for Agentic AI where speed is of less concern."
Largely backwards on both ends.
- Training is Cerebras's weaker case, not its stronghold. The industry and the company have moved to inference, and the IPO is pitched as an inference bet (Futurum, Investing.com). Third-party guides list "workload is mainly training" as a reason not to choose Cerebras, because of CUDA's ecosystem, custom-op flexibility and installed GPU base (Introl).
- Answer (chat) inference is useful, but speed matters least there. Once text streams faster than a person reads, extra tokens/s add little value. The exception is reasoning models that generate long hidden "thinking" before answering.
- Agentic AI is where speed matters most, not least. Agents chain many sequential model calls and generate many intermediate tokens (plans, tool calls, retries), so latency compounds across steps (Futurum, citing Cerebras CTO Sean Lie). The market evidence agrees:
- OpenAI's first Cerebras product is Codex-Spark for agentic coding, at 1,000+ tokens/s (OpenAI, Cerebras).
- NVIDIA launched Groq 3 LPX specifically for agentic inference (Techzine).
- Futurum suggests Codex-Spark alone could absorb the 750 MW OpenAI commitment in 18–24 months (Futurum; analyst projection, not a disclosure).
Where the statement has a grain of truth. Not all agentic work is latency-bound. Background or batch agents (overnight jobs, large-scale data processing) care about cost per token and throughput, where GPUs at high batch size usually win. Agents also tend to carry very long contexts, which hits the KV-cache weak spot from Q3.
Better phrasing: "Cerebras is strongest for interactive, latency-bound inference: coding agents, human-in-the-loop agents, reasoning and voice. It is weakest for training, offline batch agents and very-long-context or high-concurrency serving, where GPUs (and now NVIDIA's GPU+LPX combination) compete on cost."
Sources
- Cerebras S-1/A #2: https://www.sec.gov/Archives/edgar/data/2021728/000162828026033143/cerebras-sx1a2.htm
- Cerebras S-1 (Apr 2026): https://www.sec.gov/Archives/edgar/data/2021728/000162828026025762/cerebras-sx1april2026.htm
- S. Rajgopal, "209 Pages. Zero Discount.": https://thegoogly.substack.com/p/209-pages-zero-discount
- S. Rajgopal, Forbes, 21 May 2026: https://www.forbes.com/sites/shivaramrajgopal/2026/05/21/the-permabulls-blind-spot-what-the-cerebras-ipo-tells-us-about-the-super-ipo-season-ahead/
- Futurum S-1 teardown: https://futurumgroup.com/insights/cerebras-s-1-teardown-is-the-23b-wafer-scale-ipo-the-end-of-gpu-homogeneity/
- StorageReview, Groq 3 LPX: https://www.storagereview.com/news/nvidia-groq-3-lpx-everything-we-know
- DCD, Groq 3 LPU: https://www.datacenterdynamics.com/en/news/nvidia-announces-groq-3-lpu-ai-inference-chip-plans-256-lpu-rack/
- Techzine, Groq 3 for agentic AI: https://www.techzine.eu/news/infrastructure/139653/nvidias-groq-3-lpu-targets-agentic-ai-inference-at-gtc-2026/
- CNBC, 24 Aug 2026: https://www.cnbc.com/2026/08/24/nvidia-says-groq-racks-will-be-online-this-year-after-20-billion-deal.html
- Reuters via Yahoo, NVIDIA–Groq deal: https://finance.yahoo.com/news/nvidia-buy-ai-chip-startup-210634730.html
- Cerebras blog, Codex-Spark: https://cerebras.ai/blog/openai-codexspark
- OpenAI, Codex-Spark: https://openai.com/index/introducing-gpt-5-3-codex-spark/
- Introl, WSE-3 guide: https://introl.com/blog/cerebras-wafer-scale-engine-cs3-alternative-ai-architecture-guide-2025
- Investing.com: https://www.investing.com/analysis/cerebras-48-billion-ipo-tests-the-markets-inference-bet-200680080