Competition from NVIDIA's LPX
What happened
- Dec 2025: NVIDIA struck a ~$20B deal to license Groq's technology and hire its top executives, including founder Jonathan Ross. It was structured as a license plus acqui-hire rather than an acquisition (Reuters via Yahoo Finance, BNN Bloomberg).
- Mar 2026 (GTC): NVIDIA launched the Groq 3 LPX, its first non-GPU rack. It has 256 LP30 chips, each with 500 MB SRAM at 150 TB/s and 1.2 PFLOPS FP8, for 128 GB of SRAM per rack. NVIDIA claims "up to 35x higher inference throughput per megawatt." Target timing is H2 2026, sold first to model builders and service providers (StorageReview, DCD). NVIDIA pitched it explicitly at agentic inference (Techzine).
- Aug 2026: NVIDIA said Groq racks will be online this year (CNBC, headline only; the page would not load).
Why it matters for Cerebras
- It removes the "GPUs can't do fast tokens" narrative. Cerebras's IPO story is fast decode. NVIDIA now offers an SRAM-based fast-decode product inside its own rack, software stack (CUDA, Dynamo) and sales channel. A buyer who wants speed no longer has to leave NVIDIA to get it.
- LPX is designed around the weakness of SRAM-only machines. LPX does not run the whole model. Rubin GPUs do prefill and hold the KV cache in HBM, and LPX runs only the FFN/MoE layers from SRAM, token by token ("attention-FFN disaggregation"). LPX can also act as a speculative-decoding draft model (StorageReview). That sidesteps the KV-cache capacity limit that constrains Cerebras on long-context, many-user workloads (see Q3).
- Distribution and bundling. Hyperscalers and neoclouds already buy NVIDIA. LPX can be priced as part of a Vera Rubin deployment. Cerebras must win standalone deals, and its OpenAI agreement reportedly bars it from selling certain products to named OpenAI competitors (Rajgopal, Forbes). That narrows Cerebras's addressable buyers just as NVIDIA enters its niche.
- Customer diversification gets harder. In 2025, 86% of revenue came from two related UAE parties (MBZUAI 62%, G42 24%) (S-1/A). The bull case needs many new customers beyond the UAE and OpenAI, and LPX now competes for exactly those customers.
Where Cerebras still has an edge
- It is shipping and in production now. OpenAI's GPT-5.3-Codex-Spark runs on Cerebras at 1,000+ tokens/s (Cerebras, OpenAI). LPX performance is not independently validated yet (StorageReview).
- Fewer, bigger devices. WSE-3 has 44 GB SRAM at 21 PB/s on one wafer (Introl). Futurum argues a 2T-parameter model needs ~23 wafers versus ~2,000 networked LPUs, and that small chips bring back scheduling and interconnect penalties (Futurum). Futurum's argument is analysis, not a benchmark.
Bottom line: LPX turns the contest from "architecture" into "distribution and cost per token." Cerebras can keep a raw-speed lead and still lose share or pricing power if NVIDIA's speed is "good enough" and bundled. Things to watch: LPX volume shipments and third-party speed benchmarks; Cerebras's per-token pricing and gross margin; new customers outside MBZUAI, G42 and OpenAI.