# 挑战GPU、绕过HBM，深挖史上最大芯片背后的Cerebras

硅谷101 · 2026-08-07

<https://thevalley101.podhood.com/0f86cc5e-f737-43d8-b085-5c151b9c4105>

本期节目讲述AI芯片公司Cerebras的崛起史：从2015年五位工程师在白板上立下目标，到2026年5月上市首日股价上涨68%、市值一度逼近千亿美元。节目采访了早期投资人周楠、百度AI实验室研究员Greg以及Cerebras产品高管Angela，揭示百度因率先发现Scaling Law而成为最早相信Cerebras的投资人之一，并解释了Cerebras如何通过将计算、内存和互联集成到一整片晶圆上，解决数据传输瓶颈，其WSE3芯片面积是最大GPU的57倍。节目还分析了Cerebras与G42的合作如何带来商业规模但引发CFIUS审查，以及2025年平安夜与OpenAI签署的超过200亿美元的推理部署合同，如何标志其从训练转向推理时代。最后，节目探讨了Cerebras面临的供应链、竞争和利润率挑战，以及GPU、HBM和晶圆级芯片三种AI计算思路的差异。

## Questions this episode answers

### Cerebras的晶圆级芯片是什么？它和传统GPU有什么不同？

Cerebras的晶圆级芯片（WSE）是将整张晶圆做成一个巨大的AI处理器，面积是最大GPU的58倍。传统GPU需要将计算芯片和HBM内存分开，数据传输耗时；而Cerebras将内存和计算放在同一片硅上，内存带宽比GPU高数千倍。这解决了AI训练和推理中的数据搬运瓶颈。

[10:49](https://thevalley101.podhood.com/0f86cc5e-f737-43d8-b085-5c151b9c4105?t=649000)

### 为什么百度会成为Cerebras最早的投资人之一？

百度AI实验室在2014年发现Scaling Law，预测大模型将需要巨大算力，而GPU集群无法满足。研究员Greg The Amos和投资人周楠在尽调中验证了Cerebras架构的可行性，认为它能加速AI进展。周楠花了四周时间将风险拆解为可验证的技术问题，最终百度共同领投了Cerebras的C轮融资。

[14:17](https://thevalley101.podhood.com/0f86cc5e-f737-43d8-b085-5c151b9c4105?t=857000)

## Key moments

- **[0:00] 开场**
- **[1:20] 规模定律**
  - [1:20] 2014年吴恩达离开斯坦福加入百度，投入3亿美元成立硅谷AI实验室，团队后来发现Scaling Law：模型越大、数据越多、算力越强，能力提升可预测。
  - [3:28] Dario Amodei将百度团队发现的“可预测Scaling Law”带往OpenAI，成为GPT核心逻辑；百度同时预见到未来模型将撞上GPU算力瓶颈。
- **[4:28] 创业起点**
  - [4:28] 2015年，Andrew Feldman与四位曾在AMD/CMicro共事的工程师在白板写下目标，成立Cerebras，押注晶圆级引擎为AI重新设计计算机。
  - [6:32] GPU集群 vs 晶圆级引擎：行业用高速互联把GPU拼成集群解决AI算力瓶颈，Cerebras把计算、内存和互联全部集成到一整片晶圆上，声称减少跨芯片数据移动。
- **[8:24] 投资尽调**
  - [8:45] 2016年周楠加入百度AI实验室，第一项任务就是寻找Nvidia之外能为深度学习提供算力的方案，他用四周把尽调变成科学项目。
  - [10:50] 周楠用四周把投资尽调变成科学项目：拆出良率、首次流片、缺陷绕行、编译器映射、框架兼容五个可验证的风险问题。
  - [12:25] Greg Diamos拜访Cerebras办公室，看到了实验笔记本、巨大散热金属块和一块已造出的晶圆，才意识到Cerebras在做比GPU大得多的芯片。
- **[14:14] 良率难题**
  - [14:40] Cerebras解决良率难题：把每颗核心缩到0.05平方毫米（H100核心的1%），再加约1%-1.5%备用核心和动态片上网络，46个缺陷不再致命。
- **[17:50] 量产曙光**
  - [17:50] 2016到2019年，Cerebras每月烧钱约800万美元，每六周向董事会汇报‘芯片还是不能用’，三年烧掉近3亿美元。
  - [18:58] 2019年夏天Cerebras第一代WSE成功运行，工程师沉默盯屏半小时；WSE1面积超4.5万平方毫米、40万AI核心，是当时最大GPU的56.7倍。
- **[20:05] 首批客户**
  - [20:17] Cerebras高管：“我们解决了计算机行业最难的物理问题，但没人关心；第一代大约卖出12台，第二代约300台。”
  - [21:14] Cerebras首批真正客户是美国国家实验室：阿贡用它分析新冠变异并获2022 Gordon Bell奖，国家能源技术实验室测出比自家超算快约500倍。
- **[22:28] 救命客户**
  - [22:33] G42的Condor Galaxy超算合同超10亿美元，2023年为Cerebras贡献83%营收、2024上半年87%；但这一客户关系让首次IPO遭CFIUS审查而撤回。
- **[24:50] 平安夜合同**
  - [25:25] 2025年平安夜，英伟达与Groq签下200亿美元交易，OpenAI同一天与Cerebras签下超200亿美元推理部署合同，最高扩至2030年2吉瓦。
  - [26:44] Angela说2025年夏天与Sam Altman交流后，OpenAI意识到快速推理价值；双方感恩节前夜签term sheet，四周后在平安夜签下200亿美元master agreement。
- **[28:30] 推理时代**
  - [28:30] Q: 为什么Cerebras在推理时代突然迎来机会？——训练是大规模并行，推理decode必须串行生成，Cerebras把计算和内存放同一晶圆，推理速度比GPU快15倍。
  - [31:54] Cerebras将产品从超级计算机转成云推理服务：API兼容OpenAI，客户换API key即可接入，平均输出超2000 tokens/秒，传统方案仅200-300。
  - [34:37] Andrew Feldman的策略：推理分prefill与decode，prefill可并行交给AWS，decode由Cerebras大芯片承担，并宣布与除英伟达外所有Hyperscaler合作。
- **[36:00] 产品与上市**
  - [37:44] 周楠回忆IPO首日：定价区间120-160美元，盘中从200涨到380美元，收于311美元涨68%，市值一度接近千亿美元，是职业生涯最疯狂IPO需求。
- **[39:08] 未来挑战**
  - [39:29] Cerebras供应链依赖台积电5nm，但产品避开了当下AI芯片最紧张的HBM、CoWoS先进封装和3nm产线三大瓶颈。
  - [40:49] OpenAI与Broadcom发布自研推理芯片Jalapeno，计划掌握从模型到芯片全栈；Cerebras必须证明其推理速度长期不可替代，不能只靠OpenAI合同。
  - [41:55] Cerebras上市后首份财报：Q1收入超华尔街预期，但Q2毛利率指引仅36%-38%（Q1为47%），股价下跌15%并一度跌破发行价。

## Speakers

- **Angela** (guest)

## Topics

人工智能

## Mentioned

AWS (company), Baidu (company), Cerebras (company), G42 (company), NVIDIA (company), OpenAI (company), CS1 (product), CS2 (product), Cerebras Inference (product), Condor Galaxy (product), Deep Speech (product), GPT (product), GPU (product), HBM (product), Scaling Law (product), WSE (product), WSE2 (product), WSE3 (product)

## Transcript

### 开场

**Host** [0:01]
十年前这是一家没人看好的公司 。

**Host 2** [0:04]
从成立第一天 ， 它就只想做一种从来没有被造出来过的芯片 。 它试图挑战英伟达的霸主地位 ， 却在成立之初屡屡碰壁 ， 多次濒临倒闭 。

十年磨一剑 ， 今年 5 月它终于上市了 ， 市值一度接近千亿美元 。

**Host** [0:21]
CEO 说呢 ， 我们愿意和每一家 Hyperscaler 合作 ， 提升 AI 推理速度 ， 除了英伟达 。

**Guest** [0:28]
And they say it can't be solved, you say in your head you can't solve it.

**Host 2** [0:31]
我们有幸采访了 Cerebras 的早期投资人和最早证明 Cerebras 产品可以落地的研究人员 。 我们还和 Cerebras 内部负责产品的高管讨论了近两年来 Cerebras 的转型和 IPO 之后的动作 。

**Host** [0:44]
那串起这些蛛丝马迹之后啊 ， 我自己呢也很惊讶地发现 ， 要理解 Cerebras 上市的辉煌 ， 我们需要回到十年前 。

**Host 2** [0:52]
回到一块白板 ， 一个快要倒闭的创业公司 ， 一段几乎没有人相信的物理学赌局 ， 还有一个拥有当时全球最大语言模型的 AI 实验室 ， 如何在一个不可能的时间节点 ， 成为了这个故事里最关键的那一页 。

**Host** [1:10]
这是 Cerebras 的故事

。

**Host 2** [1:20]
2014 年， 温达离开斯坦福 AI 实验室 ， 加入百度 。 受到 Google Brain 启发 ， 百度投入了 3 亿美元 ， 成立了位于硅谷的 AI 实验室 。

### 规模定律

**Host** [1:29]
他们最主要的目的是让深度学习发挥商业价值 ，而对于当时的互联网公司来说 ， 这大部分呢都是物理层面的应用 。

**Host 2** [1:38]
自动驾驶 、 机器人 、 语音识别 。

**Guest** [1:40]
The reason why I was so excited about this and I kind of, you know, and Nvidia was like, let's pivot the entire CUDA roadmap to, uh, deep learning because AI is finally working,right? It was originally, uh, demonstrated for vision, like image recognition, but then, you know, after Android moved over to Bidu, we're looking at this and saying like, he's like, let's do something that has product impact, let's do something that people would actually use.

That's why we were starting with speech, because speech had way more users than like image search.

**Host** [2:06]
这是 Greg The Amos，他在独家采访当中告诉我们硅谷 101 啊 ， 温达当时为百度带去了一批顶尖的 AI 研究员 ，而 Greg 就是其中之一 。在此之前啊 ，Greg 在英伟达做了多年的 CUDA 软件架构 ，而当年的 CUDA 还是一个无人问津的内部项目 。

**Guest** [2:25]
Back then, CUDA was kind of a joke and nobody used it.

**Host 2** [2:29]
AI 实验室成立后不久呢 ， 百度团队就发表了一篇知名的论文 《Deep Speech》， 端到端的语音识别模型 ， 允许人们点击麦克风标志对搜索引擎说话 。他们呢还训练了世界上最初的几个深度神经网络 ， 下一步的计划是训练语言模型 。

**Guest** [2:46]
We had this, uh, crazy curly-haired researcher named, uh, Dara Emodi, and we'd worked on language models and scaling language models. Um, we had the biggest, uh, CUDA cluster for deep learning in the world for, for a time there, and, um, basically that became pretty famous.

**Host 2** [3:02]
百度团队研究的语言模型呢 ，在当年是全球最大的 ， 拥有 3 亿参数 ， 训练这个模型需要三个月的时间 。在这个过程中啊 ，他们发现了一个规律 ， 模型越大 ， 数据越多 ， 算力越强 ， 模型就越聪明 ，而这个提升呢是可以预测的 ，是线性的 。

这个规律被称为 Scaling Law。在几年后， 这个经过验证的结果被百度团队发表在了一篇论文里 ， 叫做 "Deep learning scaling is predictable empirically."

Dara Emodi 后来将这个发现带去了 OpenAI， 成为了 GPT 的核心逻辑 。

**Guest** [3:38]
Sometimes the way I thought of it when I first saw it was, it's a recipe for intelligence. I can't believe there's a recipe for intelligence. Like, imagine you saw that. Like, you know, I was raised in the era of, um, computer science, you know, requires a lot of work, you know, all the gains are hard won.

Certain problems are just off-limits, like things related to speech and language and vision were really, really difficult, and this was like, it just worked,right? And it was simple, and all you needed to do was just train a bigger language model.

**Host** [4:08]
但问题呢也随之而来 ， 如果 Scaling Law 是对的 ， 那么要训练将来真正有用的模型所需要的算力 ， 将是现有的 GPU 集群完全无法提供的量级 。

百度的研究员们意识到 ，他们需要一种新的硬件

。在百度团队发现 Scaling Law 的同一年， 有五位硅谷工程师决定离开 AMD。在此之前啊 ，他们在一家叫做 CMicro 的微服务器公司共事 ，2012 年被 AMD 以 3.34 亿美元收购 。

### 创业起点

**Host** [4:42]
这五个人的名字是 Andrew Feldman、Gary Lauterbach、Michael James、Sean Liu、JP Fricker，其中呢 ，Andrew 和 Gary 是 CMicro 的联创 。 第一次聚在一起的时候 ， 五个人在一块白板上写下了他们的愿望 ，他们希望硅谷的计算机历史博物馆有一天会出现他们的名字 。

从第一天起 ，Cerebras 的目标就是造一个为 AI 而生的硬件 。

**Guest** [5:08]
In 2015, and the credit goes to Gary and Sean and JP and Michael, my co-founders, they saw on the horizon the rise of AI. What that meant was there'd be a new problem for computers, that what the AI software would ask from the underlying chip, the processor, would be different.

We came to believe that we could build a better machine for that problem.

**Host 2** [5:35]
我们知道最初的 AI 训练靠 CPU 和 GPU 搭配 ，CPU 是总指挥 ， 负责调度任务 ，GPU 则像一座巨大的工厂 ， 里面有成千上万个小工位 ， 可以进行大规模并行计算 。

但 AI 训练呢 ， 远不止一次计算啊 ， 参数 、 梯度 、 激活值全得在显存 、 缓存 、 计算单元和 GPU 间来回倒腾 。

一旦模型大到单卡塞不下， 就得切成几块分给多卡 ， 让它们互相配合 。

**Guest** [6:05]
And what we saw was that this was going to be the hard problem, and that if we could solve for that problem, we would build an AI computer that was faster and used less power.

**Host** [6:16]
所以啊 ， 真正拖慢 AI 的 ， 很多时候呢 ，是数据在不同芯片 、 不同内存 、 不同机器之间来回移动的过程 。

**Host 2** [6:24]
许多芯片公司解决这个问题的思路呢 ，是把 GPU 拼成一个集群 ， 再花大量工程解决它们之间怎么沟通的问题 。

而 Cerebras 的思路是 ， 将尽可能多的计算单元 、 内存和互联放在一整片晶圆上， 称为晶圆级引擎 。 普通芯片的制造啊 ，是在一张晶圆上同时印出几百个小芯片 ， 然后切开一块一块单独使用 。

但 Cerebras 呢 ，是把整张晶圆做成一个巨大的 AI 处理器 ， 这片晶圆的面积是最大 GPU 的 58 倍 。 很多原本需要跨 GPU、 跨服务器 、 跨网络完成的数据移动 ， 现在都可以在晶圆内部完成 。

**Guest 2** [7:05]
Whereas with a traditional GPU architecture, you would have HBM, which is kept separate from your compute chip, and any time data has to transfer between the memory and the compute takes a long time. It's called memory bandwidth. Because we have memory and compute on the same piece of silicon, we have extremely high memory bandwidth, on the order of thousands of times higher memory bandwidth compared to GPUs.

**Host 2** [7:35]
2016 年 3 月 ，Cerebras 正式成立 ，由 Benchmark 领投 Series A， 融资 2,700 万美元 。

**Host** [7:41]
说到这里啊 ， 我们不得不感叹 Andrew 的远见 。在 2015 年底 ，Cerebras 决定押注晶圆级引擎的时候 ，Transformer 架构甚至都还没有出现 ，而今天这套以大模型为核心的 AI 产业 ，也还远远没有成型 。

**Host 2** [7:56]
当时啊 ， 深度学习当然已经在爆发 ， 英伟达也已经开始把 GPU 推向数据中心和神经网络训练 ，但主流路线呢 ， 仍然是把 GPU 拼成越来越大的集群 ， 再用软件和高速互联解决训练速度的问题 。

**Host** [8:12]
Cerebras 的不同之处在于啊 ， 它从一开始呢 ， 就问了一个更加底层的问题 ： 如果 AI 会成为一种全新的计算负载 ， 那能不能为它重新设计一台机器呢 ？

**Host 2** [8:24]
有投资人回忆 ，2015 年底与 Andrew 见面的时候 ， 发现他当时有一页 slide， 列出了深度学习面临的七个核心问题 ，而这些问题本质上都指向同一件事 ： 训练时间太长 。

### 投资尽调

**Host** [8:37]
这也是上一章结尾我们提到的百度团队所面临的瓶颈 。

这是投资人周楠 ，2016 年他加入百度 AI 实验室 ， 成为了 Greg 的同事 。他的第一个任务 ， 就是寻找一种为深度学习打造的计算方案 。

**Guest 3** [8:57]
As an investor, the task I was assigned by the researchers was to find a very powerful AI compute other than Nvidia, because they predict that in 5 to 10 years, actually more, I guess realistically, probably in 10 years, there will be a very big model coming out, larger than 500 parameters, um, model, and it will be like a Wikipedia type of model.

They all say, okay, we need a different type of GPU, we need a different type of architecture to make sure that the model can scale, can iterate faster, because GPU back then was not the optimized structure for deep learning.

**Host 2** [9:37]
其实那个时候啊 ， 挑战通用 GPU 的尝试呢 ， 还不少 。 谷歌呢已经在内部使用 TPU， 这是一种所谓的 ASIC，是专门为了机器学习里的矩阵计算而设计的 。

英特尔收购了深度学习硬件创业公司 Nirvana， 英国的 Graphcore 在做 IPU，Groq、Hibana、Samba Nova 这些硬件公司呢 ， 基本都是在这个阶段成立的 。

它们沿着不同的方向试图重新设计 AI 芯片 。Cerebras 的方案在这之中是最激进的 。

**Guest 3** [10:07]
Cerebras at that time was absolutely a very bold answer to that problem. Uh, it was literally saying that if AI is going to keep scaling, the compute system has to be reinvented, and the way they do it is to do this wafer-scale compute,right, wafer-scale engine.

So at that time, the question and the risk is, if the compute becomes a bottleneck for the AI progress, is this structure can really work? Can this architecture really accelerate AI compute? Um, so the idea sounds very crazy, but the diligence, entire diligence that I focused on back to 10 years ago, I turned it into a very science-driven project.

But the first principle is, will this architecture work,right? So we break these, all these risks into a lot of small components, small questions.

**Host** [11:00]
接下来啊 ， 周楠花了整整四周的时间 ， 把尽调变成了一个科学项目 ， 把所有的投资风险拆分成了一个个可以验证的技术问题 。

**Host 2** [11:11]
第一 ， 良率能不能控制 ？ 晶圆制造一定会有缺陷 ， 如果一整片晶圆就是一颗芯片 ， 那这些缺陷会不会让整颗芯片报废呢 ？

第二 ， 第一次流片能不能成功呢 ？ 在芯片行业啊 ，Teapod 意味着把设计交给工厂生产 ， 一次失败就可能意味着几千万美元的额外成本和几个月的时间损失 。

第三 ， 如果晶圆上有一部分区域坏掉 ， 系统能不能识别绕开并且继续运行呢 ？ 第四 ， 真实的 AI 工作负载能不能被编译器映射到这样一套全新的架构上 ？

第五 ， 它能不能接入客户已经在使用的机器学习框架 ，而不是要求所有人从零开始重写代码 ？

**Host** [11:56]
尽调过程当中的一个关键转折点呢 ，是在 Greg 见到 Cerebras 团队之后 。 那其实在 Scaling Law 的初步发现以后啊 ，Greg 已经见了大概 20 多家大大小小的做芯片的初创公司 。

这 20 家公司给他的 pitch 呢 ， 都大同小异 ， 基本上就是说 ， 我们要做比英伟达更好的 GPU。 一开始呢 ，Cerebras 对他来说也是这 20 多家当中平平无奇的一家 ， 直到有一天 ，Greg 来到了 Cerebras 的办公室 。

**Guest** [12:25]
Right, like it drove over to the Cerebras office, and then the CEO is like, he has this experimental notebook, and he's like, I figured out how to actually integrate more processors together. And he has this, uh, like a bunch of the, um, thermal engineers, they, they bring out this giant block of metal.

It's like, this is the, um, heat exchange that we connect to the chip, um, and we have to cool it. And I was like, wait a minute, GPUs don't need that much power, you must be thinking of something else.

And so then they bring out this giant wafer, and I'm like, oh, now I get it, you're actually trying to build a bigger chip. And not only are you trying to build it, you've actually succeeded in fabbing one of them.

**Host** [13:03]
那一刻他意识到 ，Cerebras 不是一个山寨版的英伟达 ，他们在做一件完全不同的事情 ， 那就是造一块大到前所未有的芯片 。

如果 Cerebras 真的能造出来这样一块芯片啊 ， 就会解决百度团队的核心问题 ： 速度 。

**Host 2** [13:19]
周楠告诉我们呢 ， 团队将 Cerebras 的编辑器接入了百度自研的 3E 参数语言模型 ，在 Cerebras 的模拟器上跑了一遍测试 ， 得到了不错的结果 。

要知道啊 ， 当时 Cerebras 还没有真正出货的产品 ，他们只有一张实验室里的晶圆模型和一台模拟器 。

**Host** [13:38]
四周后， 周楠得出了他的结论 。

**Guest 3** [13:41]
Yes, this architecture, it will work. It matters a lot. It will truly accelerate the speed that AI will progress. Um, and if this, you know, at that time, that's before Transformer, but we do see that the language model size is going to be really, really large, way above beyond 500 million parameters.

**Host** [14:00]
其他的技术风险呢 ， 当然依旧存在 ，但周楠认为啊 ， 它们都是可以被解决的 。

**Guest 3** [14:06]
And also, if anything happens, as long as the company has enough funding, it's solvable. So then we decided to make this aggressive bet.

**Host** [14:14]
在这里呢 ， 我们可以展开讲讲 Cerebras 是如何解决良率这个问题的 。

### 良率难题

**Host 2** [14:19]
在台积电 5 纳米工艺下， 每平方毫米大约会出现 0.001 个缺陷 。 哎 ， 听起来很小是吧 ？ 但 Cerebras 的 WSE3 芯片面积是 46,225 平方毫米 。

算一下啊 ， 这样一来 ，在一整片晶圆上就大约会有 46 个缺陷 。在一块传统芯片上呢 ， 如果没有容错设计 ， 这 46 个缺陷里 ， 任何一个都足以让整块芯片报废 。

那 Cerebras 的解法呢 ， 包括两个设计 ： 第一 ， 把核心做到极小 。 传统 GPU 的一个计算核心呢 ， 特别大 ， 英伟达 H100 的一个 SM 核心啊 ， 大约 6 平方毫米 。

一旦缺陷落在上面呢 ， 这 6 平方毫米就全废了 。 而 Cerebras 把每个核心缩小到了 0.05 平方毫米 ， 大约是 H100 核心的 1%。

这意味着同样一个缺陷 ， 砸在 H100 上呢 ， 要损失 6 平方毫米 ，但砸在 Cerebras 上只会损失 0.05 平方毫米 。 仅仅是把核心做小这一点啊 ， 就让 Cerebras 的芯片在面对缺陷时的容错能力 ，是传统 GPU 的百倍量级 。

第二招是冗余加智能路由 。Cerebras 在整片晶圆上多放了大约 1% 到 1.5% 的备用核心 ， 平时这些核心是关闭的 ， 闲置着 。

芯片上还内置了一套可以动态重新配置的片上网络 。 制造测试阶段 ， 系统可以识别哪些区域有缺陷 ， 然后把这些坏掉的核心屏蔽掉 ， 通过片上网络绕开它们 ， 把工作负载映射到正常区域 。

从软件的角度看呢 ， 整片晶圆啊 ， 始终是一块完美无缺的芯片 ， 那 46 个缺陷被悄悄地隐藏掉了 。

这个思路呢 ，其实不是 Cerebras 自己发明的 。 内存芯片厂商和 GPU 厂商呢 ， 几十年来都在用冗余核心来处理缺陷 。

英伟达的 H100 需要 132 个核心工作 ，但芯片上实际造了 144 个 ， 允许最多 12 个是坏的 。Cerebras 真正的创新 ，是把这个冗余的逻辑从一块小芯片放大到了整整一张晶圆的尺度 ， 用几十万个微小核心 ， 加成一套能在整片晶圆上动态绕涨的路由网络 ， 把一个 70 年无解的良品率难题 ， 变成了一个可以管理的工程问题 。在尽调之后， 百度共同领投了 Cerebras 的 Series C。其实啊 ，在做出投

资决定之前呢 ， 周楠还见到了 Andrew 本人， 创始人的个人魅力也是说服投资人的重要原因之一 。

**Guest 3** [16:52]
He's a great storyteller. He is bold, he has this vision, um, and conviction. Ten years ago, when I was doing the diligence, about four weeks, I spent almost two hours every day with him on the phone, um, asking all the, every single detail about the risks.

So he answered everything very well. You can tell that he has this engineering, very critical engineering mindset, and he knows exactly how to solve all the problems.

**Host** [17:21]
当时啊 ，Cerebras 的团队有 70 多名 PhD，其中大部分人呢 ， 都在 CMicro 共事过 。 最后呢 ，Greg 告诉我们 ， 就算 Cerebras 找不到客户 ，他也相信这个产品是有价值的 。

**Guest** [17:33]
It's going to end up rolled up into Nvidia or AMD or Intel or something, even if, even if they, um, can't find customers. So I was like, and, and also, by the way, this is an important problem for Baidu, so let's sign up and be a customer.

**Host** [17:50]
虽然有周楠这样的投资人选择相信 Cerebras，但那段时间 Cerebras 实际上举步维艰 。

### 量产曙光

**Host 2** [17:56]
从 2016 年成立到 2019 年中旬 ，Cerebras 完全隐身 ， 没有产品 ， 没有客户公告 ， 没有新闻稿 。 这个在外界眼里神秘莫测的硬件公司 ， 内部却一度非常危险 。

那个时候啊 ，Cerebras 每个月的烧钱速度大约是 800 万美元 ， 每六周开一次董事会 ， 它都要汇报同一个坏消息 ： 芯片还是不能用 。

**Host** [18:21]
三年， 也就是将近 3 亿美元 ， 都花在一个没人见过的产品上 。

**Host 2** [18:26]
他们要解决的问题呢 ，不只是芯片本身 。 一张晶圆级的芯片 ， 功耗高达 25 千瓦 ， 大约是一个普通家庭年用电量的两倍多 ， 集中在一块餐盘大小的面积上 。

这意味着他们还需要重新发明散热系统 ， 重新设计供电结构 ， 重新设计与服务器机架的接口 。 上一章里我们提到的技术问题 ， 都是壁垒 。2019 年夏天 ，Cerebras 终于让第一代 WSE 成功量产 ，并开始正常工作 。

那天啊 ， 几位工程师坐在老 Saltos 市中心一间临时搭建的办公室里 ， 盯着电脑屏幕 ， 看着这个过去被认为不可能的系统真的跑起来 。

半小时没人说话 。

**Guest 2** [19:11]
No, we didn't been able to do this, and it's working, and we did this.

**Host** [19:15]
2019 年 8 月 ，Cerebras 终于走出了隐身模式 。

**Host 2** [19:19]
他们在 Hall Chips 大会上发布了 WSE1， 全球第一款晶圆级芯片 。WSE1 的规格在当时看起来近乎荒诞 ， 面积超过 4.5 万平方毫米 ，是当时最大 GPU 的 56.7 倍 ，40 万个 AI 计算核心 ，18GB 的片上 SRAM 内存 。

相比之下呢 ， 英伟达同期 GPU 的片上内存只有几十兆字节 。 从具体产品上看 ，Cerebras 会整体交付一台完整的系统 CS1， 就是这台冰箱大小的机器 ， 把芯片 、 散热 、 供电接口全部整合在一起 ， 插进数据中心的机柜就可以直接使用 。

同年啊 ，Cerebras 完成了 2.72 亿美元的一轮融资 ， 正式成为独角兽 。

**Host** [20:05]
但问题来了 ， 谁会第一个使用它呢 ？ 商业 AI 客户在 2019 年还没有形成规模 ， 大家呢 ，也都在使用英伟达 。 市场上没人看好 Cerebras。

### 首批客户

**Guest 2** [20:17]
We solved it, and we solved this sort of the hardest problem in the computer industry, and nobody cared. Nobody. It was like, you know, the first gen we might have sold a dozen, the second gen we probably sold 300, and now we're still going to sell tens of thousands.

And the third gen, we had a two or three-year period where we were ahead of the market, and absolutely nobody cared that we were blisteringly fast.

**Host** [20:42]
Cerebras 的第一批真正的客户 ， 来自于一个意想不到的地方 ： 美国能源部旗下的国家实验室 。

**Host 2** [20:50]
阿贡国家实验室用 Cerebras 的系统分析新冠病毒的变异序列 。 后来啊 ， 这个实验协助获得了 2022 年的 Gordon Bell 特别奖 ， 那是高性能计算领域最重要的奖项之一 。Lawrence Livermore 国家实验室用它来跑核聚变模拟 。

国家能源技术实验室测试的结果呢 ，是 Cerebras 系统比实验室自己的超级计算机快了大约 500 倍 。

**Host** [21:14]
这些客户用真实严肃的科学任务 ， 证明了这个技术可以落地 。

**Host 2** [21:19]
2021 年 WSE2 发布 ， 升级至 7 纳米制程 ， 计算核心增至 85 万 ， 晶体管数量达到 2.6 万亿 。2022 年，Cerebras 用一台 CS2 在单台机器上训练了参数量高达 200 亿的模型 ， 创下了当时的记录 。

这年， 美国计算机历史博物馆为 Cerebras 的 WSE2 揭幕了一个永久展览 ， 题为 《The Biggest Chip in the World》。 世界上最大的芯片 。

**Host** [21:48]
距离 2015 年， 五位创始人在白板上立下那个想要被写进计算机算力历史的目标 ， 过去了整整七年 。

**Host 2** [22:01]
2021 年 11 月 ，Cerebras 完成 Series F 融资 ， 估值 40 亿美元 。2026 年 5 月 ，Cerebras IPO， 估值 950 亿美元 。 五年间 ， 估值涨了将近 25 倍 ，其中啊 ， 一半的功劳都归功于一个客户 ： 他在三年时间里 ， 把 Cerebras 从一家只有政府实验室客户的公司 ， 变成了一家有真正商业规模的企业 。

**Host** [22:28]
但也是这个客户啊 ， 让 Cerebras 的第一次上市计划彻底泡汤 。

### 救命客户

**Host 2** [22:33]
G42 是阿布扎比的一家 AI 科技集团 ， 与阿联酋政府深度关联 ， 同时是微软的战略合作伙伴 。2023 年，G42 和 Cerebras 宣布了一个名为 Condor Galaxy 的合作 ， 基于 Cerebras 的系统呢 ， 建造一个大规模超算网络 ， 最大的一台超级计算机 ， 算力达到 4 个 Exaflopse， 合同规模超过 10 亿美元 。

**Host** [22:57]
其实我自己啊 ，也是通过这个项目认识 Cerebras 的 。2023 年的夏天呢 ，是我接触科技报道的第一年 。在 Cerebras 位于 Santa Clara 的数据中心里 ， 我亲眼见到了这台 6.5 英尺高的超级计算机 。

它在白色的机箱里有序地运作着 ， 平静的白噪音闪烁的蓝光 ， 对我来说也是一个十分难忘的场景 。

**Host 2** [23:19]
Andrew 当时告诉我呢 ， 建造这台以 WSE 为内核的超级计算机 ， 只花了 10 天 。

**Host** [23:25]
那 Cerebras 和 G42 的合作呢 ，是于 2021 年， 通过 Condor Galaxy，G42 想做一个阿拉伯语版的 ChatGPT。

**Guest 3** [23:34]
Technical validation is not, is not enough,right? As a deep tech hardware company, you really need a large customer willing to deploy the systems to pay the real money to build around the technology. So G42 gave Cerebras that. I, I think that really helped Cerebras to, to, to bring real revenue to Cerebras.

That's very important for them.

**Host 2** [23:55]
财务数字就是最好的证明 。2022 年 Cerebras 营收 2,460 万美元 ，2023 年跳至 7,870 万美元 ，G42 占比超过 83%。2024 年上半年，G42 的占比进一步升至 87%。

**Host** [24:13]
直到 2024 年 9 月 ，Cerebras 向纳斯达克提交了 S1 招股书 ， 准备上市 。

**Host 2** [24:19]
美国外国投资委员会 CFIUS 在看到 Cerebras 的大批营收来自阿联酋客户 G42 之后， 几乎立即启动了国家安全审查 。 最后啊 ，Cerebras 在无法获得 CFIUS 放行的情况下， 主动撤回了 S1 申请 。G42 持有的股份呢 ， 被转换为无投票权股份 ， 双方关系重新定性 。

最终 ，CFIUS 在 2025 年 3 月 31 日正式给 Cerebras 上市放行 。

**Host** [24:46]
虽然撤回了 S1 啊 ，但这段时间的等待并没有白费 。

### 平安夜合同

**Host 2** [24:50]
2025 年里 ，Cerebras 完成了两轮大规模私募融资 。2025 年 9 月的 Series G 融资 11 亿美元 ， 估值 81 亿美元 。2026 年 2 月的 Series H 融资 10 亿美元 ， 估值 230 亿美元 。在等待上市的这段时间里 ，Cerebras 用私募市场完成了再一轮资本化 。2025 年 12 月 24 日的平安夜 ， 本应是安静的一天 ，但那一天啊 ， 发生了两件事 ， 改变了整个 AI 芯片产业的格局 。

第一件事啊 ，是英伟达和 Groq 签署合同 ， 达成了 200 亿美元规模的技术授权和资产交易 。Groq 是 Cerebras 在推理芯片市场最直接的竞争对手 ， 它做的是 LPU 语言处理单元 ， 专门为固定大小的模型设计 ， 追求极致的确定性和低延迟 。

**Guest** [25:45]
I, I was always a little bit less enthusiastic about Groq because I always in the back of my mind was thinking like, well, if Nvidia really wanted to, they could do the same thing.

**Host 2** [25:53]
第二件事呢 ， 就是 Cerebras 与 OpenAI 宣布合作 ，他们签署的合同总价值同样超过 200 亿美元 。OpenAI 承诺部署 750 兆瓦的 Cerebras 计算能力 ， 用于 OpenAI 的推理业务 ， 最大可扩展至 2030 年的两吉瓦 。

**Host** [26:10]
这是全球历史上规模最大的 AI 推理部署合同 。 哎 ， 大家捋一捋这个时间线啊 ，Cerebras 呢 ，是英伟达的挑战者 。

而就在英伟达吃下 Groq 的同一天 ，OpenAI 选择把最大的推理计算合同给了 Cerebras。其实啊 ， 这个合作也并不令人惊讶 。Sam Altman 呢 ，是 Cerebras 最早期的投资人之一 ，而两家公司团队从 2017 年起呢 ， 就频繁交流 。他们啊 ， 都相信一点 ： 随着模型规模扩大 ， 硬件架构和模型需求迟早会汇合 。

**Guest 2** [26:44]
I, I think I spoke to Sam in, in sort of middle of the summer in, in '25, and he said for the first time, we've been trying so hard just to keep up with demand. We, we now see the importance of fast inference.

That produced, uh, a set of trials and some testing that, that was done, um, and we were so much faster than, than the competition. It felt really good. Because I, I, I really want super smart customers who are doing really interesting things with our stuff.

And so we got in with, um, some of their guys and they were like, whoa, this is, we understand now.

**Host 2** [27:23]
于是 OpenAI 和 Cerebras 在感恩节前夜签了 term sheet，在四周后的平安夜签了 master agreement。 一个 200 亿美元以上的交易能在四周内完成啊 ，其实呢 ，是很罕见的 。

**Guest 3** [27:36]
So the simple reason for OpenAI to choose Cerebras is speed and also supply diversification. Uh, OpenAI needs more compute than just a single supplier that can easily provide,right? It also wants fast inference for real-time products, and it doesn't want to just rely on one compute, which is Nvidia, for inference.

**Host** [27:55]
这个合同落地的速度也超出了预期 。

**Angela** [27:59]
We shipped the first integration in February. So within a few weeks of announcing, we shipped our first integration with Codex. It was very well received. We, we launched the Codex Spark model, which was particularly optimized for speed and interactivity, uh, with Cerebras on the back end and made it available to the world.

And, you know, that was a way for many developers and individuals, not just companies, to experience Cerebras firsthand.

**Host** [28:30]
那在讲述 OpenAI 的合作的时候啊 ， 我们经常提到的一个关键词呢 ，是推理 。在最开始几章呢 ， 我们关注的是 Scaling Law，是模型的训练 。

### 推理时代

**Host** [28:39]
这也是 2024 年之前 Cerebras 和大部分硬件公司都在关注的方向 。 训练是建造模型 ， 你要把海量的数据喂给一个神经网络 ， 让它学会了语言的规律 、 世界的知识 。

这是一个巨大的 、 一次性的计算过程 。 那 GPU 呢 ， 非常适合做训练 ，是因为矩阵乘法是高度可以并行化的 。

你可以把几千块 GPU 连起来同时计算 。 但训练完成 ，AI 被我们用上的时候 ， 发生的就是推理这个过程 。

那从硬件角度看啊 ， 推理是一个非常特殊的计算过程 。

**Host 2** [29:14]
代语言模型生成回答的时候啊 ，不是一次性把整段话吐出来 ，而是生成第一个词 ， 才能知道第二个词该是什么 。

生成了第二个词呢 ， 才能继续生成第三个词 。 也就是说啊 ， 推理有很强的串行性 。 训练的时候 ， 你可以把海量数据分成很多批次 ， 扔给几千张 GPU 并行计算 。

但推理的时候 ， 尤其是生成答案的 decode 阶段 ， 模型必须沿着时间顺序往前走 。 每生成一个 token 都要进行一次模型计算 ， 都要在模型权重 、 缓存和计算单元之间搬动大量的数据 。

所以呢 ，在推理时代 ， 关键词是速度 。 模型能不能足够快地读到数据 ， 足够快地生成下一个 token。

**Host** [29:58]
这个时候我们才能看到 Cerebras 的 WSE 和 GPU 相比所产生的优势 。

**Angela** [30:04]
And with inference, it's all about memory bandwidth. So our fundamental innovation of having compute and memory co-located on a silicon wafer is what delivers very high memory bandwidth and ultimately delivers very high performance, high-speed inference for end-user applications.

**Host** [30:23]
换句话说啊 ，Cerebras 原本想解决的是 AI 训练过程当中的数据搬运问题 ，但当 AI 进入推理时代之后呢 ， 反而更加的利好 Cerebras。

你想想 ， 你每次和 AI 对话调用 agent 的时候 ，是不是希望速度快一点 、 更快一点 ？ 推理也从一个辅助性的需求变成了 AI 基础设施的核心 。

那我也问了 Angela，Cerebras 为什么选择从训练市场转而关注 AI 推理市场呢 ？

**Angela** [30:51]
So we began in training in part because back in 2016, most companies were focused on training foundation models. Now it's different. We have frontier models from OpenAI, from Anthropic, from many open-source companies that are powerful enough to be used broadly for applications.

From a Cerebras perspective, we were good for both. But when we saw how quickly the inference market was expanding and we had this particular advantage of 15 times faster speed for inference compared to GPUs, we knew that there would be a very good opportunity here for us to enable all of the many companies using inference to grow in ways that they weren't able to do before.

**Host** [31:32]
当然啊 ， 推理市场的需求呢 ，也改变了 Cerebras 的产品形态 。

**Host 2** [31:36]
之前啊 ， 就像我们讲的 G42 这个例子里面 ，Cerebras 卖的是超级计算机 ，但现在呢 ， 它开始把自己的硬件能力包装成云服务 。

客户通过 API 发送请求 ，Cerebras 在后端完成模型推理 ， 把结果返还给客户 。 那 Angela 告诉我们啊 ，他们已经把整个产品做成 OpenAI API 兼容 。

如果一家公司已经在用 OpenAI 或者其他的兼容接口 ， 就不需要重写整套代码 ， 换成 Cerebras 的 API key， 就可以把请求送到 Cerebras 后端 。在推理服务上呢 ，Cerebras 目前平均可以看到每秒超过 2,000 个 token 的输出速度 。

相比之下， 很多传统方案可能是每秒 200 到 300 个 token。

**Host** [32:19]
只要大家用过 agent 工作流啊 ， 就会知道这样速度的提升呢 ， 会让使用体验大幅优化 。 更有意思的是 ， 速度更快的时候 ，也会创造出全新的工作场景 。

哎 ， 这里我们来看几个例子 。

**Host 2** [32:31]
有一家 Cerebras 合作伙伴做了这样一个功能 ： 当程序员把鼠标悬停在一段代码上， 系统可以立刻显示这段代码在整个项目里被哪些地方调用 ， 依赖关系是什么 。

那这个功能呢 ， 本质上需要模型快速读懂代码上下文 ， 再生成一个完整的调用图 。 如果推理慢啊 ， 这个功能呢 ， 就很尴尬 。

如果光标要停留 5 到 10 秒才有响应 ， 那开发者的思路其实已经断了 。 但如果响应几乎是瞬时的 ， 它就不再像一个查询工具 ，而更像是 IDE 里自带的一部分 。

**Host** [33:08]
另一个例子是实时教育 。

**Host 2** [33:10]
想象你需要让 AI 给你解释一个复杂的概念 ，AI 一边在虚拟黑板上画图 ， 一边回答用户的追问 。 那用户呢 ， 可以直接圈出图上的某一部分 ， 让 AI 再讲一遍 。

如果 AI 要停顿很久 ， 这就只是一个普通教学视频加聊天框 。 但是啊 ， 如果 AI 能够立刻回应 ， 它就变成了一个真正的老师 ， 可以实时对话 。

**Host** [33:32]
第三个例子呢 ，是 AI 科学家 。

**Host 2** [33:35]
有 Cerebras 的合作伙伴呢 ，在做一个面向药物研发的 AI 科学家 。 想象一下 ，在一个新药研发的会议里 ， 团队正在讨论一个假设 ： 这个 AI 员工呢 ， 可以在会议进行的同时， 实时查找竞争对手研究 、 检索论文 、 测试假设 ， 然后直接参与讨论 。

**Angela** [33:54]
So all of these are examples of experiences that would not be possible with slower inference, but are made possible with much faster inference like Cerebras.

**Host** [34:04]
这就是推理时代最重要的变化 ：agent 直接进入实时交互场景 ， 代码 、 教育 、 科研 、 客服 、 语音助手 ， 这些应用的共同点都是互动 ，而互动的本质就在于速度 。

这也解释了为什么 Cerebras 会在 2026 年这个时间点迎来它的 IPO 窗口 。在宣布和 OpenAI 合作后不久 ，Cerebras 又在 2026 年 3 月宣布与 AWS 合作 。CEO Andrew Feldman 在采访中解释了这个合作的逻辑 ，并且宣布啊 ， 将会和所有的 Hyperscaler 合作 ， 除了英伟达 。

**Guest** [34:38]
What is the essence of the inference problem? And it's comprised of two parts. It's comprised of a part called, uh, processing the prompt. And then there's a second part, which is generating the answer. So you process the prompt and you generate the answer.

We call the first part prefill and the second part decode. And it turns out they have some very different compute characteristics. And so we, we thought to ourselves that there are machines that are better than us at this prefill.

It is a paralyzable problem. It has fundamentally different characteristics than the decode, which is a strictly sequential problem. And so we took this observation and we went to, to, to AWS and said, we can use your training part to do the prefill and we'll use our big chip to do the decode.

And what we will get is this extraordinary solution. And, uh, it turned out to be really well received. And we're now engaged in, in that, uh, process of using other people's part for part of the problem and our part for another part of the problem, uh, sort of with all members of the community.

**Guest 2** [35:50]
Other hyperscalers?

**Guest** [35:51]
That aren't, uh, that aren't Nvidia. So everybody but them.

**Host** [36:00]
那在我们继续讲 Cerebras 的 IPO 故事之前啊 ， 我想先花一分钟把 Cerebras 这十年做的产品解释清楚 。 第一层产品是芯片 。

### 产品与上市

**Host 2** [36:09]
2019 年，Cerebras 做出第一代晶圆级引擎 WSE One。 这是世界上第一款晶圆级芯片 ， 大小是当时最大 GPU 的 56.7 倍 。

**Host** [36:20]
但我们说过啊 ，Cerebras 卖给客户的不是芯片 ，而是机器 。 这就是 Cerebras 的第二层产品 ， 系统 。

**Host 2** [36:28]
WSE One 被装进 CS One 系统里面 ， 这是一台可以直接放进数据中心的 AI 超级计算机 。2021 年，Cerebras 发布第二代 WSE Two。2022 年 11 月 ，Cerebras 用 16 台 CS Two 拼成了 Andromeda 超级计算机 ，1,350 万个 AI 核心 ， 超过一个 Exaflop 的 AI 算力 ， 功耗 500 千瓦 ， 远低于同级别的 GPU 集群 。

这展示了多台晶圆级系统可以组成集群的能力 ， 让大语言模型训练实现接近现今的扩展 。2024 年，Cerebras 发布第三代 WSE Three。

这是当时最大 GPU H100 的 57 倍 。

**Host** [37:09]
这就是 Cerebras 的前三层产品 ： 芯片 、 系统和超级计算机集群 。 最后一层呢 ， 就是我们上一章的时候提到的云端推理业务了 。

**Host 2** [37:19]
2024 年 8 月 ，Cerebras 发布 Cerebras Inference。 客户不用买一台 CS 系统 ，而是直接通过 API 调用云服务 。 至此 ，Cerebras 从一家硬件公司转型为一家推理基础设施公司 。

**Host** [37:38]
2026 年 4 月 17 日，Cerebras 重新提交 S One，5 月 14 日正式上市 。

**Guest 3** [37:44]
And I remember that the day before IPO, the pricing per share, uh, you know, according to Banker, is around $120 to $160-ish per share. And I was like, wow, $160, that's already kind of, uh, yeah, over 25x of the return compared to the CRC price.

I was like, oh, this is already a big home run. But, you know, on the IPO day, when I was just sitting behind the NASDAQ trading desk, I saw that number just going up faster from $200 per share to $300 per share and then to $380 per share.

I was like, wow, this is the craziest IPO demand I've ever witnessed in my entire career.

**Host 2** [38:24]
首日收盘 311 美元 ， 涨幅 68%， 融资规模 55.5 亿美元 ， 市值在盘中一度触及千亿美元 。

**Angela** [38:34]
The IPO is about taking this to the next level and really allowing us to scale. In order for us to grow with how fast the market is growing, we have huge investments to make in our data centers, in our manufacturing, in our product roadmap, and of course in our entire software stack in order to make this product everything that it is.

So, um, that's what I see as, you know, our top priorities for the next year is scale to where the market is and, and distribute our product to as many as we can.

**Host** [39:08]
但可想而知啊 ， 上市并不是 Cerebras 故事的终点 。 上市后呢 ，Cerebras 其实面临着更多的挑战 。

### 未来挑战

**Host 2** [39:15]
首先呢 ， 我们来看供应链的问题 。Cerebras 的整条产品链啊 ， 包括 750 兆瓦的 OpenAI 承诺 ， 全部建立在台积电 5 纳米产能上 。

台积电是目前世界上唯一能够制造 Cerebras 芯片的代工厂 。 那么 Cerebras 在未来能拿到多少市场份额 ， 会不会取决于它能在台积电拿到多少产能呢 ？Andrew 在 Outlaws 播客上啊 ， 没有直接回答这个问题 。他承认呢 ， 台积电当然是 Cerebras 供应链里极其重要的一环 ，但他也说到 Cerebras 有几个意外的优势 ， 比如说啊 ， 它避开了当下 AI 芯片行业最紧张的三个供应瓶颈 ：HBM、 高带宽内

存 、 台积电的先进封装公司 CoWoS 和台积电最先进的 3 纳米产线 。 这三个竞争激烈的供应链瓶颈 ，Cerebras 的产品呢 ， 都用不到 。

**Host** [40:07]
第二是激烈的竞争格局 。

**Host 2** [40:10]
英伟达现在和 Groq 合作 ，有了自己的推理优化技术 。Broadcom 在帮各大云厂商设计定制 ASIC， 谷歌的 TPU、 亚马逊的 Trainium 都在做推理 。

一批新的推理芯片创业公司也正在融资 。

**Host** [40:26]
对此呢 ，Angela 是这样回答我们的 。

**Angela** [40:28]
I feel very confident about Cerebras's outlook because of our full stack capabilities that I've mentioned and our ability to innovate at every level, no matter what's needed in order to, um, continue to meet the market where it's going.

**Host** [40:42]
由竞争带来的第三个风险 ，是重要客户 OpenAI 充满变数的战略走向 。

**Host 2** [40:49]
就在 6 月 ，OpenAI 和 Broadcom 发布了 OpenAI 的第一颗自研推理芯片 Jalapeno。 这是一个专门为语言模型推理设计的 ASIC，也就是专用芯片 。

那 OpenAI 说啊 ，Jalapeno 是它长期全栈 AI 基础设施策略的一部分 。 未来呢 ，OpenAI 想把模型 、 软件 、 网络服务器 ， 甚至是芯片本身 ， 都更深地掌握在自己手里 。

**Host** [41:13]
当然啊 ， 这并不意味着 OpenAI 会立刻放弃 Cerebras，因为 OpenAI 的算力需求太大了 。 但 Cerebras 必须证明 ， 它能够提供的推理速度是长期并且不可替代的产品 。6 月 24 日，Cerebras 发布了上市后的第一份财报 ， 股价下跌 15%， 甚至一度跌破 IPO 发行价 。

**Host 2** [41:35]
表面上看 ， 这份财报并不差 ， 第一季度收入甚至高于华尔街预期 。 但 Cerebras 给出的第二季度指引中， 调整后毛利率只有 36% 到 38%， 低于第一季度的 47%。

也就是说啊 ， 收入增长留下的利润空间正在被压缩 。 对于一家上市公司来说呢 ， 市场看到的是一个现实问题 ， 包括数据中心 、 电力 、 冷却等在内的成本 ， 会不会吃掉技术优势 。

对此 ，Andrew 是这样回应的 。

**Guest** [42:04]
What we did is we put forward a plan in, in the start of '26. We shared it with investors as we went public. And we're ahead of plan. We beat margin, uh, consensus substantially. And then we guided for full year that gross margins would be 10% better than planned.

We also shared that in Q2 and Q3, we, we would go back to, to some of our customers and we would rent back gear that we'd sold them to try and keep up with demand. And that would have a margin impact on the order of 10 or 15 points.

We did that to, to keep our customers close, to be sure we could keep up with their extraordinary demand for, for our product, for fast inference. And so, uh, that, that, that was the, the, the story. Um, on every metric we put out, we're, we're ahead of plan.

**Host** [43:01]
但市场会不会买单啊 ， 就是另一回事了 。 那投资人对于 Cerebras 长期稳定盈利的关注呢 ，其实也是这个故事最有意思的地方 。

它不是一个英伟达挑战者上市的爽文 。Cerebras 的一波三折 ，其实在很大程度上是对 AI 时代基础设施需求变迁的回应 。

当我们对模型的需求从规模变成推理 ， 从能力变成速度 ， 我们到底需要什么样的计算机呢 ？

**Host 2** [43:28]
英伟达的答案是 GPU 集群 ，OpenAI 在自研 ASIC，Groq 选择了 LPU，而 Cerebras 给出的答案是一整片晶圆 。

**Host** [43:39]
十年前它是硅谷疯子 ，而今天这个餐盘一样大小的芯片 ， 就在数据中心里面 ， 服务着一系列重要的客户 。

**Host 2** [43:48]
这个曾经的科幻电影变成了可以登陆纳斯达克的现实 。 真实到资本市场开始用毛利率来审判它 。 这或许才是一个硬科技公司真正进入主流世界的时刻 。

**Host** [44:03]
好了 ，以上就是 Cerebras 崛起的故事 。 我是硅谷 101 的伊雯 ， 大家不要忘记关注我们 ，不要错过来自硅谷的一线分享 。

大家的点赞 、 转发和留言 ，是我们做好深度科技和商业内容的最佳动力 。 那我们就下期再见啦

。

---

本节目库由 PodHood（https://podhood.com）提供支持——播客网站平台。
