Dan Luu 写系统问题时,习惯先测量、再对照、最后才下结论。把「Why use ECC?」放到智能体、评测和线上系统里,真正要问的是:默认做法会不会系统性失败。下面用中文整理成可执行的工程笔记,去掉原站导航和无关链接。
1. Google didn't use ECC in 1999
Jeff Atwood, perhaps the most widely read programming blogger, has a post that makes a case against using ECC memory . My read is that his major points are:
Not too long after Google put these non-ECC machines into production, they realized this was a serious error and not worth the cost savings. If you think cargo culting what Google does is a good idea because it's Google, here are some things you might do:
2. Most RAM errors are hard errors
Articles are still written today about what a great idea this is, even though this was an experiment at Google that was deemed unsuccessful. Turns out, even Google's experiments don't always succeed. In fact, their propensity for “moonshots” in the early days meannt that they had more failed experiments that most companies. Copying their failed experiments isn't a particularly good strategy.
Part of the post talks about how awesome these servers are:
3. Due to advances in hardware manufacturing, errors are very rare
Some people might look at these early Google servers and see an amateurish fire hazard. Not me. I see a prescient understanding of how inexpensive commodity hardware would shape today's internet. I felt right at home when I saw this server; it's exactly what I would have done in the same circumstances
The last part of that is true. But the first part has a grain of truth, too. When Google started designing their own boards, one generation had a regrowth 1 issue that caused a non-zero number of fires.
4. If ECC were actually important, it would be used everywhere and not
BTW, if you click through to Jeff's post and look at the photo that the quote refers to, you'll see that the boards have a lot of flex in them. That caused problems and was fixed in the next generation. You can also observe that the cabling is quite messy, which also caused problems, and was also fixed in the next generation. There were other problems as well. Jeff's argument here appears to be that, if he were there at the time, he would've seen the exact
One generation of Google servers had infamously sharp edges, giving them the reputation of being made of “razor blades and hate”.
Conclusion
From talking to folks at a lot of large tech companies, it seems that most of them have had a climate control issue resulting in clouds or fog in their datacenters. You might call this a clever plan by Google to reproduce Seattle weather so they can poach MS employees. Alternately, it might be a plan to create literal cloud computing. Or maybe not.
Note that these are all things Google tried and then changed. Making mistakes and then fixing them is common in every successful engineering organization. If you're going to cargo cult an engineering practice, you should at least cargo cult current engineering practices, not something that was done in 1999 .
Appendix: security
When Google used servers without ECC back in 1999, they found a number of symptoms that were ultimately due to memory corruption, including a search index that returned effectively random results to queries. The actual failure mode here is instructive. I often hear that it's ok to ignore ECC on these machines because it's ok to have errors in individual results. But even when you can tolerate occasional errors, ignoring errors means that you're exposing yo
Google has great infrastructure. From what I've heard of the infra at other large tech companies, Google's sounds like the best in the world. But that doesn't mean that you should copy everything they do. Even if you look at their good ideas, it doesn't make sense for most companies to copy them. They created a replacement for Linux's work stealing scheduler that uses both hardware run-time information and static traces to allow them to take advantage of n
值得单独记下的观察
- Google didn't use ECC when they built their servers in 1999
- Most RAM errors are hard errors and not soft errors
- RAM errors are rare because hardware has improved
- If ECC were actually important, it would be used everywhere and not just servers. Paying for optional stuff like this is "awfully enterprisey"
落地时建议先做的 5 件事
- 用自己的真实负载测,而不是只用公开榜或厂商数字。
- 把评测设计成能抓到失败模式:平均分好看但尾部崩溃,仍然算失败。
- 智能体默认不会好好用测试;要写进流程,而不是写在口头规范里。
- 性能和正确性都要有基线,改模型或改语言前后必须能对比。
- 结论写成可回滚的决策:哪一版配置、哪一版评测集、谁签字。
和智能体产品怎么接
龙虾PRO做 OpenClaw 落地时,同样吃「先测量再扩面」这条纪律:技能、数字员工和网关都要有可复现评测,而不是只看一次演示通过。
本文侧重全链路风控方法论。落地时请用自身业务单据做回放验证,不要把示例阈值直接当生产策略。 相关:风控体检 · 方案资源
常见问题 FAQ
什么是AI智能系统?
「AI智能系统」可概括为:Jeff Atwood, perhaps the most widely read programming blogger, has a post that makes a case against using ECC memory . My read is that his major points are: 本文从定义、方法与实践要点展开说明。
为什么要关注AI智能系统?
关注AI智能系统,是因为它直接影响效率、风险与可复制性。文中指出:Jeff Atwood, perhaps the most widely read programming blogger, has a post that makes a case against using ECC memory . My read is that his major points are:
如何落地AI智能系统?有哪些关键步骤?
建议按以下路径推进AI智能系统:1) Google didn't use ECC when they built their servers in 1999;2) Most RAM errors are hard errors and not soft errors;3) RAM errors are rare because hardware has improved;4) If ECC were actually important, it would be used everywhere and not just server…;5) 用自己的真实负载测,而不是只用公开榜或厂商数字。。细节见正文对应章节。
AI智能系统适合哪些人或团队?
AI智能系统更适合:产品/技术负责人、运营与增长团队、需要落地智能体或自动化的中小团队、关注「AI智能系统」方向的读者。若你只需要单次聊天式问答,可先读概念;若要上生产,请重点看步骤、权限与风控相关段落。
关于「1. Google didn't use ECC in 1999」,本文给出了什么结论?
「1. Google didn't use ECC in 1999」是理解全文的关键切片:建议先读该节的结论句与列表项,再对照前后章节形成闭环。
关于「2. Most RAM errors are hard errors」,本文给出了什么结论?
在「2. Most RAM errors are hard errors」部分,要点是:'t always succeed. In fact, their propensity for “moonshots” in the early days meannt that they had more failed experiments that most companies. Copying their failed experiments isn't a particularly good strategy. Part o