Maggie Appleton 写设计、人类学和智能体界面时,习惯先画清楚隐喻,再谈工具。把「Humanity's Last Exam」整理成可阅读的中文笔记:问题在哪、界面默认会把人带去哪、落地时该改什么。原站导航、评论回响和广告已去掉。
问题怎么被看见
We have a new(ish) Okay, it’s not that new – created in September 2024 – but we’ve only recently seen companies using when they announce new models. benchmark, cutely named “Humanity’s Last Exam.”
If you’re not familiar with benchmarks, they’re how we measure the capabilities of particular AI models like o1 or Claude Sonnet 3.5. Each one is a standardised test designed to check a specific skill set.
默认界面在强化什么
When you run a model on a benchmark it gets a score, which allows us to create leaderboards showing which model is currently the best for that test. To make scoring easy, the answers are usually formatted as multiple choice, true/false, or unit tests for programming tasks.
Among the many problems with using benchmarks as a stand-in for “intelligence” (other than the fact they’re multiple choice standardised tests – do you think that’s a reasonable measure of human capabilities in the real world?), is that our current benchmarks aren’t hard enough.
可以怎么改
New models routinely achieve 90%+ on the best ones we have. So there’s a clear need for harder benchmarks to measure model performance against.
Made by ScaleAI and the Center for AI Safety, they’ve crowdsourced “the hardest and broadest set of questions ever” by experts across domains. 2,700 questions at the moment, some of which they’re keeping private to prevent future models training on the dataset and memorising answers ahead of time. Questions like this:
落到产品里的动作
So far, it’s doing it’s job well – the highest scoring model is OpenAI’s Deep Research at 26.6%, with other common models like GPT-4o, Grok, and Claude only getting 3-4% correct. Maybe it’ll last a year before we have to design the next “last exam.”
When people make sweeping statements like “language models are bullshit machines” or “ChatGPT lies,” it usually tells me they’re not seriously engaged in any kind of AI/ML work or productive discourse in this space.
值得单独记下的观察
- MMLU (Massive Multitask Language Understanding) measures understanding across 57 academic subjects including STEM, social science, and the humanities.
- HumanEval measures code generation skills.
- GPQA (Graduate-Level Google-Proof Q&A Benchmark) measures correctness on a set of questions written by PhD students and domain experts in biology, physics, and chemistry.
落地时建议先做的 5 件事
- 先写清这个工具在强化思考还是在替人思考。
- 默认交互不要只会谄媚:该追问、该给反例、该要求证据。
- 智能体界面要露出推理步骤,而不是只给一个光滑答案。
- 知识发布用可生长的笔记,而不是一次性营销长文。
- 改产品前先画隐喻:用户以为自己在做什么,系统实际在做什么。
和智能体产品怎么接
龙虾PRO做 OpenClaw 落地时,最该从这类笔记里拿走的是「别把聊天框当唯一界面」。数字员工要能追问、能验证、能把过程摊开,而不是只负责说好话。
本文侧重全链路风控方法论。落地时请用自身业务单据做回放验证,不要把示例阈值直接当生产策略。 相关:风控体检 · 方案资源
常见问题 FAQ
什么是AI智能系统?
「AI智能系统」可概括为:We have a new(ish) Okay, it’s not that new – created in September 2024 – but we’ve only recently seen companies using when they announce new models. benchmark, cutely named “Humani 本文从定义、方法与实践要点展开说明。
为什么要关注AI智能系统?
关注AI智能系统,是因为它直接影响效率、风险与可复制性。文中指出:We have a new(ish) Okay, it’s not that new – created in September 2024 – but we’ve only recently seen companies using when they announce new models. benchmark, cutely named “Humanity’s Last Exam.”
如何落地AI智能系统?有哪些关键步骤?
建议按以下路径推进AI智能系统:1) MMLU (Massive Multitask Language Understanding) measures understanding across 5…;2) HumanEval measures code generation skills.;3) 先写清这个工具在强化思考还是在替人思考。;4) 默认交互不要只会谄媚:该追问、该给反例、该要求证据。;5) 智能体界面要露出推理步骤,而不是只给一个光滑答案。。细节见正文对应章节。
AI智能系统适合哪些人或团队?
AI智能系统更适合:产品/技术负责人、运营与增长团队、需要落地智能体或自动化的中小团队、关注「AI智能系统」方向的读者。若你只需要单次聊天式问答,可先读概念;若要上生产,请重点看步骤、权限与风控相关段落。
关于「问题怎么被看见」,本文给出了什么结论?
在「问题怎么被看见」部分,要点是:named “Humanity’s Last Exam.” If you’re not familiar with benchmarks, they’re how we measure the capabilities of particular AI models like o1 or Claude Sonnet 3.5. Each one is a standardised test designed to check a spec
关于「默认界面在强化什么」,本文给出了什么结论?
在「默认界面在强化什么」部分,要点是:atted as multiple choice, true/false, or unit tests for programming tasks. Among the many problems with using benchmarks as a stand-in for “intelligence” (other than the fact they’re multiple choice standardised tests –