-
自建专家值不值得:外包省不下判断
An alternate title for this post might be, “Twitter has a kernel team!?”…
-
读懂「Distinguished Lecture Series II: Charles Fefferm」:先把定义、反例和适用范围钉死
On Thursday, Charlie Fefferman continued his lecture series on interpola…
-
智能体栈里的Multi-Head Attention:延迟、成本和翻车点
单头 attention 只有一组 softmax 权重,只能在一种相似度度量下做一次聚合。Multi-Head Attention 通过多套独…
-
2026机器人量产落地实操指南 4个核心策略搞定供应链与基建痛点
机器人商业化量产落地,核心不在于极致的单机技术优化,而在于供应链降本、无改造部署、场景复用、数据迭代四大核心维度的平衡。摒弃定制化研发、前置基建…
-
为什么永远无法就「允许什么」达成一致
On large platforms, it’s impossible to have policies on things like mode…
-
「Distinguished Lecture Series III: Charles Feffer」笔记:问题从哪来、证明卡在哪
Today, Charlie wrapped up several loose ends in his lectures, including …
-
智能体栈里的Scaled Dot-Product:延迟、成本和翻车点
> 本文从零推导注意力机制点积方差的来源,解释缩放因子如何防范梯度弥散,并作为大模型 Scaling Laws 数值稳定的基石。
-
HN 评论被低估了
HN comments are terrible . On any topic I’m informed about , the vast ma…
-
「PCM article: Compactness and compactification」笔记:问题从哪来、证明卡在哪
I’m continuing my series of articles for the Princeton Companion to Math…
-
智能体栈里的Self-Attention:延迟、成本和翻车点
从 cross-attention 到 self-attention 的退化路径 → 为什么 self-attention 是 O(1) 跳数 …