nobody really teaches you research. you get a desk, a problem someone else picked, and a vague instruction to produce something novel. so most people reverse-engineer the job from what they can see, which is papers, threads, and announcements, and what they end up learning is how to look like a researcher rather than how to be one. the actual skill is a stack of smaller skills, and almost every one of them can be deliberately trained. 很少有人真正教你做研究。你得到一张办公桌、一个别人选定的问题,以及一句“做出些新颖成果”的模糊指令。于是大多数人通过可见的产出(论文、讨论串和公告)来逆向拆解这份工作,最终学到的只是“如何看起来像研究者”,而非“如何成为研究者”。真正的专业能力是一系列小技能的叠加,而其中每一项都可通过刻意训练来掌握。
pick your own problems 选择你自己的问题
richard hamming had a habit at bell labs that made him unpopular at lunch. he'd ask whoever sat near him what the important problems in their field were, then ask why they weren't working on them. people changed tables. the question stings because most of us have no good answer. we don't choose problems, we absorb them, from an advisor, from whatever a big lab announced last quarter, from the paper everyone is quote-tweeting this week. 理查德·汉明在贝尔实验室有个午餐时令人不快的习惯:他会询问邻座的人所在领域的重要问题是什么,接着追问对方为何不研究这些问题。人们常因此换桌就座。这个问题如此扎心,是因为我们大多数人都给不出像样的答案。我们并非主动选择课题,而是被动接受:从导师那里、从大型实验室上季度公布的课题中、或是从本周被广泛引用转发的论文中。
the trouble with an absorbed problem is that you hold the conclusion without the reasoning. you know some famous lab cares about a direction, but not why, not what they expect to find, not what would make them drop it. when they pivot, you find out a year later. and on a problem that's already fashionable, you're racing a thousand people who started earlier and have more compute than you. 被动接受课题的麻烦在于,你只知结论不知推理。你知道某个著名实验室关注某个方向,却不知其缘由,不知他们期待发现什么,不知何种结果会让他们放弃该课题。当他们转向时,你可能一年后才知晓。而对于已然热门的课题,你要与成千上万更早起步、拥有更多算力的人竞争。
john schulman's guide to ml research splits the work into two modes. in one, you read the literature and hunt for things to improve. in the other, you choose an outcome you genuinely want to exist and reason backwards to the experiments. he argues for the second, and the quiet reason is that it manufactures originality. a goal you actually care about will drag you into territory no survey paper covers. 约翰·舒尔曼的机器学习研究指南将研究工作分为两种模式。一种是阅读文献,寻找可改进之处;另一种是选择一个你真正希望实现的目标,然后反向推导出实验方案。他主张后者,其背后的原因在于这种方法能催生原创性。一个你真正关心的目标会引领你涉足任何综述论文都未曾覆盖的领域。
taste, meanwhile, gets discussed like a gift. it behaves more like a muscle. predict the result of every experiment before you run it. cover a paper's results section and guess the numbers from the method alone. mark down which of this month's releases will matter in two years and check your hit rate later. a forecast plus a correction, repeated a few hundred times, is how every good model gets trained, including the one in your head. 与此同时,“品味”常被视为天赋,但它更像是肌肉。每次实验前都先预测结果,遮住论文的结果部分,仅凭方法猜测数据。标记出本月发布的哪些成果两年后仍有价值,事后验证命中率。预测加上修正,重复数百次——这就是训练每个优秀模型的方法,包括你大脑中的那个。
upgrade your inputs 升级你的输入
shared reading lists produce shared ideas. if your information diet is the trending page of arxiv plus whatever survives the group chat filter, you will reliably reach the same conclusions as everyone else, at the same time, which makes those conclusions worth approximately nothing. 共享的阅读清单会孕育共享的观点。如果你的信息来源仅仅是 arXiv 上的热门推荐加上群聊筛选后的残羹冷炙,那么你注定会和所有人同时得出相同的结论,而这些结论的价值近乎为零。
old material is criminally underpriced. this field reruns its own past on a delay: mixture of experts dates to 1991, lstms to 1997, backprop went mainstream in 1986. rich sutton needed about a thousand words in 2019 to write the bitter lesson, and it predicts the shape of the field better than surveys ten times its length. claude shannon gave a talk on creative thinking in 1952 where his opening move was to shrink a problem until it's nearly trivial, crack the small version, then reintroduce the difficulty one piece at a time. that single trick will carry you through more walls than any modern productivity advice. 旧知识的价值被严重低估了。这个领域总在以滞后的方式重演过去:混合专家模型(Mixed of Experts)可以追溯到 1991 年,LSTM 可追溯到 1997 年,反向传播算法(Backprop)在 1986 年就已成为主流。理查德·萨顿(Rich Sutton)在 2019 年仅用千字写就《苦涩的教训》,其对领域未来走向的预见性,胜过篇幅十倍于它的综述文章。克劳德·香农(Claude Shannon)在 1952 年关于创造性思维的演讲中,开场演示是将问题缩减至近乎平凡,破解简化版本后再逐步重新引入难度。仅此一招,比任何现代生产力建议更能助你突破思维壁垒。
range matters as much as depth. interpretability borrows shamelessly from neuroscience. eval design is mechanism design wearing a lab coat. a working sense of how gpus actually move memory tells you which architecture papers are doomed before the benchmarks do. and honest statistics might be the rarest skill in ml, where a lot of published rigor is vibes with error bars. 知识广度与深度同等重要。可解释性研究毫不避讳地借鉴了神经科学。评估设计本质上就是穿着实验服的机制设计。对 GPU 内存运作原理的真切理解,能在基准测试之前就预判哪些架构论文注定失败。而扎实的统计学素养或许是机器学习领域最稀缺的技能,因为很多已发表的严谨性研究不过是带误差线的主观臆断。
one more thing. read the paper itself, not the thread summarizing it. the appendix is where the bodies are buried, and the limitations section is usually the most honest paragraph in the document. 还有一点:务必阅读论文原文,而非总结推文。附录才是真正埋藏细节之处,而局限性章节通常是全文最诚实的段落。
write everything down 把一切都写下来
paul graham points out that an idea can feel fully formed right up until you try to put it into words. the page finds gaps your head papers over: the assumption you never tested, the step that doesn't actually follow, the two claims that quietly contradict each other. 保罗·格雷厄姆指出,一个想法在被付诸文字之前,往往让人觉得已经成熟完整。但文字会暴露那些被大脑轻易忽略的缺口:你从未验证的假设、无法自圆其说的步骤,以及两个悄然矛盾的观点。
feynman's rule was that the first person you must avoid fooling is yourself, because you're the easiest target. writing is the cheapest defense ever invented. darwin went further and made it procedural. any fact that cut against his theory got written down on the spot, because he'd caught his own memory deleting inconvenient evidence faster than the convenient kind. your memory does the same thing to your failed runs. keep a log: hypothesis, setup, expectation, result, updated belief. rereading last month's entries is humbling in a way no reviewer can match. 费曼的原则是,你必须避免欺骗的首要对象是自己,因为你最容易被蒙蔽。写作是人类发明的最廉价的防线。达尔文更进一步,将其程序化:任何与自己理论相悖的事实都被当场记录,因为他发现记忆删除不利证据的速度总比便利证据更快。你的记忆也会对你失败的尝试做同样的事情。所以,要做好记录:假设、设置、预期、结果、更新后的认知。重读上个月的日志记录,会带来一种任何审稿人的批评都相形见绌的谦卑感觉。
then put some of it in public. olah and carter's research debt essay makes the case that fields choke on undigested ideas, and that a clear explanation is a genuine contribution rather than a service job. a lot of people working in interpretability today found the field through readable posts, not conference papers. a body of public writing also doubles as the strongest credential you can hold, because it's an unfakeable sample of how you think. 然后把其中一部分公开出来。Olah 和 Carter 的《学术债务》一文指出:许多领域因未消化的观点而窒息,而清晰的阐释本身就是实质性的学术贡献,并非服务性工作。当今可解释性领域的众多研究者,正是通过可读的公开文章而非会议论文接触到该领域。此外,系统的公开写作同时也是你所能拥有的最有力资质证明,因为它是你思维方式的不可伪造的样本。
tighten the loop 收紧环路
the stories about alec radford rarely involve a single stroke of genius. they involve volume. more runs per day, more wrong ideas discarded per week, a model of reality that updated faster than anyone else's. that's the actual game. research speed is mostly the speed at which you discover you're wrong. 关于亚历克·拉德福德的故事很少提及灵光一现的天才时刻,它们讲述的是数量。每天更多次的尝试,每周淘汰更多错误的想法,建立比其他所有人都更及时更新的现实模型。这才是真正的较量。研究的速度,本质上是你发现自己错了的速度。
which makes tooling a first-class research activity. launching a run should be one command. plotting it should be one more. every experiment should be reproducible from its config, and comparing two runs should take seconds, not an afternoon of archaeology. karpathy's recipe for training neural networks has a step that pays for itself a hundred times over: overfit a single batch before training at scale. thirty seconds, half your bugs, gone. shrink everything until it's cheap, get it right, then spend the compute. 这使得工具开发成为首要研究活动。启动实验运行应只需一条命令,绘制图表再加一个命令。每个实验都应从配置文件复现,两次运行的比较应几秒完成,而非耗费大半天考古式钻研。Karpathy 的神经网络训练秘诀中有一步能带来百倍收益:在大规模训练前先对单个批次过拟合。三十秒,解决一半错误,搞定。将一切压缩到成本最低,确保正确,然后投入算力。
and retire the idea that engineering is the junior partner here. at the frontier the two jobs have fused. the researcher who can build the harness, the eval, and the data pipeline is the one whose hypotheses actually get tested. everyone else is waiting in a queue. 彻底摒弃「工程研究仅是附属角色」的陈旧观念。在探索前沿,二者已融为一体。那些能构建实验平台、评估体系与数据管道的研究员,其假设才真正得以验证。其余人则都在排队等待。
stare at the outputs 审视输出结果
a descending loss curve is not analysis, it's reassurance. your experiments throw off far more information than you consume: transcripts, failure cases, the strange tail of the distribution. most of it dies unread in a logs folder. 下降的损失曲线并非分析,而是安慰剂。你的实验抛出的信息远比你所能处理的要多:录音转录、失败案例、异常分布的长尾。它们大多在日志文件夹中无人问津地死去。
karpathy's recipe starts before any training code gets written, with hours spent on the raw data by hand. most ml bugs live in the data, and they fail silently. nothing crashes. you simply get a mediocre model and a wrong theory about why. Karpathy 的秘诀始于任何训练代码编写之前,需要花费数小时手工处理原始数据。大多数机器学习错误隐藏在数据中,且悄无声息地失效。程序不会崩溃,你只会得到一个平庸的模型,以及关于其原因的错误理论解释。
andrew ng has taught the same unglamorous move for over a decade because nothing beats it. pull a hundred failures, read all of them, sort them into piles, attack the biggest pile. it works on models and it works on evals, where a benchmark you've never read transcripts from is a benchmark you don't actually understand. one transcript of genuinely strange behavior will teach you more than the next decimal of accuracy ever will. 吴恩达坚持传授这一不起眼的招数已逾十年,因为没有比这更有效的方法了。把一百个失败案例拉出来全部阅读,将它们分成几堆,然后集中攻克最大的那堆。这套方法对模型有效,对评估同样有效。若你从未阅读过某基准测试的原始记录,那你实际上并未理解该基准。一份真实展现异常行为的记录,其教学价值远超模型准确率小数点后再多一位的提升。
wander on purpose 故意闲逛
your first subfield is an accident of timing, so treat it like one. spend real time in interpretability, in evals, in rl, in systems, before deciding where you live. somewhere in this field is a corner where your specific weirdness is an unfair advantage, and the only way to locate it is to pay tuition in several places. nobody waives the tuition. 你选择的第一个研究领域往往是时机使然,不妨就把它当成一次偶然。在确定最终方向之前,不妨在可解释性、评估体系、强化学习、系统设计等领域都花些真功夫深入探索。这个领域中总有某个角落,你的独特特质会成为不公平的竞争优势。而找到那个角落的唯一方法,就是在多个领域先“交学费”。没有人能免除这笔学费。
run the disposable version of every idea first and let most of them die young. tune your baselines until it hurts, because the graveyard of ml is full of gains that evaporated against a properly tuned baseline, and a reviewer is the worst possible person to learn that from. ablate until you know which component carries the result. it's usually one, and it's usually not the one in the title. 先为每个想法运行其试用版,让大多数想法夭折。反复调校你的基准线直至近乎苛刻,因为机器学习的坟墓里满是那些在严苛基准面前烟消云散的成果,而审稿人正是最不该让你从他们那里学到这点的人。逐一排除直到你明确哪个组件承载了结果。往往只有一个,但往往不是标题里的那个。
breadth is also insurance. subfields saturate, all of them, usually right after they peak on twitter. the people who keep producing through those transitions are the ones who already know their way around the neighboring territory. 广度也是一种保障。所有子领域都会饱和,通常是在它们在推特上达到热度顶峰之后。那些能够持续产出成果的人,往往是已经熟悉邻近领域情况的人。
find your people 找到你的同路人
hamming noticed a pattern in who ended up doing important work. colleagues with closed office doors got more done in any given year, and colleagues with open doors did the work that mattered, because the interruptions carried information about what the world actually needed. your open door is probably an inbox. keep it that way. 汉明注意到一个现象:最终完成重要工作的人群中,关门办公的同事每年能产出更多成果,而开门办公的同事则完成了真正重要的工作。因为那些打扰中包含了世界实际需求的信息。你的开放办公室很可能就是个收件箱,保持这个状态就好。
generosity compounds in research like nothing else. replicate a result and publish what you find. release the tool you built for yourself. explain something hard in plain language. the returns arrive sideways, months later, as the collaboration or the reference or the role you couldn't have applied for. float your half-formed ideas in public too, because being wrong on the timeline is far cheaper than being wrong in print. and the collaborator who tells you an idea is bad before you sink three months into it is worth more than compute. that relationship can't be bought, only earned. 慷慨在研究领域会复利累积,胜过其他任何事。复现结果并发表你的发现,发布你为自己打造的工具,用平实语言阐释艰深概念。回报会以迂回方式降临:数月后化作合作机会、推荐引荐或你原本无从争取的职位。也请将半成品想法抛入公共领域,因为时间线上的错误远比印刷品的谬误代价低廉。而那个在你投入三个月前就直言想法糟糕的合作者,比算力更珍贵。这种关系无法购买,只能赢得。
the long game 长期博弈
pasteur said luck favors the prepared mind, and hamming built a whole career philosophy on top of it: knowledge and productivity compound like interest. the daily edges look trivial in isolation. what you read, what you record, how fast your loop runs, who you argue with. give them a few years and they produce careers that look like luck from the outside. start compounding earlier than feels necessary. future you already knows this was the cheap part. 巴斯德曾言,幸运眷顾有准备之人。汉明则以此为基石构筑了毕生信条:知识与生产力如利息般复利增长。日常的边际收益孤立来看似乎微不足道。你阅读的文献、记录的笔记、工作流程的迭代效率、思辨交流的对手。假以数年,这些积累将孕育出旁观者眼中如同天赐的事业成就。要在感觉“为时过早”之前就开始积累复利,未来的自己会知道这不过是最初成本最低的投资阶段。