Every Problem Solvable, No One to Conjecture万题可解,无人去猜
On the Value of Conjecture论猜想的价值
On May 20, 2026, an unreleased model overturned a conjecture from 1946.2026 年 5 月 20 日,一个未公开的模型推翻了一个 1946 年的猜想。
The question Paul Erdős asked back then was simple enough to write on a napkin: place n points in the plane — at most how many pairs of them can be exactly distance 1 apart? He gave a construction for the lower bound and guessed the true answer lay close to it. For eighty years, nobody proved it and nobody refuted it. That day OpenAI announced: the model had produced a counterexample, using tools from algebra and number theory — a direction no one had ever successfully carried into this geometric problem. The Harvard mathematician who joined the external review called it a beautiful piece of mathematics. Within weeks, human mathematicians had substantially improved the result.保罗·厄多斯当年问的问题简单到能写在餐巾纸上:平面上放 n 个点,最多有多少对点之间的距离恰好是 1。他给出了一个下界的构造,猜测真实答案离这个下界很近。八十年里,没有人证明它,也没有人推翻它。那天 OpenAI 公布:模型给出了反例,用的是代数与数论里的工具——一个此前从没有人成功搬进这个几何问题的方向。参与外部评审的哈佛数学家说这是一段漂亮的数学。几周之内,人类数学家又把这个结果大幅改进了。
All the coverage was about the machine. How smart it is, whether it truly understands, whether its proof counts as creation.所有的报道都在谈机器。它多聪明,它是不是真的理解,它的证明算不算创造。
I want to talk about something else.我想谈另一件事。
Erdős guessed wrong.厄多斯猜错了。
The conjecture was false. It has been sentenced to death. Yet in the eighty years it lived, it organized the labor of an entire subfield — the best minds in discrete geometry, across generations, working around this one sentence. The methods, bounds, counterexamples, and byproducts it spawned outweigh most theorems that were ever proven true, combined. Today it is dead, and it is one of the most valuable statements in twentieth-century discrete geometry.这个猜想是假的。它被判了死刑。可是在它活着的八十年里,它组织了一整个子领域的劳动——离散几何里最好的一批头脑,几代人,围着这一句话工作。它催生的方法、边界、反例和副产品,比绝大多数被证明为真的定理加起来还多。今天它死了,而它是二十世纪离散几何里最有价值的陈述之一。
An erroneous statement, worth eighty years.一个错误的陈述,值八十年。
That sentence punches through one thing: if value equals correctness, if value equals verifiability, those eighty years are inexplicable.这句话把一件事击穿了:如果价值等于正确性,如果价值等于可验证性,那么这八十年无法解释。
So our whole way of pricing research value is wrong at the root.所以我们对研究价值的整个定价方式,从根上就是错的。
◇
I. Liquidation: One Man's Conjectures, a Whole Industry's Firing Range一、清仓:一个人的猜想,一整个产业的靶场
The unit-distance conjecture is no isolated case. It was the loudest shot in a general clearance.那个单位距离猜想不是孤例。它是一场清扫里最响的一枪。
Erdős posed thousands of problems in his lifetime. Later, someone collected more than twelve hundred of them into a public database, erdosproblems.com, each problem tagged with a status: open, or solved.厄多斯一生提了几千个问题。后来有人把其中一千二百多道整理成一个公开数据库,erdosproblems.com,每道题标一个状态:未解,或已解。
That website is now the hottest firing range in the world.这个网站现在是全世界最热的一块靶场。
The timeline runs roughly like this. After October 2025, activity on the site exploded — by a page Terence Tao maintains, roughly a hundred Erdős problems have since moved into the "solved" column, a good share with AI involvement. On January 4, 2026, problem #728 was solved by GPT-5.2 Pro; the participants used Harmonic's Aristotle to formalize and verify the proof in Lean, and Tao vouched for its autonomous character — this is widely regarded as the first non-trivial Erdős problem completed autonomously by an AI. The same month, a twenty-four-person team led by DeepMind announced: four solved, and forgotten old solutions recovered for nine more. Another team of twenty-one reported that their strongest agent autonomously solved nine out of three hundred fifty-three open problems, at a few hundred dollars per problem. May 20 was the unit-distance conjecture. On August 1, ten more advances, including three Erdős problems.时间线大致是这样。2025 年 10 月之后,站上的活动开始暴涨——据陶哲轩维护的一个页面统计,此后大约一百道厄多斯问题被移进了"已解"一栏,其中相当一部分有 AI 参与。2026 年 1 月 4 日,第 728 号问题被 GPT-5.2 Pro 解出,参与者用 Harmonic 的 Aristotle 在 Lean 里把证明形式化并验证通过,陶哲轩为它的自主性质背书——这被普遍视为第一个由 AI 自主完成的、非平凡的厄多斯问题。同月,一个由 DeepMind 领衔的二十四人团队公布:解掉四道,另外为九道找回了被遗忘的旧解。另一个二十一人的团队报告:他们最强的智能体在三百五十三道开放问题里自主解掉九道,单题成本几百美元。5 月 20 日是那个单位距离猜想。8 月 1 日,又是十项推进,其中包含三道厄多斯问题。
As of today, of the twelve-hundred-plus problems in the database, nearly half are solved.到今天,这个数据库里一千二百多道题,已解的接近一半。
Three things deserve to be pulled out separately.三件事值得单独拎出来。
First, the strongest models are not the main force. What surprised the database's maintainers: although OpenAI, DeepMind, and a crowd of startups are all watching this range, most of the new results come from amateurs and undergraduates, using public models anyone can buy. Which means the bottleneck was never on the capability side. The bottleneck is whether anyone thinks to feed a given problem in.第一,最强的模型不是主力。 让数据库的维护者感到意外的是:尽管 OpenAI、DeepMind 和一批创业公司都在盯着这块靶场,大部分新结果却来自业余爱好者和本科生,用的是任何人都能买到的公开模型。这说明瓶颈从来不在能力那一侧。瓶颈在有没有人想到把哪一道题喂进去。
Second, the accuracy rate is frighteningly low — and verification is being taken over by machines too. In one large-scale sweep, DeepMind reported that of two hundred candidate solutions flagged right-or-wrong, only 6.5% were meaningfully correct. One of the most active participants admitted his own mathematics was not up to verifying the solutions he had coaxed out of the models; he could only look for willing mathematicians to check them. And the answer to that link in the chain has already appeared: a Lean formal proof containing no sorry is the insurance policy — even "checking whether it's right" is being eaten by formal systems.第二,正确率低得吓人,而验证也在被机器接管。 DeepMind 在一次大规模扫描里报告:两百个被标记为正确或错误的候选解当中,只有 6.5% 是有意义地正确的。一位最活跃的参与者坦承,他自己的数学水平不足以验证他从模型里哄出来的那些解,只能去找愿意帮忙的数学家来查。而这一环的答案也已经出现了:一个不含 sorry 的 Lean 形式化证明,就是那道保险——连"查对不对"这件事,也正在被形式化系统吃掉。
Third — and this is the vital point — this firing range is a dead man's list.第三,也是要害——这块靶场,是一个死人列的清单。
728, 1051, 1196, 1217, 92, 164, 401, 729, 858… Behind every number is a sentence one man guessed out by hand, one at a time, across decades. He could not solve them. He could only pose them — and then, out of his own pocket, put a price on each.728、1051、1196、1217、92、164、401、729、858……每一个编号背后,都是一个人在几十年里,一道一道亲手猜出来的句子。他解决不了它们。他只能提出它们,然后自掏腰包给每一道标上价钱。
Now that list is being swept as a benchmark by an entire industry; the cost per problem has fallen to a few hundred dollars; every lab has clocked that the site can serve as a benchmark; the remaining open problems have become targets in a public race.现在这份清单被整个行业当成基准在清扫,单题成本已经降到几百美元,各家实验室都明确意识到这个网站可以当 benchmark 用,剩下的开放问题成了一场公开竞赛的靶子。
So:那么:
What happens when the sweeping is done?清扫完了以后呢?
Twelve hundred problems; half already fallen. At current speed and cost, the other half will not take another eighty years. And no one is replenishing the list.一千二百道题,已经掉了一半。以现在的速率和成本,剩下的一半不会再用八十年。而没有人在补充这份清单。
We are consuming, at industrial speed, the conjecture inventory of one man's lifetime — and meanwhile, not a single new one is in production.我们正在以工业化的速度,消耗一个人一生的猜想库存;与此同时,一道新的也没有在生产。
That is the literal meaning of "every problem solvable, no one to conjecture." It is not a figure of speech.这就是"万题可解,无人去猜"的字面意思。它不是修辞。
◇
II. The Machine's Boundary: The Verifiable Is Destined to Be Eaten, and Machines Do Not Conjecture二、机器的边界:可验证的注定被吃掉,而机器不会猜
First, the most counterintuitive step, because nearly everyone's first reaction points the wrong way.先把最反直觉的那一步说清楚,因为几乎所有人的第一反应都反了。
Watching the flood of papers, the collapse of peer review, AI-generated reviews reviewing AI-generated papers, one's first reaction is: verification will become the new scarcity. If generation is free, then culling the false becomes the only act worth money.看到论文洪水、评审失守、AI 生成的评审去评审 AI 生成的论文,人的第一反应是:验证会变成新的稀缺品。生成免费了,那么把假的剔出去就成了唯一值钱的动作。
For the next three to five years that judgment is correct, and as a business it is correct. But as a long-term position it is wrong — dangerously wrong.这个判断在未来三到五年是对的,作为一门生意也是对的。但作为一个长期立场,它是错的,而且错得很危险。
Because verification is an act with a well-defined objective function. And in this era, every act with a well-defined objective function is standing in the automation queue — the only question is its place in line. That Lean proof with no sorry from the last chapter is the queue's progress bar.因为验证是一个有明确目标函数的动作。而在这个时代,凡是有明确目标函数的动作,都在被自动化的队列里排队,只是先后而已。上一章里那个不含 sorry 的 Lean 证明,就是这条队伍的进度条。
The order of automation's advance says the same. It does not advance along "hard"; it advances along verifiability. Mathematics and formal proof first, because judging right from wrong is nearly free; domains with reliable simulators next; domains where the physical world must speak for itself, last. Erdős's problem became the first historic named problem to fall not because it was easy — eighty years — but because once its answer appeared, humans could confirm it within days.看自动化的推进顺序也知道。它不沿着"难"推进,它沿着可验证性推进。数学和形式化证明最先,因为对错的判定近乎免费;有可靠模拟器的领域其次;必须让物质世界亲自开口的领域最后。厄多斯的问题之所以是第一个倒下的历史性名题,不是因为它简单——八十年——而是因为它的答案一旦出现,人类可以在几天之内确认它。
This whole mechanism even has an engineering name. In recent years, the explosion in model capability has concentrated almost entirely in domains with verifiable rewards: math, code, reasoning with reference answers. The direction of the entire field is selected not by "what matters" but by "what can be scored." Wherever scoring is possible, everyone swarms; wherever scoring is hard, whole regions are abandoned.这一整套机制在工程上还有个名字。过去几年,模型能力的爆发几乎全部集中在有可检验奖励的领域:数学、代码、有标准答案的推理。整个领域的方向不是被"什么重要"选出来的,是被"什么能被打分"选出来的。凡是能打分的地方,所有人蜂拥而至;凡是难打分的地方,被整片放弃。
The conclusion, then, is cold:于是结论是冷的:
Everything you can verify, the machine will eventually verify for you. Every researcher who camps inside verifiability is camping on the machine's line of march.你能验证的一切,机器终将替你验证。凡是把自己安放在可验证性之内的研究者,都是在机器的行进路线上扎营。
And what of the "new" that machines themselves produce?那机器自己产出的"新"呢?
In March 2026, a group of researchers proposed a framework called "the alien space of science." Their observation: scientific discovery is constrained not only by "what is true," but by what is cognitively accessible to the current population of researchers. And language models inherit that bias wholesale — asked to think of new ideas, they interpolate between the literature's high-density regions, recombining fashionable concepts and recurring patterns.2026 年 3 月,一组研究者提出了一个叫"科学的异乡"的框架。他们的观察是:科学发现受限的不只是"什么是真的",还有什么对当前这批研究者是认知上可及的。而语言模型完整继承了这个偏置——被要求想新点子时,它在文献的高密度区域之间插值,在时髦概念和反复出现的套路之间重组。
Their remedy is itself the verdict: split papers into "idea atoms," train two models — one to learn which combinations are coherent, one to learn which combinations the existing research community might plausibly propose — then sample by maximizing coherence and minimizing availability.他们的解法本身就是判词:把论文拆成"想法原子",训两个模型,一个学哪些组合是连贯的,一个学哪些组合是现有研究者群体可能提出的,然后最大化连贯性、最小化可用性去采样。
In other words: to make a machine novel, you must first compute what humans would not think of, and then aim there deliberately. Novelty cannot be sampled out of the distribution.也就是说:要让机器新,你得先算出人类不会想什么,然后专门往那里去。新颖性无法从分布里采样出来。
And the more fundamental reason hides in a detail almost every report skipped.而更根本的原因,在一个几乎所有报道都跳过的细节里。
On that Erdős problem, mathematicians broadly believed the conjecture was true. The model did not go prove it — the model went hunting for a counterexample. In hindsight, its path looked like extremely stubborn search and optimization, not any kind of insight.在厄多斯那个问题上,数学家们普遍相信猜想是真的。而模型没有去证明它,模型去找反例了。事后看,模型的路径更像是极其顽固的搜索与优化,而不是某种洞察。
The reason is simple: the model has no beliefs.原因很简单:模型没有信念。
It does not believe the conjecture is true, nor that it is false. It has no position to be overturned, no reputation to forfeit, no eighty years to waste. It merely exhausts a solution space someone has already defined.它不相信这个猜想是真的,也不相信它是假的。它没有立场可以被推翻,没有名声可以被赔掉,没有八十年可以被浪费。它只是在一个已经被定义好的解空间里穷尽。
And conjecture requires belief. The definition of belief is: staking yourself on one side when the evidence is insufficient.而猜想需要信念。信念的定义就是:在证据不足的情况下,把自己押在某一边。
This the machine cannot do — not for lack of capability, but because it has nothing to lose. What cannot lose does not bet; what does not bet does not conjecture.这件事机器做不到,不是因为能力不够,而是因为它没有可以输的东西。一个不会输的东西不会赌,一个不会赌的东西不会猜。
The machine exhausts; the human stakes.机器负责穷尽,人负责押注。
The machine sweeps flat the verifiable space; the human decides where — in what is not yet even defined as a space — to drive the first stake.机器把可验证的空间扫平,人决定往哪个还没有被定义为空间的地方,先插下一根桩子。
And in a world where machine labor is nearly free, a conjecture is the only interface that converts unlimited compute into meaningful labor.而在一个机器劳动近乎免费的世界里,一个猜想,是把无限算力转化成有意义劳动的唯一接口。
So upstream there is exactly one thing: the conjecture. And it is not "a sentence not yet proven."所以上游只有一个东西:猜想。而它不是"一句还没被证明的话"。
A conjecture is a statement that, at the moment it is proposed, no one can decide true or false — and whose proposer nonetheless treats it as a proposition that can be sentenced to death, and lays it out in public.猜想是一个在被提出的那一刻,任何人都无法判定其真假的陈述,而提出者仍然把它当作一个可以被判死的命题,公开摆出来。
Its value comes precisely from that undecidability. This is not a defect; it is the source of the pricing. A statement verifiable today settles all its value today; a statement decidable only fifty years from now spends those fifty years continuously organizing labor, drawing boundaries, generating methods, manufacturing adversaries. Its value is not at the settlement date. It is in the holding period.它的价值恰恰来自那个"无法判定"。这不是缺陷,这是定价的来源。一个今天就能验证的陈述,全部价值在今天就结算完了;一个五十年后才能判定的陈述,在这五十年里持续地组织劳动、划定边界、生成方法、制造对手。它的价值不在结算日,在持有期。
Put more precisely, in the language of finance:用金融的语言说得更准一点:
A conjecture is a long-dated option written on reality. Verification is exercise; conjecture is issuance.猜想是一份对现实开出的长期期权。验证是行权,猜想是发行。
An option's value rises with time to expiry and with volatility — that is, with unverifiability.期权的价值随到期时间和波动率上升——也就是说,随不可验证性上升。
A conjecture decidable at once is an exercise problem. A conjecture never decidable is an attitude. True conjecture lives in the stretch between — sharp enough that one day it must be sentenced, remote enough that today no one can sentence it.一个立刻能被判定的猜想是一道习题。一个永远无法被判定的猜想是一种态度。真正的猜想活在中间那段——足够锋利以至于总有一天会被判死,足够遥远以至于今天没有人能判。
◇
III. Six Entries in the Ledger, and a Ruler That Cannot Measure Them三、六笔账,和一把量不出的尺子
We have never seriously kept books on conjecture. Because the way we keep books presupposes one thing: value equals being right.我们从来没有认真给猜想记过账。因为记账的方式本身预设了一件事:价值等于正确。
Remove that presupposition, and recompute six entries.把这个预设拿掉,重算六笔账。
Entry one: Fermat's marginal note — three hundred fifty-seven years.第一笔:费马的一句边注,三百五十七年。
In 1637, in the margin of Diophantus's Arithmetica, Fermat wrote: for n greater than 2, the equation has no positive integer solutions. He added that he had discovered a truly marvelous proof, which this margin was too narrow to contain.1637 年,费马在丢番图《算术》的页边写下:当 n 大于 2,方程没有正整数解。他补了一句,说自己发现了一个绝妙的证明,可惜这里的空白太窄,写不下。
He almost certainly did not have that proof.他几乎肯定没有那个证明。
The sentence lived in mathematics for three hundred fifty-seven years. In the nineteenth century, Kummer, attacking it, discovered that unique factorization fails in cyclotomic integer rings, and invented "ideal numbers" — the apparatus that later grew into all of algebraic number theory. In the twentieth, Frey connected it to elliptic curves, Ribet proved the link, and in 1994 Wiles closed it out via modular forms.这句话在数学里活了三百五十七年。十九世纪,库默尔为了攻它,发现分圆整数环里的唯一分解不成立,于是造出了"理想数"——那套东西后来长成了整个代数数论。二十世纪,弗雷把它接到椭圆曲线上,里贝证明了那个联结,1994 年怀尔斯经由模形式完成收官。
And the theorem itself is perfectly useless. No person, no engineering project, no discipline needs to know whether that equation has integer solutions. The algebraic number theory that grew over three hundred fifty-seven years — that is the entire figure on the books.而这个定理本身毫无用处。没有任何人、任何工程、任何学科,需要知道那个方程有没有整数解。三百五十七年里长出来的代数数论,才是账面上的全部数字。
First: the value of a conjecture is independent of the content it asserts.第一条:猜想的价值,与它所断言的内容无关。
Entry two: Hilbert's error is the computer's birth certificate.第二笔:希尔伯特的错,是计算机的出生证。
Paris, 1900. One lecture, twenty-three problems, organizing an entire century of mathematics.1900 年巴黎,一场演讲,二十三个问题,组织了整整一个世纪的数学。
September 1930, Königsberg. Hilbert gave his retirement address at the annual meeting of the Society of German Natural Scientists and Physicians, ending with the sentence later carved on his tombstone: We must know; we will know. The speech went out over the radio.1930 年 9 月,柯尼斯堡。希尔伯特在德国自然科学家与医师协会的年会上做退休演说,结尾是那句后来刻上他墓碑的话:我们必须知道,我们必将知道。演讲由电台播了出去。
And the day before, in the same city, at a roundtable of the concurrently held epistemology conference, a twenty-four-year-old Gödel, in an offhand remark, spoke incompleteness aloud in public for the first time.而就在前一天,同一座城市、同期举行的认识论会议的一场圆桌讨论里,二十四岁的哥德尔在一段随口的发言中,第一次公开说出了不完备性。
Hilbert guessed wrong, and thoroughly. Mathematics cannot be fully formalized; the consistency of arithmetic cannot be proven from within itself. In 1936, Turing, to answer the Entscheidungsproblem — the last plank of Hilbert's program — first had to define "mechanical computation" precisely; to say what is not computable, he invented an imaginary machine.希尔伯特猜错了,而且错得彻底。数学不能被完全形式化,算术的相容性不能在其自身内部被证明。1936 年,图灵为了回答判定问题——希尔伯特纲领的最后一块——必须先给"机械地计算"下一个精确定义;为了说清什么是不可计算的,他造出了一台假想的机器。
The machine you are reading this essay on is the byproduct of a falsified conjecture.你此刻读这篇文章用的这台机器,是一次被证伪的猜想的副产品。
Second: a wrong conjecture can be worth more than a whole generation of correct work — provided it is wrong deeply enough.第二条:一个错误的猜想可以比一整代正确的工作更值钱——只要它错得足够深。
Entry three: Riemann — one hundred sixty-six years unsettled, paying interest the whole time.第三笔:黎曼,一百六十六年没结算,一直在付息。
In 1859, Riemann wrote an eight-page paper, dropped in a passing guess about the zeros of the zeta function, remarked that the point was not essential to his immediate purpose, and set it aside.1859 年,黎曼写了一篇八页的论文,顺手放进去一句关于 zeta 函数零点的猜测,并且说这一点对他手头的目的不重要,就搁下了。
One hundred sixty-six years on, no one has proven it, and no one has refuted it.一百六十六年过去,没有人证明它,也没有人推翻它。
Yet there is an entire genre of theorems in mathematics that open with "assuming the Riemann Hypothesis…" — thousands of them. That work is not waiting for redemption — it is already spending an unredeemed thing.而数学里有一整个类别的定理,写法是"在黎曼猜想成立的前提下……",数以千计。这些工作不是在等待兑现——它们已经在使用一个尚未兑现的东西了。
Third: an option never exercised has been paying interest for one hundred sixty-six years.第三条:一份从未行权的期权,已经付了一百六十六年的利息。
Entry four: the aether — a false belief that drew the correct blueprint.第四笔:以太——错误的信念,画出了正确的图纸。
Nineteenth-century physics was certain light needed a medium, and the medium was called the aether. In 1887, Michelson and Morley built an interferometer at the precision limit of the age to measure Earth's drift relative to the aether.十九世纪的物理学确信光需要介质,那个介质叫以太。1887 年,迈克尔逊和莫雷造了一台当时精度极限的干涉仪,去测地球相对以太的漂移。
The result was zero. That zero is the doorway to special relativity.结果是零。这个零,是狭义相对论的入口。
The crux: without the false conjecture of the aether, no one would have built that interferometer. A thing that does not exist specified a real, extremely expensive measurement. The experiment's design drawing grew out of a false belief.要害在于:如果没有以太这个错误的猜想,没有人会去造那台干涉仪。 一个并不存在的东西,指定了一次真实的、极其昂贵的测量。实验的设计图,是从一个错误的信念里长出来的。
Fourth: a conjecture's first function is not to assert the truth — it is to specify a measurement worth making.第四条:猜想的第一功能不是断言真相,是指定一次值得做的测量。
Entry five: a paragraph forced out by a rejection, which ended up mobilizing tens of billions of dollars.第五笔:一段被拒稿逼出来的话,最后调动了上百亿美元。
In the summer of 1964, Peter Higgs wrote a second, very short paper and sent it to Physics Letters. It was rejected — the gist of the reason: no obvious relevance to physics.1964 年夏天,彼得·希格斯写了第二篇很短的论文,寄给《物理快报》。被拒了,理由大意是:看不出与物理有什么明显的相关性。
Higgs added a paragraph on the mechanism's possible applications, noting at its end: a massive spin-zero particle would be left behind. The revision went to Physical Review Letters and was published on October 19, 1964 — two pages. He later said that the reason the boson carries his name is probably that paragraph.希格斯加写了一段,谈这个机制可能的应用,并在那一段的末尾提到:这里会留下一个有质量的自旋为零的粒子。改稿寄给《物理评论快报》,1964 年 10 月 19 日发表,两页。他后来说,他之所以被冠名到那个玻色子上,大概就是因为那一段。
And that paragraph was forced out by a rejection — a rejection whose reason was precisely a verifiability screen: "no visible relevance."而那一段,是被拒稿逼出来的;拒稿的理由,恰恰是一次"看不出相关性"的可验证性筛查。
Forty-eight years later, on July 4, 2012, CERN announced it had found the particle, on a machine twenty-seven kilometers around, costing on the order of ten billion dollars. Higgs, eighty-three, sat in the auditorium.四十八年后,2012 年 7 月 4 日,欧洲核子研究中心宣布在一台周长二十七公里、耗资以百亿美元计的机器上找到了它。八十三岁的希格斯坐在礼堂里。
Fifth: the conjecture is the strongest capital-organizing instrument humanity has ever known. A sentence unverifiable at the time it was made, nearly screened out by a journal, ultimately mobilized dozens of nations, thousands of scientists, and decades of time. No business plan has ever had that fundraising power — and we have never priced it.第五条:猜想是人类已知最强的资本组织工具。 一句在提出时无法验证、且差点被期刊筛掉的话,最终调动了几十个国家、数千名科学家和几十年的时间。没有任何一份商业计划书具备这种募资能力——而我们从来没有为它定过价。
Entry six (the reverse side): use verifiability as a gate, and you kill the true conjectures.第六笔(反面):把可验证性当门槛,会杀死真的猜想。
In 1912, Wegener proposed continental drift. The evidence was hard: the fit of coastlines across the ocean, fossils and strata matching shore to shore. But he could offer no mechanism — he could not say what force could push continents.1912 年,魏格纳提出大陆漂移。证据很硬:两岸海岸线的拼合、隔洋对应的化石与地层。但他给不出机制——他说不清是什么力量能推动大陆。
So he was ridiculed for nearly fifty years. In 1930 he died on the Greenland ice, forty-nine years old, having waited for nothing. In the sixties, seafloor spreading and paleomagnetism supplied the mechanism, and plate tectonics became the bedrock of the earth sciences.于是他被嘲笑了将近五十年。1930 年他死在格陵兰的冰上,四十九岁,什么也没等到。六十年代,海底扩张与古地磁给出了机制,板块构造成为整个地球科学的基石。
He was right. The only reason he was voted down is that his conjecture ran ahead of the era's means of verification.他是对的。他被否掉的唯一理由,是他的猜想领先于当时的验证手段。
Sixth: making verifiability the entrance requirement means systematically screening out the truths that run ahead of the means of verification.第六条:把可验证性当作准入门槛,等于系统性地筛掉那些跑在验证手段前面的真理。
◇
Set the six entries side by side: Fermat — value independent of content; Hilbert — the deeper the error, the larger the dividend; Riemann — never settled, paying interest throughout; the aether — a false belief drawing a correct blueprint; Higgs — one sentence organizing decades and tens of billions; Wegener — verifiability as gate kills the true.六笔账并排放着:费马——价值与内容无关;希尔伯特——错得越深,红利越大;黎曼——从未结算,一直付息;以太——错误的信念画出正确的图纸;希格斯——一句话组织了几十年和上百亿;魏格纳——可验证性做门槛,会杀掉真的。
Not one entry's value came from "this sentence was right."没有任何一笔的价值,来自"这句话是对的"。
Distilled, three functions — none of which requires truth:蒸馏出来是三种功能,没有一种需要它为真:
Coordinates. Before a conjecture is posed, that region is not "unsolved" — it does not exist: no one knows there is a question there to ask. A proof adds a brick to a place that already exists; a conjecture creates the place.坐标。 在猜想被提出之前,那片区域不是"没被解决",而是不存在——没有人知道那里有一个可以问的问题。证明是往一个已有的地方添一块砖;猜想是创造那个地方。
Organization. A good conjecture aligns the labor of strangers across generations toward a single direction. It needs no funding, no institution, no authority — only enough people who believe the sentence deserves to be sentenced to death. It is the only technology humanity has for organizing large-scale intellectual labor without power.组织。 一个好的猜想能把互不相识、跨越世代的人的劳动对齐到同一个方向上。它不需要资金,不需要机构,不需要权威——只需要足够多的人相信这句话值得被判死。这是人类唯一一种不靠权力就能组织大规模智力劳动的技术。
The stake. The proposer wagers himself. Erdős funded bounties for his problems out of his own pocket, from tens of dollars to thousands. That was not performance art; it was pricing: with his own money he marked the weight each problem carried in his eyes, converting judgment into liability.赌注。 提出者押上了自己。厄多斯为他的问题自掏腰包设立悬赏,从几十美元到几千美元不等。这不是行为艺术,这是定价:他用自己的钱标注了每个问题在他眼里的分量,把判断变成了负债。
◇
And our ruler measures none of the three.而我们的尺子,量不出这三样中的任何一样。
In 2026 there was an exquisitely designed study. A Stanford group had first found, in 2024, that LLM-generated research ideas were rated more novel than human experts' in blind expert review — 5.64 versus 4.84 on a ten-point scale, statistically significant. Then they completed the other half: recruit forty-three experts, randomly assign the ideas to be actually executed — a hundred-plus hours each — written up, and blind-reviewed again. The AI ideas collapsed across the board — novelty down 1.05, excitement down 1.76, effectiveness down 1.88; the human ideas barely moved: −0.01, +0.08, −0.05. The rankings inverted.2026 年有一项设计极好的研究:斯坦福的一组作者先在 2024 年发现,LLM 生成的研究点子在专家盲评里比人类专家的更新颖——十分制下 5.64 对 4.84,统计显著。然后他们补上另一半:招募四十三位专家,把点子随机分配下去真做,每人投入一百小时以上,写成短文,再交给专家盲评。结果 AI 点子全线塌方——新颖性掉 1.05,兴奋度掉 1.76,有效性掉 1.88;人类点子几乎不动,分别是 -0.01、+0.08、-0.05。排名翻转。
The experiment is usually read as "AI isn't strong enough yet." I want to point out its other meaning: it measures ideas that can be executed to completion within a hundred hours.这个实验通常被读作"AI 还不够强"。我要指出它另一层意思:它测量的,是能在一百小时内被执行完的点子。
Fermat's sentence took three hundred fifty-seven years; Riemann's has not settled to this day. Any evaluation instrument with a settlement window is structurally blind to conjecture — it does not measure conjectures as bad; it simply cannot see them. And all of our evaluation machinery — review cycles, grant cycles, graduation clocks, assessment periods — comes with windows.费马那句话用了三百五十七年,黎曼那句话到今天还没结算。任何带有结算窗口的评估工具,对猜想都是结构性失明的——不是它测出猜想不好,是它根本测不到猜想。而我们所有的评价机制——评审周期、基金周期、毕业年限、考核周期——全部带窗口。
The cost is already written in the data. Bloom, Jones, Van Reenen, and Webb conclude: research inputs have risen sharply while research productivity falls about 5% a year, halving every thirteen years; sustaining Moore's Law today takes more than eighteen times the researchers it took in the early 1970s. Park, Leahey, and Funk computed sixty years of the CD index — which measures whether a work consolidates the existing network or disrupts it — over forty-five million papers and 3.9 million patents: a pervasive, sustained decline across fields (the metric has methodological disputes; follow-up work suggests truncation bias may inflate the magnitude, but the direction has not been overturned). Chu and Evans examined 1.8 billion citations among ninety million papers across two hundred forty-one disciplines: a flood of papers brings no turnover of core ideas, only the ossification of the canon; in a field publishing a thousand papers a year, the year-to-year stability of the most-cited list is 0.25 — at a hundred thousand a year, 0.74.代价已经写在数据里。Bloom、Jones、Van Reenen 和 Webb 的结论是:研究投入大幅上升,而研究生产率大约每年下降 5%,每十三年腰斩;维持摩尔定律今天需要的研究者,是 1970 年代初的十八倍以上。Park、Leahey 和 Funk 用四千五百万篇论文和三百九十万项专利算了六十年的 CD 指数——衡量一项工作是巩固既有网络还是打断它——结论是各领域普遍且持续的下降(这个指标有方法论争议,后续研究指出截断偏差可能夸大了幅度,但方向没有被推翻)。Chu 和 Evans 检视两百四十一个学科、九千万篇论文之间的十八亿次引用:论文洪水不会带来核心思想的更替,只会带来经典的骨化;一个每年发一千篇论文的领域,最高被引名单的年际稳定度是 0.25,每年十万篇时是 0.74。
The mainstream reading is "ideas are getting harder to find."主流的读法是"想法越来越难找了"。
My reading: the CD index is the conjecture index. It measures precisely whether a work spared its successors the need to cite its predecessors — whether it broke open new ground. Its sixty-year slide measures not that ideas got harder to find, but that we stopped conjecturing.我的读法是:CD 指数就是猜想指数。 它测的正是一项工作有没有让后人不必回头引它的前辈——有没有开出一块新地。它六十年的下滑,测的不是想法变难找了,是我们停止了猜。
Before AI arrived, the bottleneck was already not output. And the first thing AI does is multiply output by a hundred. That is not acceleration; that is doubling the only known pathogen.在 AI 到来之前,瓶颈就已经不是产量。而 AI 做的第一件事,是把产量乘以一百。那不是加速,那是把唯一已知的病因加倍。
◇
IV. Only Conjecture Reaches the Great Questions of the Age四、只有猜,才能碰到时代的大命题
The first three chapters said why conjectures are valuable. This chapter says something harder: there is a class of problems that has no entrance except conjecture.前三章讲的是猜想为什么值钱。这一章讲一件更硬的事:有一类问题,除了猜,没有别的入口。
Kuhn divided science into two states. Normal science solves puzzles within a given paradigm — the problems are clear, what counts as a solution is clear, the work is filling in the puzzles one by one. But the paradigm itself is never a product of normal science; it comes from one person's re-declaration of the shape of the world.库恩把科学分成两种状态。常规科学在既定范式内解谜——问题是清楚的,什么算解答是清楚的,工作是把谜一个一个填上。而范式本身从来不是常规科学的产物,它来自某个人对世界形状的一次重新宣称。
AI is the strongest normal-science machine humanity has ever built. Its puzzle-solving power within a given frame will soon exceed the sum of all humans. And normal science never produces paradigms.AI 是人类造出的最强的常规科学机器。它在给定框架内解谜的能力,很快会超过所有人类的总和。而常规科学从不产生范式。
Worse follows. The birth of a paradigm requires anomaly to be treated as signal, not noise — someone must recognize, in unclean data, that a certain aberration is not error but a door. Yet a system trained on the distribution of existing literature will naturally file anomalies under known categories — because in the literature, the label on anomaly is "error."更糟的还在后面。范式的诞生需要异常被当成信号,而不是噪音——需要有人在不干净的数据里认出某个反常不是误差,而是一扇门。可是一个在既有文献分布上训练出来的系统,天然会把异常归进已知类别:因为异常在文献里的标签,就叫误差。
The machine not only produces no paradigms — in passing, it tidies away the raw material a paradigm's birth requires. That sixty-year curve of canon ossification is the early reading.机器不仅不产生范式,它还会把范式诞生所需的原料,顺手清理干净。 经典骨化那条六十年的曲线,就是早期读数。
◇
Lay out the truly great questions of this age: What is consciousness? Is aging reversible? How does life arise from chemistry? What are dark matter and dark energy? How does gravity reconcile with quantum mechanics? What is the principle of intelligence itself? What keeps a ten-billion-scale hybrid carbon-silicon system from collapsing?把这个时代真正的大问题摆出来看:意识是什么。衰老是不是可逆。生命如何从化学里起来。暗物质和暗能量是什么。引力怎样和量子力学相容。智能本身的原理是什么。一个碳基与硅基混合的百亿级体系靠什么不崩溃。
What they share is not difficulty. It is that not one of them has a benchmark today.它们的共同点不是难。是它们今天一个都没有 benchmark。
No scoreable objective function, no gradient; no gradient, and an army of machines that advances by gradient structurally cannot get there. This is not a capability problem but a direction problem: machines can only enter spaces that humans have already named.没有可打分的目标函数就没有梯度,没有梯度,一支按梯度前进的机器军团在结构上到不了那里。这不是能力问题,是方向问题:机器只能进入已经被人命名的空间。
Erdős's problem fell first precisely because it is the extreme opposite of this — objective clear, boundary clean, verification cheap. And the remaining half of the list is likewise all named.厄多斯的问题之所以第一个倒下,正因为它是这件事的极端反面——目标清晰,边界干净,验证便宜。而清单上剩下的那一半,也全都是被命名过的。
The entrance to the great questions is not compute. It is naming. And naming is conjecture.大命题的入口不是算力,是命名。而命名,就是猜。
◇
A great question cannot itself be attacked. "What is consciousness" admits no experiment. It must first be cut, by some conjecture, into a shape that can be sentenced to death — into a measurement that can be performed.大命题本身无法被攻击。"意识是什么"没法做实验。它必须先被某个猜想切成一个可以判死的形状——切成一次可以做的测量。
The aether serves a second time here. The nineteenth century's great question was "what, in the end, is light" — unattackable directly; it was the aether, a conjecture later falsified, that translated it into a measurable quantity — Earth's drift relative to the medium. Hence the interferometer, hence that zero, hence relativity.以太在这里第二次发挥作用。十九世纪的大命题是"光到底是什么",无法直接攻;是以太这个后来被证伪的猜想,把它翻译成了一个可测的量——地球相对介质的漂移。于是有了干涉仪,有了那个零,有了相对论。
In 1944 Schrödinger wrote a very thin little book, What Is Life?, and in it guessed that genetic information is stored in an "aperiodic crystal." Wholly unverifiable at the time. And Watson, Crick, and Wilkins all later acknowledged its influence — the little book translated the unattackable great question "how does heredity work" into a structural problem. Structural problems can be photographed. Nine years later, the double helix.1944 年薛定谔写了一本很薄的小册子《生命是什么》,在里面猜测遗传信息存储在一种"非周期性晶体"里。这在当时完全无法验证。而沃森、克里克、威尔金斯后来都承认受过它的影响——这本小册子把"遗传如何工作"这个不可攻的大命题,翻译成了一个结构问题。结构问题可以照相。九年后是双螺旋。
A conjecture does not answer the great question. It translates the great question into something that can be done.猜想不回答大命题。它把大命题翻译成一件可以被做的事。
Without that act of translation, any amount of compute just spins on old questions already translated.没有这一步翻译,再多的算力也只是在已经翻译好的旧题上打转。
◇
A load-bearing conjecture meets three hard standards.能承重的猜想有三条硬标准。
Killable — it must be capable of being sentenced to death. A statement no imaginable evidence could kill is not a conjecture; it is a posture. A century-old misreading must be dismantled here: what Popper demanded was "capable of being killed," not "verifiable today." Wegener's continental drift was perfectly killable — the means of the day simply couldn't kill it yet. Swapping "killable" for "currently verifiable" is the deepest mistranslation in the modern institution of research.可错——它必须能被判死。没有任何可以想象的证据能杀死它的,不是猜想,是姿态。这里有一处百年误读要拆开:波普尔要求的是"可被杀死",不是"今天能被验证"。魏格纳的大陆漂移完全可被杀死,只是当时的手段还杀不动它。把"可被杀死"偷换成"当下可验证",是现代科研制度最深的一处错译。
Arable — when it turns out wrong, the tilled land remains. This is the only standard separating good conjecture from wild guessing. A conjecture's value should be appraised by what still stands after it is falsified.可耕——它错了之后,耕出来的地还在。这是区分好猜想和瞎猜的唯一标准。一个猜想的价值,应该按它被证伪之后仍然留下的东西来估。
Bettable — the proposer must stake name, money, years, a public position. An unstaked conjecture is noise, because it costs nothing, and what costs nothing can carry no information.可赌——提出者要押上名字、钱、年限、公开的立场。没有押注的猜想是噪音,因为它没有成本,而没有成本的东西无法承载信息。
Conversely: a "new direction" a model interpolates in a high-density region is not a conjecture; something an experiment could settle today is not; something never settleable is not; something no one has to answer for, is not.反过来说,模型在高密度区插值出的"新方向"不是猜想;今天就能跑个实验判定的不是;永远无法判定的不是;不需要任何人为它负责的,不是。
◇
Institutionally, then, four things can be done:制度上因此有四件事可做:
Issuance. Erdős's bounties were already a complete financial structure — the proposer marks the difficulty in cash, the solver takes the prize, and all the years in between are farmed for free by the community. erdosproblems.com proves it fully workable as engineering: a numbered list, a status column, a forum — enough to organize the world's attention. Today we can do better: open registries, tradable bounties, reputation records weighted by time-to-falsification, and retroactive prizes for conjectures that were overturned but tilled new ground — rewarding only those who guessed right means rewarding only the short-term, the safe, the immediately verifiable.发行。 厄多斯的悬赏就是一个完整的金融结构——提出者用现金标注难度,解决者拿赏金,中间的所有年份由共同体免费耕作。erdosproblems.com 证明了它在工程上完全可行:一个编号清单、一个状态栏、一个论坛,就够组织全世界的注意力。今天可以做得更好:公开登记、可交易赏金、按证伪时间加权的信誉记录,以及为"被推翻但耕出了地"的猜想设立回溯性奖励——只奖励猜对的人,等于只奖励短期、安全、马上能验证的猜。
Output. What a researcher owes in a year is not three settleable papers, but at least one killable, arable, bettable conjecture — together with what he is willing to stake on it.产出。 一个研究者一年该交出的,不是三篇可结算的论文,而是至少一个可错、可耕、可赌的猜想,附上他愿意为此押上什么。
The ledger. An institution's true balance sheet is not how much it has published, but which unexpired conjectures it currently holds, how much is staked on each, and how long to expiry. One glance at this sheet tells you whether it is tilling land or collecting rent.账本。 机构真正的资产负债表不是发表了多少,而是现在持有哪些未到期的猜想、各押了多少、还有多久到期。这张表能一眼看出它是在耕地还是在收租。
Bearings. Automation advances along the contour lines of verifiability — so the fields at the bottom of the prestige ladder for being "not rigorous enough" — wet labs, clinical work, ecology, fieldwork — are in fact the last and most valuable ground. Not because they are deeper, but because the world refuses to answer them cheaply.方位。 自动化沿可验证性的等高线推进,所以那些因"不够严格"而位于鄙视链下游的领域——湿实验、临床、生态、田野——反而是最后、也最值钱的阵地。不是因为它们更深刻,是因为世界拒绝便宜地回答它们。
Three falsifiable predictions: before 2029, the first mainstream mechanism appears that pays for conjectures rather than results, with retroactive prizes for "falsified but ground-tilling" conjectures; before 2030, the first institutional annual report appears whose core disclosure is its "portfolio of unexpired conjectures"; before 2031, at least one widely acknowledged major advance traces its origin to a statement proposed by a human, unverifiable at the time, ultimately proven false. If none of the three occurs, this essay's judgment is wrong.三条可证伪的预言:2029 年前出现第一个专门为猜想而非成果付钱、且对"被证伪但耕出了地"设有回溯性奖励的主流机制;2030 年前出现第一份以"未到期猜想组合"为核心披露项的机构年报;2031 年前至少一项被广泛承认的重要进展,其源头可追溯到一个由人类提出、当时无法验证、最终被证明为假的陈述。三条都不发生,本文的判断就是错的。
◇
V. And Conjecture Can Become Beautiful五、而猜,可以变得很美
The last chapter was about conjecture's uses. This one is about its other face — the one I believe is the true one.上一章讲的是猜想的用处。这一章讲它的另一面,也是我认为真正的那一面。
Because if conjecture were merely a high-yield asset, then the moment a market is built for it, it gets optimized, gamed, arbitraged — one more citation count. It has not been arbitraged away, because its root is not in institutions.因为如果猜想只是一种高收益资产,那么一旦为它建立市场,它就会被优化、被刷分、被套利,变成另一个论文数。它没有被套利掉,是因为它的根不在制度里。
Erdős was an atheist who called God the "Supreme Fascist." But he believed there was a book — THE BOOK — in which the most beautiful proof of every theorem is written, and the mathematician's work is to glimpse a page of it now and then.厄多斯是无神论者,管上帝叫"至高法西斯"。但他相信有一本书——天书——里面记着每一个定理最漂亮的那个证明,数学家的工作是偶尔窥见其中一页。
That half-joke is mathematics' aesthetic axiom: there exists a standard higher than "correct." A proof can be correct, and ugly.这句半玩笑的话是数学的美学公理:存在一个比"正确"更高的标准。 一个证明可以是正确的,而且是丑的。
Which gives the 2026 event its sharpest reading. The proof that overturned the unit-distance conjecture is generally described as won by stubborn search rather than insight. It is correct. It is probably not in THE BOOK.于是 2026 年那件事有了它最尖锐的读法。推翻单位距离猜想的那个证明,被普遍描述为靠顽固的搜索而非洞见得来的。它是对的。它大概不在那本书里。
When "correct" has been automated, the residue that remains is beauty.当"正确"被自动化之后,剩下的残差就是美。
◇
This is not a littérateur's sigh; it is a working method at the front line. Dirac, in 1963, was blunt: it is more important to have beauty in one's equations than to have them fit experiment — if they disagree, some secondary factor has probably not been handled. Weyl said he had always tried to unite the true with the beautiful, and when forced to choose, he usually chose the beautiful. Hardy said ugly mathematics has no permanent place in the world.这不是文人的感慨,是一线的工作方法。狄拉克在 1963 年说得毫不含糊:让方程漂亮比让它符合实验更重要,如果和实验不完全一致,多半是些次要因素还没处理好。外尔说他一直试图把真与美统一起来,当必须二选一时,他通常选美。哈代说丑陋的数学在世上没有永久的位置。
They were using beauty to choose direction.他们是在用美来选方向。
Why is beauty method rather than ornament? Because in a space of nearly infinite coherent statements, you cannot enumerate. Beauty is the human being's extremely low-cost judgment of high-dimensional structure — a compression: this structure is "right," before you can prove it. And it is the only search heuristic that requires no reward signal.为什么美是方法而不是装饰?因为在一个连贯陈述近乎无限的空间里,你不可能穷举。美是人类对高维结构的一种极低成本的判断——一种压缩:这个结构"对",在你能证明它之前。而且它是唯一一种不需要奖励信号的搜索启发式。
Which is exactly its position today. In recent years, everything that can be written as a reward function has been optimized to the limit; every direction that can be scored has been claimed first.这正是它今天的位置。过去几年,凡是能写成奖励函数的都被优化到了极限,凡是能被打分的方向都被抢先占满。
Beauty is the only prior that cannot be written into a reward function — and therefore the only prior not yet arbitraged away.美是唯一一个写不进奖励函数的先验,因此也是唯一一个还没有被抢先套利的先验。
◇
Two conflated words must be pulled apart here.这里要把两个被混用的词分开。
What models learn is the average taste of the human corpus, so they excel at the good-looking — fluent, balanced, familiar, expectation-conforming. Good-looking is the center of the distribution.模型学的是人类语料的平均品味,所以它极其擅长好看——流畅、匀称、熟悉、符合期待。好看是分布的中心。
But beauty lives at the distribution's fracture lines. Frontier beauty is almost always first taken for ugliness: Higgs's paragraph was judged "no visible relevance"; Wegener was ridiculed for fifty years; non-Euclidean geometry was long treated as pathological.而美在分布的断裂处。前沿的美几乎总是先被当成丑:希格斯那一段被判为"看不出相关性",魏格纳被嘲笑了五十年,非欧几何很长时间里被当成病态。
A system that samples on density can make infinitely many good-looking things, and never once strike the fracture.一个在密度上采样的系统,能造出无穷多好看的东西,永远撞不上断裂。
◇
Its danger must also be stated plainly, or this chapter becomes superstition.它的危险也必须说清楚,否则这一章就变成迷信。
In 1596 Kepler explained the spacing of the planetary orbits by nesting the five Platonic solids. It was a construction of pure aesthetic and theological conviction, and it was wrong through and through. But it was precisely the false belief that "the cosmos must be harmonious" that gave him the patience to grind Tycho's Mars data for years — grind until he was forced to abandon the circle — and got the ellipse.1596 年开普勒用五个柏拉图立体嵌套解释行星轨道的间距。那是一个纯出于美学与神学信念的构造,而且彻头彻尾是错的。但正是"宇宙必须是和谐的"这个错误信念,让他有耐心把第谷的火星数据磨了很多年,一直磨到被迫放弃圆——然后拿到了椭圆。
And the same thing's other face: the aesthetic conviction "the circle is the perfect shape, therefore heavenly bodies must move in circles" trapped humanity for fifteen hundred years. In recent years, physicists have also systematically criticized this — arguing that the obsession with "natural law must be elegant" led fundamental physics into decades of idling.而同一件事的另一面是:"圆是最完美的形状,所以天体必须走圆"这个美学信念,把人类困了一千五百年。 近年也有物理学家系统地批评过这件事,认为对"自然律必须优美"的执念,把基础物理带进了几十年的空转。
So the conclusion must be drawn precisely:所以结论必须下得精确:
Beauty does not guarantee truth. Beauty guarantees "worth a try." It is a heuristic for the search, not a criterion of truth.美不保证真。美保证的是"值得一试"。它是搜索的启发式,不是真理的判据。
Confuse the two, and you fall from Kepler's ellipse into a mathematics of self-admiration; abandon it entirely, and you hand the choice of direction to the only thing that can still keep score — that is, to the region the machines already occupy.混淆两者,就是从开普勒的椭圆掉进一片自我欣赏的数学;完全放弃它,就是把方向的选择权交给唯一还能打分的那个东西——也就是交给机器已经站满的那片区域。
◇
One last thing: institutions cannot produce conjectures.最后一件事:制度生产不了猜想。
Erdős had no house, no position; he slept on other people's sofas; his luggage was one suitcase. He conjectured not because anyone paid for conjectures — then as now, nobody pays for conjectures. The four measures of the last chapter are all, without exception, concrete forms of "just don't kill it."厄多斯没有房子,没有职位,睡别人的沙发,行李是一个手提箱。他猜,不是因为有人为猜想付钱——那时候和现在一样,没有人为猜想付钱。上一章那四件事,全部只是"别把它杀掉"的具体形式。
Then why does a person conjecture at all?那么一个人为什么要猜?
Because conjecture is the finite being's only way of reaching toward the whole.因为猜是有限者对整体唯一的伸手方式。
A mortal being, granted sight of only the smallest patch in a lifetime, declares that this world he will never finish seeing has a certain shape — it is the least reasonable thing a human can do, and the most human. It does not grow out of reasonableness; it grows out of a person's believing the world deserves to be understood, and being willing to lose a lifetime on that belief.一个会死的、一辈子只能看见极小一块的存在,宣称这个他永远看不完的世界具有某种形状——这是人能做的最不合理的一件事,也是最像人的一件事。它不是从合理里长出来的,是从一个人相信世界值得被理解、并且愿意为这个相信赔上一生里长出来的。
And what science leaves behind, in the end, was never a pile of true propositions.而科学最后留下的,也从来不是一堆真命题。
It is a world become more intelligible.是一个变得更可理解的世界。
Truth is only one path to intelligibility, not the only one. Fermat was wrong, Hilbert was wrong, the aether was wrong, Kepler's solids were wrong — and because of them the world became more intelligible, and better looking.真只是通往可理解性的一条路径,不是唯一一条。费马错了,希尔伯特错了,以太错了,开普勒的立体错了——而世界因为它们变得更可理解,也更好看。
This is the exact meaning of "beyond verifiability": verifiability measures propositions; and the product of science is not propositions.这就是"超越可验证性"的确切含义:可验证性衡量的是命题;而科学的产物,不是命题。
◇
Erdős posed thousands of problems in one lifetime, crossed the world, owned no house, held no post, lived off other people's sofas and his own pockets. He funded bounties for his problems out of his own money — he could not solve them; he could only price them. He believed in a book of heaven that records the most beautiful proof of every theorem.厄多斯一生提了几千个问题,走遍世界,没有房子,没有职位,用别人的沙发和自己的口袋。他自掏腰包给他的问题悬赏——他不能解决它们,他只能标价。他相信有一本天书,里面写着每个定理最漂亮的那个证明。
Eighty years later, a machine solved one of them, and proved he had guessed wrong. It is correct. It got there by stubbornness, not insight. It is probably not in THE BOOK.八十年后,一台机器解掉了其中一道,并且证明他猜错了。它是对的。它靠的是顽固,不是洞见。它大概不在那本书里。
And what was advancing all through those eighty years was never his answers. He had no answers.而八十年里一直在被推进的,从来不是他的答案。他没有答案。
It was his conjecturing.是他的猜。
In a world where answers are free, answers are the denominator.在一个答案免费的世界里,答案是分母。
Every problem solvable; no one to conjecture.万题可解,无人去猜。
That is all.这就是全部。
◇
Further Reading延伸阅读
- erdosproblems.com itself, and the "AI contributions to Erdős problems" wiki on teorth/erdosproblems — a list being emptied in real time; watch the rate of change in the status column.erdosproblems.com 本身,以及 teorth/erdosproblems 上的"AI 对厄多斯问题的贡献"维基——一份正在被实时清空的清单;建议直接看状态栏的变化速率。
- "Resolution of Erdős Problem #728" (arXiv:2601.07421) and LeanMarathon (arXiv:2606.05400) — the first widely acknowledged autonomous solution; and the source of that 6.5% figure, and of why formalization became the only insurance.《Resolution of Erdős Problem #728》(arXiv:2601.07421)与 LeanMarathon(arXiv:2606.05400)——前者是第一例被广泛承认的自主解,后者给出那个 6.5% 的数字,以及为什么形式化成了唯一的保险。
- Kuhn, The Structure of Scientific Revolutions (1962) — the division of labor between normal science and paradigm; read it to understand why an extremely strong puzzle-solving machine may be exactly what postpones the revolution.库恩《科学革命的结构》(1962)——常规科学与范式的分工;读它是为了理解为什么一台极强的解谜机器,恰恰可能让革命被推迟。
- Popper, Conjectures and Refutations (1963) — note his preference for "high information content, low prior probability," and the distinction between "killable" and "currently verifiable."波普尔《猜想与反驳》(1963)——注意他对"高信息含量、低先验概率"的偏好,以及"可被杀死"与"当下可验证"的区别。
- Schrödinger, What Is Life? (1944) — how a slim booklet translated an unattackable great question into a structural problem that could be photographed.薛定谔《生命是什么》(1944)——一本小册子如何把一个不可攻的大命题,翻译成一个可以照相的结构问题。
- Dirac, "The Evolution of the Physicist's Picture of Nature" (Scientific American, 1963) — the original source of "beauty in one's equations matters more than fitting experiment."狄拉克《物理学家的自然图景之演变》(Scientific American, 1963)——"让方程漂亮比让它符合实验更重要"的原始出处。
- Aigner & Ziegler, Proofs from THE BOOK — the mortal edition of Erdős's book of heaven.Aigner & Ziegler,《Proofs from THE BOOK》——厄多斯那本天书的人间版本。
- Hossenfelder, Lost in Math (2018) — the counter-testimony to Chapter V; must be read alongside it, or the heuristic gets mistaken for a criterion.Hossenfelder,《Lost in Math》(2018)——第五章的反面证词,必须一并读,否则容易把启发式当成判据。
- Si, Yang, Hashimoto, "Can LLMs Generate Novel Research Ideas?" (arXiv:2409.04109) and Si, Hashimoto, Yang, "The Ideation-Execution Gap" (arXiv:2506.20803) — must be read as a pair; together they expose the blind spot of the evaluation window itself.Si, Yang, Hashimoto,《Can LLMs Generate Novel Research Ideas?》(arXiv:2409.04109)与 Si, Hashimoto, Yang,《The Ideation-Execution Gap》(arXiv:2506.20803)——必须成对读;两篇合起来暴露的是评估窗口本身的盲区。
- Artiles et al., "The Alien Space of Science" (arXiv:2603.01092) — redefines "novel" from "absent from the literature" to "unlikely to occur to the community of researchers."Artiles 等,《The Alien Space of Science》(arXiv:2603.01092)——把"新颖"从"文献里没有"重新定义为"研究者群体不会想到"。
- Bloom et al., "Are Ideas Getting Harder to Find?" (AER, 2020); Park et al. (Nature, 2023); Chu & Evans (PNAS, 2021) — read the three curves together: they are the invoice for "we stopped conjecturing."Bloom 等《Are Ideas Getting Harder to Find?》(AER, 2020)、Park 等(Nature, 2023)、Chu & Evans(PNAS, 2021)——三条曲线合起来读,是"我们停止了猜"的账单。