Rendered at 20:15:53 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
gbjcantab 1 days ago [-]
For some reason, this approach makes me think of the difference between “wizardry” and “sorcery” in some fantasy magic systems. The magic of “wizards” is fundamentally based on a deep study and understanding of arcane things, perhaps assisted by some (necessary or helpful) tools of great power. “Sorcerers” summon supernatural beings and are able to control them, cajole them, and protect themselves and others against them (with more or less success)... but the actual desired magical effect is performed by those beings.
Computing has historically been a field of wizardry. It's... interesting (?) to see so many people pushing so hard in the direction of sorcery, and in fact applying that sorcery to other fields, in which they themselves aren't quite able to validate whether the spell worked or not.
necklesspen 24 hours ago [-]
It's a truly wonderful read, made better by the fact the author doesn't quite understand what's going on.
The usual format that fun mathematics is presented (being talked at by someone who is very well versed in the subject) comes with a heavy cognitive burden - and often I just can't really make it through.
When the author is not an expert the writing is just so much more accessible - it's easier to understand and making it through feels more of an adventure and less of a lecture.
I've never thought previously how much I would enjoy this format though. I'm here to see more amateurs stumbling through mathematics.
Also, wasn't expecting this sort of side-quest from the guy who got me into React.
SturgeonsLaw 20 hours ago [-]
Agreed! I've always been interested in the field of mathematics, only to be let down by my lack of ability to comprehend it. I can usually follow an article up until the point that it starts using formulae.
I love this approachable prose.
If anyone is aware of any other "mathematics for people who don't know mathematics" resources I'd greatly appreciate any links.
danabramov 19 hours ago [-]
OK this is probably not quite what you meant but I'll shoot: I highly recommend Terence Tao's Analysis textbook (which now also has a Lean counterpart on GitHub). You can skip Chapter 1 (it sets up the motivation but I couldn't answer most questions, which is part of the point). From Chapter 2, it builds up in a somewhat "dry" but actually very accessible and methodical way. It is "mathematical" but it is "for people who don't know mathematics" in the sense that it forces you to build the entire mathematics from scratch, including proving things like `a + b = b + a` as exercises. This is actually how I got into proofs in the first place, with later picking up Lean by doing Natural Number Game.
GPerson 15 hours ago [-]
Since you respect Terence Tao, but not the general population of mathematicians, maybe you’ll read his blog posts which describe the ways your actions are detrimental to human striving.
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
SturgeonsLaw 8 hours ago [-]
Thanks, I'll have a read
Certhas 14 hours ago [-]
Are you talking about the posted article? I agree that it was a fun read, but it did not present any mathematics at all. Like, none. Its not that the author not being an expert made it easy to understand, it's that there was nothing mathematical to understand presented.
patcon 1 days ago [-]
Heh, I like this. but it should be pointed out that from the other point of view, software developers were the supernatural beings (dare I say demons), which the sorcery of a good project manager could tame (with more or less success) to perform the desired magical effect
KyleTheDev 1 days ago [-]
All this time, I've been considering myself the warlock. When, in fact, I've simply been the Imp. Dastardly news.
dodslaser 9 hours ago [-]
It's goblins all the way down.
gbjcantab 1 days ago [-]
The same corollary occurred to me, as well!
dormento 1 days ago [-]
And now everyone is a sorcerer: they can trap small demons inside metal boxes and force them to do their bidding. As before, any supernatural effects are performed by those beings (which were willed into existence by siphoning the wisdom from the wizard's own grimoires...)
baq 1 days ago [-]
I’ve felt like a warlock for about half a year now - talking to demons which summon code from the abyss. Exhilarating and terrifying, especially when you can tell the demons get better faster.
bwfan123 23 hours ago [-]
> in which they themselves aren't quite able to validate whether the spell worked or not
Knowledge is of 2 kinds: know-that and know-how. Know-that is what LLMs are enabling such as the proof here, while know-how is more useful as that constitutes understanding and puts that knowledge to use.
adamddev1 22 hours ago [-]
Or like the difference between chemistry and alchemy? Understanding and reasoning with the building blocks as opposed to throwing random stuff together, trying different things and hoping it somehow produces gold.
Xirdus 1 days ago [-]
The sorcerers have always outnumbered the wizards. Before AI, we called them code monkeys.
gbjcantab 1 days ago [-]
That’s fair! I think what struck me, though, is that what we’re seeing is large numbers of “wizards” jump ship to and actively promote “sorcery” instead, in this sense. That is, it’s interesting to me to see Dan Abramov, whose blog primarily consists of painstaking explanations of React internals based on the deep knowledge he developed over years as the most visible member of the core team, switch over to “do a breakthrough.”
danabramov 1 days ago [-]
I try to address this in the post explicitly in a few places.
Primarily I thought of this as a sort of "epistemic performance art project", maybe similar to playing Elden Ring blindfolded having never played it before, or speedrunning a game by opening a box a thousand times and overflowing some counter. It's funny and absurd to do knowledge work without the knowledge.
I think it's also a stress test of meta skills. Like, how much can we do without knowing? What kind of processes can we set up around these demons that would constrain them into our requirements? How can we know when things are going wrong? In some sense, this isn't too different from engineering management.
Naturally, I'm also interested in how much of my role in this could've been automated away. Can there be a skill for that? Then "do a breakthrough" is an irrelevant implementation detail of that skill.
Note that "do a breakthrough" actually produced the worst results over the runs. The best results were from more directed runs like searching for first obstacle towards the next milestone.
iamflimflam1 1 days ago [-]
This reminds me of Terry Pratchett’s Sourcery.
The wizards don’t become sourcerers themselves - they become enthusiastic users of someone else’s sourcery. Their years of learning don’t protect them from mistaking access to power for mastery of it.
thesuitonym 1 days ago [-]
And before those code monkeys were making money, we called them skiddies.
Xirdus 4 hours ago [-]
Script kiddies are hacking sorcerers. Different discipline.
gchamonlive 1 days ago [-]
Differently than wizards that lock their knowledge in towers and in sects, software development has a tradition of being open for the most part, so the sorcery and wizardry analogy works more like a spectrum. It just depends on how close to the metal the apprentice would like their consciousness.
saghm 22 hours ago [-]
I thought you were going to talk about the difference between D&D wizards and sorcerers: the former are the same as you describe, whereas the latter are born with some sort of innate talent for casting spells without needing to learn it or necessarily practice any discipline. It honestly also kind of fits; there's a difference between "I worked hard to learn the skills I have so I can explicitly craft the solution to the problem I have" and "idk man I just kinda point and think 'go magic spell' and the magic happens".
Sorcerers are the vibe coders of D&D magic, which perfectly describes how wizards in-universe feel about them.
bwfan123 1 days ago [-]
> In either case I believe people who can put AI to the most value are the mathematicians themselves
The net output of math will increase, and mathematicians have more work now to unravel all this, and make it useful. AI plays the role of a monkey in the infinite monkey theorem [1]. We now need an LLM corollary - Something like: A finite number of LLM agents will almost surely find all theorems given an infinite token budget.
It's impossible for finite number of LLMs to solve all theorems. This would imply that the busy beaver sequence is computable which implies the halting problem is decidable.
For any finite program (eg some LLMs), there is a true math theorem which they cannot prove or disprove (given fixed input of the statement with no other information sources). If that weren’t true, BB would be computable.
Math is beyond computation. Since AI is just bits in bits out, it has this fundamental limitation.
Any magic of AI systems comes from the transformed meaning of its input data. With fixed weights any LLM is just an artifact. For example a human prompting an LLM constitutes an extra information source, which removes the above limitations. In theory any input from the natural world would remove the limitations too. The natural world is a black box and we don't know what kind of meaning or intelligence could underly it.
gf000 1 days ago [-]
> Math is beyond computation.
We are talking about the same thing, but I would actually put this the other way around.
Computation and computability is "the final frontier". Math is a "subset" of that. Doesn't matter if we choose ZFC or in the future discover some "better" subset of core axioms, we will always hit limits where BB will trivially skip over whatever we could prove (let alone Gödel's theorems).
> given fixed input of the statement with no other information sources
Also, this is just trivially avoidable, so not sure if we really should be concerned about this limitation. An LLM in a loop where it can write on a tape can be Turing complete, ergo it can compute anything computable and is "bigger" than math at that point.
magicalist 1 days ago [-]
> Computation and computability is "the final frontier". Math is a "subset" of that.
Maybe I'm misunderstanding you point, but I don't know how widely this would be held as true. Are you defining "math" as _only_ what can be proven under some particular formal system?
gf000 1 days ago [-]
Well, I only know how to define computability in terms of Turing machines.
For math I don't have a fix definition, but it's surely a bit more specific than that (e.g. I wouldn't consider the computation that prints a 0 at the same place for infinity math) - but of course I do see the circularity in my argument: a Turing machine is a mathematical object in and of itself. Though being able to talk about something doesn't necessarily change which is "bigger".
As for the other direction, this gets a bit more into the philosophy behind math itself. Constructive math's territory is "easy" - but I am on the opinion that if humans (or any intelligent physical entity) are at most Turing-complete [1], then any non-constructive math "steps" or thoughts must also be at most computable. Well, unfortunately I can't prove whether math done by transcendent entities are also computable, though.
In any case, I am no mathematician, so whatever I think regarding this topic may not have much relevance to anyone, only done CS course with quite a bit of math, but that's obviously not the same.
[1] I believe religion is an escape hatch here from an argument perspective
streetfighter64 1 days ago [-]
> if humans (or any intelligent physical entity) are at most Turing-complete
This is a bit of a strange assumption to make. I do agree that a human, if it had infinite memory, would be an universal machine, i.e. capable of computing any given Turing machine [0]. But would that be the limits of its capabilities? It's far from certain.
You'll get into the philosophy of free will (funnily enough, a sort of inverted Turing test), i.e. for a given human with infinite memory, is there a Turing machine that exactly replicates the behavior of that human? Is our behavior governed entirely by rules? Would that imply that a human themselves is a kind of Chinese room [1]?
> any non-constructive math "steps" or thoughts must also be at most computable.
What does it mean for a "thought" to be computable? Compare to Gödel's incompleteness theorem. Clearly the act of stating the thought, or writing down the theorem, is computable. But proving it to be true or false may very well be impossible.
> What does it mean for a "thought" to be computable?
Well, given our scientific knowledge it's a molecule-level (only important to disregard quantum physics to make the case easier) physical/chemical process, that we should in principle be able to simulate on any other medium, including a Turing machine.
Nonetheless, I can accept the definition of math where it's about "truths" and truths can obviously exist without being computable.
qarl 1 days ago [-]
> But would that be the limits of its capabilities? It's far from certain.
Do you agree that humans are physical systems?
My understanding is that any physical system can be evaluated to any degree of accuracy by a computer, no?
rsrsrs86 16 hours ago [-]
No
qarl 2 hours ago [-]
HEH. Well... the only other option is supernatural. Is that what you mean?
streetfighter64 1 days ago [-]
> any physical system can be evaluated to any degree of accuracy by a computer
That's an interesting hypothesis, but I don't know why you'd assume it to be true at face value. It's a bit unclear how you would even define "evaluated", given that we don't yet have a mathematical model of all of physics as we know it. [0] And then consider unknown unknowns.
> Do you agree that humans are physical systems?
Do you consider humans _with infinite memory_ as physical systems? Do you consider computers _with infinite memory_ as physical systems?
I agree that physics being simulated is not that easy to handwave away. That's why I mention that brains probably don't "depend" on some quantum-level behavior and a more macro view of physics could be enough. What I mean here is that while there are obviously quantum effects in play at the atomic/molecular levels, if we take the cells as a black box and replace them with statistical processes, we would probably still get a human intelligence as a result - but of course I can't prove it. As a hunch, the 100 billion neurons and their 100 trillion connections, and of course their environment (glial cells are important)'s proper Simulation is enough.
As for the infinite memory, Turing machines have this nice property that they can only visit a finite amount of memory after finite steps, no matter what. A Turing machine running for a finite time (we got this) will surely use a finite space, so being "a bit short" on infinite space is not a problem, I believe.
qarl 1 days ago [-]
> given that we don't yet have a mathematical model of all of physics as we know it
Yes... but that's in the area of the big bang and black holes. My understanding is that the chemistry of the brain is very well modeled.
So, unless we find unknown physics, and unless that physics behaves differently than every other known physics, humans are computable?
Do I have that right?
streetfighter64 24 hours ago [-]
Let me illustrate with an example. Are you familiar with with the Collatz conjecture? It's an example of a system with only one variable, and two simple rules. Are you certain that there exists a computer program that in finite time can compute where any given integer ends up?
Now consider throwing a ball in the air. Can you even write down the rules that each of the ball's subatomic particles obeys? How can you be certain there exists a computer program that in finite time can predict where any of the particles, for any ball, ends up?
> the chemistry of the brain is very well modeled
There are models, but the fact of those models is that they do not apply to "any degree of accuracy", as you claim.
Consider the ball thrown in the air again. Is the ball affected by what happened 100 years ago, inside of a black hole 100 light years away? Why would it not be affected by that? Or if you grant that it is affected by that, do we then need a model to predict those effects before we can "evaluate" them?
EDIT regarding the below linked blog post: Did you read the rest of my comment? Did you even read the blog post you linked to?
> We certainly don’t have anything close to a complete understanding of how the basic laws actually play out in the real world — we don’t understand high-temperature superconductivity, or for that matter human consciousness
Can you try to consider my central point before replying: Are the rules governing physical reality simpler or more complex than the Collatz conjecture? Does there exist a (theoretical) computer that can "evaluate the Collatz conjecture to any degree of accuracy"?
EDIT 2: I'm not the one moving goalposts. On what grounds are you classifying the question whether a given number ends at 1 or not for the Collaz conjecture as an "inifite" computation? It's a simple boolean question, yes or no. All you have to do is build a computer that can answer yes or no for each integer. Isn't that simpler than answering the position of each atom in the ball after the throw? Each is just a function, what makes one more infinite than the other?
Friend, I'm not going to chase your ever changing text. Please just use the reply button.
qarl 22 hours ago [-]
> For everyday-life purposes, we can’t get around the fact that quantum mechanics makes it impossible to predict the future robustly.
You've moved the goalposts again. I said simulate, not predict. It is possible to simulate the entire Schrödinger wavefunction.
And PLEASE - just use the reply button. It is impossible to track every time you edit your comment.
streetfighter64 21 hours ago [-]
PLEASE just stop complaining about "moving the goalposts" when the issue is your own lack of clarity of both expression and reasoning. WHAT is the distinction between "simulate" and "predict"? You said neither by the way, you said "evaluate".
> ANY physical system can be evaluated to ANY DEGREE of accuracy by a computer
It's completely SENSELESS to claim that they are distinct, because in order to EVALUATE or SIMULATE the physical system you will need a FUNCTION which COMPUTES the STATE of the system at a given point in time. The only POSSIBLE distinction between SIMULATING and PREDICTING would be the time taken for the computation, but that is COMPLETELY IRRELEVANT as long as it is finite.
Again, your own source says:
> We CERTAINLY don’t have ANYTHING CLOSE to a complete UNDERSTANDING of how the basic laws actually play out in the real world
How does that square with your claim above?
qarl 21 hours ago [-]
> You said neither by the way, you said "evaluate".
You are entirely correct. I was sloppy in my first comment. I should have said simulate. My sincere apologies if that's been the crux of our dispute.
> The only POSSIBLE distinction between SIMULATING and PREDICTING would be the time taken for the computation
No. The distinction is in determining which "you" is you. When simulating the wavefunction, every you is simulated.
> We CERTAINLY don’t have ANYTHING CLOSE to a complete UNDERSTANDING of how the basic laws actually play out in the real world
It's very understandable if you include his following sentence:
> But these are manifestations of the underlying laws, not signs that our understanding of the laws are incomplete
He's saying we don't understand emergent behavior produced by the laws - not that the laws themselves are incomplete. E.g. we don't know how/why a bag of neurons turns into a person.
qarl 23 hours ago [-]
Re: Collatz - you've moved the goalposts. Answering Collatz requires solving a halting problem. I didn't claim that I could find the end of an infinite computation. I claimed to be able to simulate a finite one.
And - you should reply to my comments rather than edit your old ones.
throw310822 24 hours ago [-]
Does that mean that humans could produce mathematical proofs that are entirely logical and verifiable by other humans, but that cannot be formalised in any automatically verifiable language such as lean?
rsrsrs86 16 hours ago [-]
You need to do some studying _without_ chat gpt if you like math.
streetfighter64 1 days ago [-]
> Computation and computability is "the final frontier". Math is a "subset" of that.
In what sense? BB(n) is a prime example of an object that can be mathematically defined, yet is not computable. Or see BBB(n) for an "even more" uncomputable function. [0]
> An LLM in a loop where it can write on a tape can be Turing complete
What does this mean? A given LLM, like a given C program, can't really be Turing complete or not in a meaningful sense. The C programming language, or the concept of LLMs in general can be said to be Turning complete or not. Do you mean to state that LLMs in general are not Turing complete, but being "in a loop" somehow makes a difference?
> it can compute anything computable and is "bigger" than math at that point
Again, in what sense is it "bigger" than math? Lots of things are Turing complete, I wouldn't classify lambda calculus as "bigger" than math.
> It's impossible for finite number of LLMs to solve all theorems. This would imply that the busy beaver sequence is computable which implies the halting problem is decidable
LLMs use RNG for sampling, so they are not pure computers.
fsmv 5 hours ago [-]
Computable includes BPP
BoredomIsFun 5 hours ago [-]
Not sure if GPT based LLMs are polynomial time.
Timpanzee 1 days ago [-]
Even if the busy beaver sequence were computable and the halting problem were decidable, Gödel's incompleteness theorems would still prevent all theorems from being solved, regardless of if one used LLMs or not.
srcreigh 1 days ago [-]
I think there's a really important sense in which Godel's argument is not the full story.
IIUC, Godel's incompleteness is less about theorems and more about axiomatic systems. Given an axiomatic system, there are statements within it which cannot be proven or disproven. It's relatively unrelated to the platonic ideal of the theorem itself. The statements it considers are axiomatic-system-specific.
Another way to view it is, who cares if we can't prove or disprove "This statement is false". Ok, the axiomatic system is incomplete; fine. What's important is can the system prove a real theorem that I care about.
The busy beaver computability argument addresses these issues. The problem format is always "For Turing machine T with no input, does T halt?". This format can encode many math problems. And we know already that BB(432) is independent of ZF, aka, there is a 432-state TMs which ZF can't prove or disprove the halting behaviour of.
So BB looks at real theorems, ranks them, and we can ask what axiomatic systems can solve them or not. Godel looks at 1 axiomatic system and produces a toy theorem which the system can't solve. That's an extremely important difference!
The core issue is that any fixed LLM can only encode so many axiomatic systems in its states, and the fixed systems implies an upper bound in terms of the BB number which it can solve. Godel is only looking at one system at a time, while BB is a way to use a common problem format to rank every axiomatic system on an infinite number line.
gf000 1 days ago [-]
But a 432-state TM is a problem that we would like to "prove" is it not? It's not even a particularly complex one to begin with, my smartwatch has orders of magnitude more state then that and yet here we see that all of our math "fails" at it.
I'm no mathematician, but this is also the crux of Gödel's theorem, he just showed it in a more "hacky" and clever way - but BB(432)'s relation to ZF is also a consequence of Gödel's more general idea, is it not?
IsTom 1 days ago [-]
Even more concretely, the halting problem for turing machines with halting problem oracle would be undecidable for them. And if you could solve that you won't believe what problem would be undecidable. It's turtles all the way up.
moomin 1 days ago [-]
Pretty sure Gödel’s theorems imply the halting problem if you squint hard enough.
ogogmad 1 days ago [-]
The problem with what you're saying is that any old random true proposition about the integers is not necessarily interesting enough to be called a theorem. GIT (or the uncomputability of the Busy Beaver problem) does not establish a limitation on proving theorems, but rather on determining whether a proposition is true or not. Most propositions are ugly and irrelevant. So GIT/Busy Beaver is irrelevant.
-----
Oh, and: All proofs are conditional on axioms. If those axioms are computably enumerable, then all of their consequences are computably enumerable too.
gf000 1 days ago [-]
> Most propositions are ugly and irrelevant.
Most propositions may be ugly and irrelevant, but how do you know how many are not so and we just can't prove it? Also, what about stuff like Continuum Hypothesis, would you add it or not?
1 days ago [-]
1 days ago [-]
johnsmith1840 1 days ago [-]
lol same with people.
"Given infinite thinking time a finite number of humans will solve all theorems"
I also love the angle that this was not intelligence just brute force. As if the mathematicians didn't reeaaally want to solve this they were just too lazy to give it a good try.
What does AI have to actually do before you realize these things are actually smart?
alansaber 1 days ago [-]
Machines have a much higher capacity for work than human beings. Saying that these proofs did not require equivelant intelligence, but benefitted from sheer volume, does not strike me as unreasonable.
johnsmith1840 1 days ago [-]
Yes my only point here is that you can argue about the semantics of how smart they really are but they are undeniably smart.
Today it cost massive effort but it's possible 10-20yrs from now an AI could solve a problem like this in under an hour with a single thread on a free subscription paid for by serving an ad.
These arguments are so weak because you'll then have to make the same one a few years from now when it does something else impossible. The argument only stands if we assume no progress will occur.
streetfighter64 1 days ago [-]
Funny that you imagine a future where AI can solve complex math quickly, but humans are still watching ads, for some reason.
I just think you and the other guy have different definitions of "smart". There's no denying that LLMs are useful, but I don't know if I'd classify them as "smart". There were probably people in the 80s saying computers were "smart" because they could compute 78971 * 12341 faster than a human.
In what sense is a LLM "undeniably smart" but a CPU from the 80s isn't? Or would you define such a CPU as "smart"?
johnsmith1840 23 hours ago [-]
ASI may kill us all 100yrs from now but ads are forever my friend.
You're right though they largely are "smart" in the 80's computer sense. This is largely due to continual learning being unsolved.
BUT the more you look at them, research, and try experiments there's something there not in a 80s computer. If I had to guess maybe 1-5% of a humans ability but it's there. They are able to do novel things but ever step outside of their distribution takes exponential effort for every small addition. There is a true ability to adapt and learn new things on the fly, things never seen before. That is the the smart part. There something hidden in these things we don't understand that allows novel insights built from in context learning.
It's actually measurable in experimental settings but even there it's hard to tease out. I saw it mostly while doing CL training experiments. But I also see it while working with them for coding novel things.
But the more power we provide and farther down the road of this we go those 1-5% are things like solving unsolved math problems. No human solved these things. You say brute force, I say it needed massive effort to break out of it's distribution and get those small insights. It's very human like when taken at scale. The scary thing is that scale is getting smaller every day.
jimmaswell 1 days ago [-]
It feels goalpost-movey to downplay exploring a large search space efficiently in regards to "intelligence". If we dug into a human genius's brain and found it was somehow trying out a million ways to solve a a problem at once, no one would seriously suggest the person isn't actually intelligent.
And our brains must something like that at some physical level. You can't have a "turtles all the way down" of reasoning - the building blocks must be simpler. It must reduce to something like pathfinding and brute force at some point, weighted by factors in the system and maybe some randomness.
alansaber 1 days ago [-]
We have a romantic view of intelligence, perhaps stemming from intuition within the context of scientific discovery. Given enough intelligence, and enough context, a brilliant person can have a stroke of inspiration that allows them to make a major leap (a-la General Relativity or Fermats last theorem). We haven't seen THAT same capacity from a machine, but we see the more ordinary, unsexy grinding type of progress that represents 99.9% of scientific reality.
jimmaswell 1 days ago [-]
I would find it very interesting to train a model on information only available prior to the discovery of e.g. relativity or calculus and see if it can invent it. My intuition is that modern frontiers absolutely could. Not to take away from their brilliance, but Newton and Einstein were brilliant people who also happened to be in the perfect place at the perfect time - there's not so much "low hanging (i.e. approachable by one brilliant individual) but immensely valuable fruit" anymore.
williamcotton 1 days ago [-]
Is there truly anything new under the sun? Hasn't all of existence alway been here? All math, all physics? We could have merely discovered it. Intuition might be nothing more than combinations of what already exists rather than some sort of divine insight that unlocks previously unknowable mysteries.
jimmaswell 1 days ago [-]
Agreed. I don't believe intuition and creativity would be more than pattern recognition, remixing ideas, and trial and error combined with a kind of "genetic algorithm" approach if you deconstructed them into what the brain is actually doing.
glimshe 1 days ago [-]
This isn't necessarily true. There may be proofs so complex, they could exceed the limit of human cognition.
pfdietz 1 days ago [-]
There certainly are such proofs. Even for simple decidable theories we have very large lower bounds on decision complexity (like double exponential), which implies large lower bounds on the function from "length of theorem statement" to "length of shortest proof".
For undecidable theories, there is no computable function bounding this blowup from theorem length to proof length (otherwise, the theory would be decidable.)
pretzellogician 1 days ago [-]
(Background: trained, published, but still amateur mathematician.)
This is a cool blog post and I think you're going the right way, and beginning to get an understanding of the proof as you go.
I'd recommend continuing on the simplification and understanding route, until you yourself can follow the proof. Some suggestions, as I did something similar:
1. See if (or ask the AIs) if individual parts of the proof can be found elsewhere, i.e., is an argument just a copy of something else? If so, it's important to attribute this, but also this usually allows simplification ("by Theorem X", etc.)
2. Look for redundant patterns and try to combine them.
3. Ask the AI to be a critical reviewer from some journal, and try to fix its criticisms.
4. Continue simplifying! Assume that the final result may actually be relatively short.
Good luck!
zozbot234 1 days ago [-]
OP has reportedly been in contact with Prof. Mantova, who actually worked (jointly with S. L'Innocente) on the key human-authored results behind this AI proof and is arguably in the best position to understand exactly what the AI added that wasn't known before. (See the OP's thread on the Lean Zulip.) So this is happening, and we might see an actual paper publication of this result down the line (possibly encompassing multiple roughly self-contained papers, building up to the final result). The current AI-written version is way too obscure for that, and the AI-written human-targeted "summaries" are not really helpful. Again the OP is quite aware of this.
xworld21 21 hours ago [-]
Vincenzo (Mantova) here: yes, I have been reading bits and pieces of the proof and I can say for sure that the method is sound, at least for the first half (power series with real exponents). I haven't even tried reading the part that mentions the Cantor-Bendixson rank yet, although given how the rest went, I'd be really surprised if there's a problem there.
As with most interesting proofs, the number of core ideas is actually small, I'd say two for the real exponents, and presumably a third idea for lifting up to omnific integers. I have been redoing the real exponents part of the proof going on the ideas only, and with a few smarter choices, I am converging on something very short. And I mean very short, which is amazing. I didn't think the answer would be this close: it 'just' needs looking at the problem from the right angle, and also make a fairly bold guess at the outcome.
Dan's current proof is of course much longer. Between the fossilized ideas that Dan mentions in the post and the formalisation of previous results, there's a lot of cruft that inflates the proof but does not really help understanding what is going on. Luckily the word 'derivation' pops up early, otherwise it would have been very challenging to wade through the lemmas to find the important points.
GPerson 16 hours ago [-]
Thanks for encouraging the end of our career.
xworld21 10 hours ago [-]
I appreciate the sentiment, especially after the escalation of the last few weeks. But I am hoping this story can become an example of why mathematicians are still very much needed and LLMs are still severely lacking when it comes to displacing scientists. Conway's conjecture has no application whatsoever, even within pure maths, and the only point in pursuing it is what you can learn in the process, and whether you can make something beautiful. LLMs are consistently unable to do the latter (and yes, some mathematicians are also not very good at it... this is an old topic of discussion, just amplified by current events). To me it feels like the technology is still where it was in 2015 when DeepDream images came out: increasingly good at pattern matching, but still a lot more like dreaming than thinking. So our job is still there somewhere.
The problem is rather how quickly we can change our ways of working to make sure the training and hiring pipeline does not collapse. That's the disastrous scenario, for both the individuals affected and the discipline, that we must avert somehow. I wish we had an easy answer to that. I certainly don't. But I like to think that at least engaging with the public in a constructive way will have a net positive effect.
GPerson 5 hours ago [-]
I’m not actually upset with your actions I was just worked up at the time. You’re in a dilemma and graciously handling it the way you are is the only sensible path.
zozbot234 15 hours ago [-]
Needless to say, I disagree that what Prof. Mantova is planning to do (digesting the proof and making it human-understandable) represents the "end of [mathematicians'] career". Systematizing has always been a key part of human mathematical work, and tidying up a raw proof can be viewed as a kind of systematizing.
My hope is also that Mantova and very possibly L'Innocente will get a substantial share of credit for their role in the resolution of this conjecture by Conway: the AI would not have embarked on this were it not for their prior work. So even human mathematicians with an inclination for more exploratory "problem solving" will have plenty to do in the future. (The story is actually not that different for the recent Navier-Stokes forced blowup result, which also built on key conceptual work from 2023 by Córdoba and Martinez-Zoroa.)
GPerson 15 hours ago [-]
Luckily for these guys the problem is famous enough that giving some kind of credit for its resolution even makes sense at all. The vast majority of published work is not like this. There’s not going to be any credit divvied up to the thousands of people who’s work was probably involved in the recent formalization of FLT, which involved formalizing 300 thousand theorems.
zozbot234 15 hours ago [-]
Kevin Buzzard has reportedly been working on a more systematized (i.e. leveraging a more modern approach) and human-targeted formalization of FLT. I certainly hope that his work continues and he gets the deserved credit. Of course there is also some possibility that it won't, and that would mean you have a point after all.
GPerson 14 hours ago [-]
You’re just mentioning established and famous people. What about the average people who are just going to get fucked for spending their lives trying to promote humanity, just to be rewarded with the risk of economic ruin and no job prospects?
zozbot234 14 hours ago [-]
If Mantova and L'Innocente (as well as Córdoba and Martinez-Zoroa) count as "famous people" now, I would say that their fame is quite deserved! Would you disagree that AI played a significant part in surfacing their work? Aside from that, I'm not really sure what to make of your comment. I do sympathize for the researchers who are at risk of having their ongoing work "scooped" by AI (and for all we know, this may even include Mantova and L'Innocente) but the way to address that is to still have "divying up" of relevant human credit.
GPerson 14 hours ago [-]
Kevin Buzzard is famous. You’re pedantically misinterpreting my comment to feel clever.
You have nothing more to make of my comment because you don’t care to consider the actual problem here.
bonoboTP 7 hours ago [-]
Keep doing things that people want to pay for. If this becomes impossible, the world will face a much broader political crisis than "mathematics careers". How concerned were you when manufacturing industries died in certain regions and jobs got wiped out? Or it only matters when it's academics?
GPerson 5 hours ago [-]
Yes I was concerned. That kind of projection is only possible in people who lack empathy.
The math career is going to killed off 2-3 years before law and medicine and Wall Street banker. How does this make me more secure in any way?
bonoboTP 5 hours ago [-]
It shows that the issue is much broader and myopically focusing on how math PhD students will be evaluated etc. entirely misses the forest for the trees, it's just not even close in proportionality.
This is indeed also coming for law, medicine, and banking, though licensed professions will hold out for longer because you need someone to put in jail when things go wrong. The problem is that all this is extremely over politicized and nobody is able to think clearly. They want to simultaneously say all this is just hype and a bubble and will go away like NFTs did, and also are starting to worry about economic replacement. Some more coherent political narrative will have to be formed.
GPerson 5 hours ago [-]
I’ll state this again. How does my economic redundancy get mitigated in anyway by being the first on the chopping block by several years? Do you think being jobless for 3 years is just missing the forest for the trees?
bonoboTP 5 hours ago [-]
Since it will be coming for everyone, the solution will be major social upheaval with some consequence for all of humanity, hopefully a good one where we somehow manage to keep on living some kind of good life.
Regarding being jobless for the 3 intervening years, it is certainly a personal concern but in this temporary phase there are still some other jobs for smart people. Once there aren't any, we are entering the part that I was talking about where you will be far from alone and you can join together to exert some kind of political pressure but it will not be about math PhDs, but employment as a whole. And it may not be very effective if AI is on the other side, not on yours. Yeah, it sounds like scifi, and people want to dismiss scifi concerns and instead focus just one inch ahead of their toes, instead of seeing the writing on the wall.
GPerson 4 hours ago [-]
For some reason you think I’m opposed to the use of AI in mathematics. This is not the case. If this guy sat down to actually learn the math and then did actual work to disseminate the knowledge in a positive sustainable way it would be a different story. His actions are just here to inspire more people to use this tool in counterproductive ways.
Xmd5a 12 hours ago [-]
I'm not working on some famous conjecture, just following my curiosity, and I think I found something (minor and specific) that should reignite a young mathematician's interest for work she did almost a decade ago, for a phd she never used (she left academia). I'm going to have flowers delivered to her, with a brief note explaining what I think I found and encouragement to resume her academic career and reach out to a her former co-authors.
unified101 14 hours ago [-]
What's this obsession with credit. From elsewhere in the thread:
> people spending their lives trying to promote humanity
What does this even mean? Become a monk? They are credited with helping humanity.
GPerson 14 hours ago [-]
You have no idea how mathematics has progressed for thousands of years. How about having some humility and understanding the culture you’re cheering on the death of? It’s a goddamn credit based system.
Xmd5a 12 hours ago [-]
> the culture you’re cheering on the death of
Ahaha. Flowers are the way to go.
bonoboTP 8 hours ago [-]
Do you buy stuff made in factories of do you open your wallet for everything handmade craft products? Or it's fine as long as craftsmen's jobs got replaced, just not when yours? Do you ever use a self-checkout? An ATM? Or do you pay extra to get cash from a human cashier at the bank?
GPerson 5 hours ago [-]
Do you actually think this is a good argument? I had literally zero say in what happened in the past. I did not support the Midwest getting gutted. I wasn’t even able to vote when that was happening? Would it shock you that I generally think capital is misaligned to humanity?
bonoboTP 4 hours ago [-]
> Would it shock you that I generally think capital is misaligned to humanity?
You still prefer to get cheaper options yourself. I know this argument gets caricatured in the "yet you participate in society" meme, but the point is that this is the aggregate result of individual humans making decisions on where to allocate their resources. You can attack this using various ideological and religious frameworks, but if it's just some stoner college freshman's communism, I'm not interested (neither if it's the more potent version that dispossessed my ancestors in Eastern Europe).
GPerson 4 hours ago [-]
Also is your position really that nobody can point to problems or criticize anything without a fully fleshed political system to replace the USA with? I don’t support the excesses of the system of the USSR or 1950s-1980s China.
bonoboTP 2 hours ago [-]
You can criticize things of course, like you can shake your fist at a cloud too when you'd rather want a sunny day, but the results may vary. You're asking people to refrain from using a cost-saving technology and instead spend more money on things that can be obtained easier with the tech. This means going against one's own rational interest. Such things usually happen willfully when the person believes in a religion like the Amish or Hasidic Jews. So having a fully fleshed political/ideological system on as similar level will be necessary here.
GPerson 4 hours ago [-]
I don’t “prefer” to get cheaper options in any meaningful sense. There is no choice of mine involved in what products corporations serve. I have to buy the cheapest option because I have a poverty wage and have to survive.
bonoboTP 4 hours ago [-]
The meaningful sense is when you open your wallet and buy one option when another option is there. Of course you have a budget and choose rationally, which is kinda like a forced choice. But you could always go and buy from the mom and pop shop and you could buy handmade clothes. Well, you don't have the money for it. Neither did people have it generations ago, they just wore worse clothes, went barefeet to school etc. You could still do that, walk barefoot until you can afford a handmade shoe and keep that for decades. It is in fact your preference not to do this, even if you want to claim that social pressure predetermines that you can't go around barefoot etc. At some point we disprove free will and agency entirely. Nobody in any historical era had more choice to express their preferences than consumers today.
GPerson 4 hours ago [-]
I am not allowed to walk into my office without shoes. My shoes have had holes for 90% of my life. I assure you I own fewer things than you are imagining.
unified101 14 hours ago [-]
WTF? Have some humility towards someone who took the time to talk about their work. And made intellectual progress.
Air your LLM greviences someplace else.
GPerson 14 hours ago [-]
Nope I’m airing them right here. This guy didn’t do any work except tell the bot to continue for a month. That’s not work, that’s just unhealthy and negative. He’s made no intellectual progress. He doesn’t even know any mathematics and has never cared to learn. He’s just lucky other people are there and kind enough to let him dump his slop onto. They could have done this and gotten a lot more out of it.
zozbot234 14 hours ago [-]
You can definitely argue that the user's direction and curation work was intellectually trivial (though there's meaningful room for disagreement even there, especially wrt. having the AI stick to established terminology/broad approaches - this is arguably a sort of successful "systematizing" work, though only in a very minimal sense) but this was not a one-shotted result. The blog post is extremely clear about that.
And of course, going by their own admission, they couldn't "have done this themselves": the most you can argue wrt. this is that Mantova and L'Innocente, or some other narrow domain experts, might have done this themselves and that AI "scooped" this result from them.
GPerson 14 hours ago [-]
Nobody said it was one-shotted? It was mindlessly “continue-shotted” except for the brilliant idea of upgrading the model. Obviously the models are going to be improved to the point where typing continue continue, how ya feeling, continue continue, okay let’s double check this, upgrade model, continue continue, will be less necessary.
danabramov 9 hours ago [-]
I was not "mindlessly continue-shotting". The way you describe it is literally the same as "one-shotting" and, as I explained in the article, it simply doesn't work. You are welcome to try it yourself on this problem to verify that.
Yes, I was not doing any mathematical work in curating the output, but the article makes it quite clear that pivots and constraints the LLM would not impose on itself were critical to actually making progress.
Also:
>He doesn’t even know any mathematics and has never cared to learn
While I don't know enough mathematics to work on this problem, claiming something like this is preposterous. As I link in the first paragraph of the article, I've been learning mathematics on my own by going through Terence Tao's Analysis book and solving exercises. I'm familiar with the concepts of mathematical definitions, proofs, etc. I've gotten about halfway through the book solving them on paper before abandoning it (and later got through the first few chapters in Lean, also solving every exercise — by hand, mind you). Sure, this doesn't make me a mathematician, but I'm closer to a dropout first-year student than to someone who has "never cared to learn".
GPerson 4 hours ago [-]
You’re a guy who benefited from easy choices which led to you a life of luxury and now you use your high perch to shit on people who were stupid enough to work hard to do something more meaningful with their lives than attain wealth and social status.
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week. I won’t hide this comment though it is shameful.
GPerson 5 hours ago [-]
I read your blog and nothing in it indicates you did anything besides mindless continue shotting. You don’t know this because you don’t know anything about the culture you’re stomping on.
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
danabramov 5 hours ago [-]
I invite you to continue-shot this result. I haven’t been able to do that with the current models in a separate control session.
It might be possible to plainly continue-shot it with more powerful models in the future. I agree that whatever I did is probably automatable.
GPerson 4 hours ago [-]
A distinction without a difference. I can’t believe you’re actually trying to take some kind of credit for this. That’s just so shameless.
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
danabramov 4 hours ago [-]
What does "taking credit" even mean here? I am not saying that I contributed to the mathematical work in the proof. I am simply disagreeing with you that the process is fair to describe as plain continue-shotting. You saying "without a difference" does not actually make your claim true. There clearly is a difference between a specific technique getting to the result and that technique not getting to the result. Whether or not you like the technique, and whether or not you consider a result obtained through that technique of any value. I thought there is some value in sharing both the technique and the result, and that's why I published a post about it. I think it's slightly different from "an AI company threw 50,000 agents on it" and it's also slightly different from "I just said Claude to work hard and it got me a solution", and that makes it worth sharing with other people.
I am not saying that my work constitutes a mathematical contribution on its own. Not any more than stumbling upon an anonymous manuscript with the solution would constitute a mathematical contribution. I do, however, think that it can lead to a mathematical contribution if any mathematicians consider it worthwhile to do something with it. Whether or not they consider it worthwhile is not up to me.
GPerson 3 hours ago [-]
There is no value in your technique when 2 months, or 6 months, or 2 years from now the model will improve. There is no skill in suggesting to double check work. Everything you did could be trivially automated with today’s models anyway, as you are aware. Taking credit is trying to pretend you had some meaningful role here.
It is a distinction without a difference because I want to live in a world where people get to fill their lives with meaningful things, and are not forced into Uber delivery driving jobs just because rich people like you think it’s fun to put their name next to something other people made prestigious.
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
danabramov 3 hours ago [-]
I'm not sure it is trivially automatable with today's models since they get "carried away" too much, and I'm not sure there's a good way to prevent the supervisor from drifting. But I would like somebody to attempt that, which is part of the reason for my posting. If someone can automate whatever I was doing, it should be possible to get to deeper results than current models allow. Or maybe the next models would just be that good, I don't know.
Re: "rich", I've essentially spent $400 on this (in subsidized subscriptions), plus my free time being a mindless drone. Given that you assume my role is automatable, it sounds like this is relatively accessible to anyone with $400 (as long as AI companies continue subsidizing the frontier models). I don't think I've had some kind of an unfair advantage beyond that. If anything, a proper mathematician would probably be able to derive the result much faster with the same tools.
I don't know how the broad availability of these tools (to mathematicians and non-mathematicians alike) will change the field, what is considered prestigious, what work gets funding, how it affects the pipeline, etc. You seem to be implying that even testing the limits of these tools, or at least publishing the results obtained with them, is unethical in itself, even though it is broadly accessible now. I can understand this point of view.
unified101 12 hours ago [-]
The direction of continue shotting will develop a new culture that lead to a much more wider understanding for humans in the field of math. Find a way to accept this new culture and you'll thrive.
GPerson 5 hours ago [-]
It won’t. It will lead to a culture which makes it impossible to actually spend a life learning mathematics deeply.
unified101 4 hours ago [-]
It will. you just lack imagination.
GPerson 4 hours ago [-]
It won’t. You just have no idea how the world actually works.
unified101 3 hours ago [-]
It will. you lack imagination, and have a unnecessary high opinion of yourself.
GPerson 3 hours ago [-]
It won’t. I don’t.
unified101 3 hours ago [-]
I apologize to GPerson for needlessly trying to convince him to my point of view. He seems to be an smart person deeply affected by how LLM are affecting his vocation and work. It was wrong of me to do this and I will take a break from this website for one week.
unholiness 1 days ago [-]
A wonderfully made introduction to the surreal numbers and their surrounding game theoretic concepts is this video on Hackenbush[0], a winner in 3Blue1Brown's Summer of Math competition.
>I’ve emailed some of the mathematicians with a few proposed typo fixes, and I got confirmation that at least a few of those fixes seemed real. However, some of the problems that weren’t backed by Lean also turned out to be misunderstandings.
I think this project is really neat, but is it appropriate to cold email specialists before you've put in enough hours of effort to describe yourself as more than an "amateur"? OP's emails may have been helpful, but billions of people use these LLMs to wade into new areas and email is already low signal-to-noise.
danabramov 1 days ago [-]
Yeah it's a pretty tough question! I've resisted doing that until I had a relatively high certainty that their published results contained minor mistakes, which I assumed they would want to know about. I've also been explicitly apologetic and tried to keep it super brief.
vessenes 1 days ago [-]
A counterpoint - I was told a story by one of my professors in the late 1990s, about one of his professors -- he'd written a thesis, gotten hired somewhere like Princeton, and taught there for a few years as Dr. <Somebody>. One day he received a letter pointing out a construction flaw in his thesis. He brought it to the department head who read the letter, and said "Well, Mr. Somebody, ..." Ultimately he fixed the proof.
Upshot, if there are real errors in published work, I think most mathematicians want to know about them.
stevemk14ebr 1 days ago [-]
Yea but that was a person who actually put in work and had to think about and understand the problem. They didn't just generate something with a magic box.
dev_dan_2 1 days ago [-]
As a software engineer, I could not care less about how a bug was found or by whom, as long as I can quickly verify it is correct, I always appreciate being able to improve my work. I don't see why it should be different for mathematicians (the ones I knew would think similarly, I would assume.)
danabramov 1 days ago [-]
If I may offer an analogy, what I tried to do is essentially reporting a bug after having a failing unit test that exercises the public API, without looking into the black box of internals.
chasd00 23 hours ago [-]
> They didn't just generate something with a magic box.
It doesn't matter who found the error nor how it was found. An error is an error.
bonoboTP 7 hours ago [-]
This issue really separates the wheat from the chaff. The dirty secret is that a lot of academic published work contains errors but the authors also have fragile egos. It reminds me of when Data Colada exposed someone for fraud and they accused the exposers of "methodological terrorism" (something in psychology or social science). Just imagine what happens once papers get exposed at scale.
Now here there was no fraud just genuine error, but it will annoy people nonetheless and scrape their ego that someone uninitiated can just type some stuff in a magic box and conclude that they, the established published, tenured mathematician with awards and medals can be wrong.
1 days ago [-]
Feathercrown 1 days ago [-]
I find the way the author communicates with the LLM fascinating. For example:
> However, I didn’t just want any result; I wanted something that pulls me.
> Initially, I asked Claude:
> Me: which unsolved problems in the Surreal Numbers research program pull you the most and why?
Note the switch from "pulls me" to "pull[s] you". What is the author's perception of the relationship/boundary between them and the LLM here?
1. Are they using it to find things it flags as interesting in hopes they might also find it interesting?
2. Do they consider "interesting" to be a universal (observer-independent) trait and are using the LLM to find things that are interesting?
3. Have they delegated their desire to find something interesting to the LLM so that it can instead find something that it flags as interesting, regardless of how the author feels?
4. Do they see it as a part of their thought process, and so do not distinguish "you" from "me"?
5. Do they see it as part of them, and are referring to the combined entity in the second person?
I would love clarification on this.
danabramov 1 days ago [-]
Hah, very interesting question!
Let me first clarify my relationship with mathematics. I think of myself as "an awestruck observer from a distance". I find some parts that I understand beautiful, and I have also tried to understand some of the basics rigorously. However, I generally just can't make my way through any serious paper, as I both lack the prerequisites and struggle with the amount of inference mathematics tends to place on the reader. That's the "from a distance" part.
Now, about picking the problem. I am genuinely "pulled by" surreal numbers themselves. I find them irresistibly beautiful. There is also a bit of bitterness around how they haven't fulfilled their promise (yet?) as Conway hoped they would be able to become a better foundation for some mathematics. But they are a bit too difficult to prove things about so far, and we know too little about them. So what "pulls me" also is a possibility of making enough dents in this that we would be able to use them more broadly, and learn even more things about them.
However, I do not know the details of the latest research. I don't know which problems have actually been solved, which pursue Conway's original vision vs narrower approaches, and which are elegant enough to feel "awestruck" enough about. So this is an invitation from me to LLM to share what it "feels pulled by" (for whatever definition; I think of it as just navigating the languagespace) , and then sifting through that list to see if something it lists makes me feel something. I would assume that with the field currently being so small (serious mathematicians mostly don't care about surreals), it's easy to get the LLM "excited" (again, just a vector in the languagespace) enough that it would give me genuinely interesting candidates. Then it's up to me to sift through them and see if they "speak" to me.
It's like asking a mathrock nerd to share their favorite mathrock albums. Niche enough that you'd likely get good results. Then you can listen and form an opinion.
In this particular example, the "ONAG birthday" and "maybe last Conway's unsolved conjecture about surreals" part spoke to me emotionally, the statement itself amazed me with its simplicity, and I felt "blood in the water" related to the recent results bringing the conjecture closer. So I felt the pull myself and went with it.
scotty79 1 days ago [-]
I use you, me, us interchangably because I don't think it really matters for anything and I don't need to reaffirm my individuality with such words.
incr_me 18 hours ago [-]
> I don't need to reaffirm my individuality with such words.
Record scratch
scotty79 56 minutes ago [-]
It's for your benefit not mine. You can guess if I meant sinfular or plural you and decide if it matters.
hlynurd 9 hours ago [-]
us really messed up there
howunfortunate 1 days ago [-]
> On the second day, there are two gaps: “between nothing and zero” and “between zero and nothing”. Two numbers spawn in those two gaps. Call them –1 and 1.
Got lost here. I think I'm officially too dumb for math.
gjm11 1 days ago [-]
I don't think this is your fault; the description isn't very explicit. Let me try to do a bit better. (I'll also try to go somewhat further, and you should not be discouraged if at some point it stops making sense.)
You can think of the "surreal numbers" as being built up step by step. We start out with no numbers at all, and then we repeatedly do a construction that makes some new numbers.
A surreal number is made from two sets of (pre-existing) surreal numbers. We typically call them L and R, for "left" and "right", and sometimes write it as L|R or {L|R} or something like that. The "left" numbers have to be smaller than the "right" numbers. The resulting number will turn out to be, in a certain sense, the "simplest" number in between all the left numbers and all the right numbers.
Now, as I said, we start out with no numbers at all. It might seem like that gives us no way to proceed, but it does: even given no numbers at all, we can still make a set of numbers, namely the empty set! So we can use that for both L and R, getting ∅|∅. Empty sets on both sides. We call this 0, and it will turn out to behave in the way you'd expect the number 0 to behave.
Now we suddenly have another set available, namely {0}, the set containing only zero. Which means that instead of being able to make one number, maybe we can make four: ∅|∅, ∅|{0}, {0}|∅, {0}|{0}. The first of these we already knew about. The last isn't actually admissible -- remember that the "left" numbers have to be smaller than the "right" numbers, which is "vacuously" true when one of those sets is empty (it means "if you have a number x in the left set, and a number y in the right set, then x<y", and if there are no numbers in the left set or no numbers in the right set then that's trivially true) but isn't true when both sets contain 0 because 0<0 is false.
So actually we get two new numbers: ∅|{0} and {0}|∅. The first fits into what OP calls the gap "between nothing and zero". The second first into what OP calls "the gap between zero and nothing". In both cases, "zero" means a number and "nothing" means a space where we don't yet have any numbers.
The number ∅|{0} is called -1 (it has to lie to the left of 0, and there's no constraint on its left, and -1 is "the simplest number less than 0") and the number {0}|∅ is called +1 (it has to lie to the right of 0, and there's no constraint on its right, and +1 is "the simplest number greater than 0").
I should explicitly acknowledge that I haven't defined what "less than" and "greater than" actually mean for these numbers, nor anything else about how they relate to one another that could possibly justify giving these things the specific names 0, -1, and +1. But there are definitions for "less than" and "greater than" and "plus" and "minus" and so forth, and the whole thing does turn out to work very nicely.
Anyway, once we've got these numbers we have eight possible sets that can go on the left or on the right. The requirement for left-things to be smaller than right-things reduces the possibilities somewhat, and the actual new numbers we get next time around are: ∅|{-1}, which turns out to be -2; {-1}|{0} which turns out to be -1/2; {0}|{+1} which turns out to be +1/2; {+1}|∅ which turns out to be +2. We also get some already-existing numbers in new ways; for instance, {-1}|{+1} is actually equal to 0 ("0 is the simplest number between -1 and +1"). Again, I should explicitly acknowlege that I haven't said anything about how you determine when two of these things are actually equal; again, it does all turn out to work properly.
If you keep going with this construction, you produce all the integers, two at a time, and also all the "dyadic rationals", meaning fractions where the denominator is a power of 2. And then, once you've got all those, at the next stage of construction you abruptly get all the real numbers -- e.g., the square root of 2 is L|R where L = {dyadic rational numbers that are negative or have a square smaller than 2} and R = {dyadic rational numbers that are positive and have a square larger than 2} -- and you also get {0,1,2,3,4,...}|∅, conventionally written as a lower-case Greek letter omega, which is an infinite number, larger than all the integers. (And its negation.) And {0}|{1,1/2,1/3,1/4,...} which is an infinitesimal number, positive but smaller than any ratio of positive integers. And you can then proceed further and construct a vast infinitude of numbers, including all the real numbers (which we've already made) and all of the so-called infinite ordinals (which you can kinda think of as being a sort of "infinite positive integer", though there's more to them than that) and much more, all in a system that lets you do arithmetic and suchlike. It's very elegant, if your brain has been twisted into the mathematician-y shape that finds such things elegant.
danabramov 24 hours ago [-]
Thank you for writing a detailed explanation! I've slightly edited mine to explain the infinity jump. Yours is, of course, much more detailed.
For the infinitesimal number, I think it makes more sense to use {0}|{1,1/2,1/4,1/8,...} since it gets born at the same day as say 1/3. So it is easier to understand how it arises without "waiting" for all reals.
gjm11 23 hours ago [-]
Oops, that was an oversight: indeed you don't get all the rationals I need for what I wrote until "one day later" (in Knuth's terminology). Regrettably I'm too late to edit what I wrote above.
atuladhar 1 days ago [-]
Thank you for this explanation! The construction is so elegant, and in a way, the basic idea is simple (?) -- I wonder why it wasn't thought up of much earlier than it was. Maybe it's a little bit like https://en.wikipedia.org/wiki/Egg_of_Columbus
gjm11 1 days ago [-]
Not only is the basic idea simple, it's a sort of generalization of two other things that were already well known but before Conway were thought of as completely independent.
First: the construction of the real numbers from (traditionally) the rational numbers by means of "Dedekind cuts" (sometimes called "Dedekind sections"). The idea is that if you're trying to build up the machinery of mathematics from scratch, it's not too hard to go step by step from (say) sets to nonnegative integers to integers to rational numbers, but it's harder to get from there to the real numbers, and Dedekind's idea is to say that e.g. the square root of 2 is the way of chopping the rational numbers into "things less than the square root of 2" and "things greater than the square root of 2".
Second: the construction of the ordinal numbers (a sort of generalization of the notion of "nonnegative integer" that allows the numbers to get very infinite) due to von Neumann: you start off saying that zero "is" the empty set, and then you repeatedly say: the next ordinal "is" the set of all the ordinals you've constructed so far. So, e.g., 1 = {0}, and then 2 = {0,1}, etc. -- but once you've constructed all the nonnegative integers you can then look at {0,1,2,...} and that's a new ordinal typically called ω, and then you can take {0,1,2,...,ω} and call it ω+1, and so on and so forth.
Both of these are special cases of what Conway does: Dedekind's is the case where all the numbers are rational numbers and you don't allow either set to be empty, and von Neumann's is where you _require_ the right-hand set to be empty.
There's a further connection, which I believe is how Conway found these things in the first place: if in the definition of surreal numbers you delete the requirement that everything in L has to be less than everything in R, then what you've got is (more or less) the definition of a position in a two-player game. L is the set of positions one player can move to, R is the set of positions the other player can move to. (I say "more or less" because e.g. in many games you're allowed to repeat positions, and games may have complicated winning conditions or involve chance or whatever.) And there's a whole rather nice thing called "combinatorial game theory" that's all about these, and from that perspective numbers are just one particular kind of (position in a) game. (Specifically, a number is a game in which at no point in the subsequent gameplay can it ever make your position better for you to make a move: you'd always rather pass if you could.)
What an interesting construction. Thank you from a curious layman for your write-up. I thought it was pretty easy to follow. I'd heard of the surreal numbers before and never knew about the construction mind-game behind them.
samiskin 1 days ago [-]
Thank you this was very well explained
dubcanada 1 days ago [-]
I think it's just we don't have a way to say/write these numbers. So you make up a way to write them (-1 and 1) and continue.
The numbers don't matter and you could replace -1 and 1 with anything. It's just easier to begin your new fake number at - 1 and 1. Because position does matter.
howunfortunate 1 days ago [-]
This is very helpful!
Basically what I take away is that we're inventing a new number system from scratch. So we're not "proving" that 1 is a number between 0 and the empty set. We're defining it as such, and it just so happens that a number system defined this way works out in convergent ways with other mathematics.
Is that roughly right?
danabramov 1 days ago [-]
Exactly.
xg15 1 days ago [-]
There has to be some procedure how to come up with "new" numbers though, if you want to have more in the end than just a fancy binary tree - in particular if you want to map your "fake numbers" to the reals, infinity, etc.
danabramov 1 days ago [-]
This procedure is enough. If you define addition and other operations in a certain way (as Conway did), it turns out that on the omega-th day (i.e. after initial infinite steps), all reals will be born.
skeledrew 1 days ago [-]
> position does matter
Only as a mental abstraction that's based on our experience/concept of space+time.
This would make a lot more sense to me if "nothing" and "nothing" were instead "-inf" and "+inf"
Some other comments clarify that "nothing" is more accurately "the empty set". This is helpful because at first I wrongly synonomized "nothing" with zero. But now I get tripped up on the "between" language. Maybe it's a lack of background in sets, but I don't know what "between" implies for an integer (zero) and a set (the empty set).
danabramov 1 days ago [-]
I didn't want to introduce the notion of infinity because there are actual "infinite numbers" on the surreal number line. I've kind of tried to have both the simplicity of set-theoretic definition and the intuition of the number line, and slightly bungled the exposition. I hope the newly added diagram helps.
I'm not a mathematician and "nothing" doesn't really make sense for me either. But I guess the problem with inf might be that it'd be strange to get +2 as the next thing between +1 and +inf, while getting +1.5 between +1 and +2.
danabramov 1 days ago [-]
It's more that surreal numbers already include infinities like ω, ω + 1, and so on, so "inf" felt like a concept I want to avoid. The actual description is set-theoretic and uses empty sets there ("nothing to the left", "nothing to the right") so I took that a bit too literally.
beingforthebene 1 days ago [-]
Wow thanks. I'm a mathematician and also got lost at the step. This illustration makes the construction much more clear. The text isn't really describing this process well
danabramov 1 days ago [-]
No problem! I've added it to the article, appreciate the feedback.
bmacho 23 hours ago [-]
> On the second day, there are two gaps: “between nothing and zero” and “between zero and nothing”. Two numbers spawn in those two gaps. Call them –1 and 1.
That's a mistake. They should've written "empty set of surreal numbers" and not "nothing".
It is a constructive theory, like sets/ordinals. For ordinals you can use ∅, { }, ∪ and you construct
∅, {∅}, {{∅},∅}, ... (von Neumann ordinals).
For surreal numbers you use the form { A | B } where A and B are sets of surreal numbers. Some restrictions apply so not all of these forms will be surreal numbers.
You build up surreal numbers as
{∅|∅}, {∅|{∅|∅}}, {{∅|∅}|∅}
and so on.
---
Edit: A happy accident: I denoted the "empty set of surreal numbers" with the symbol ∅. It works, and it gives you the surreal numbers. But if you think of ∅ as the empty set in set theory, then the same construction (using an ordered-pair construction) gives you the surreal numbers as sets!
danabramov 19 hours ago [-]
What is the problem with saying “nothing” to mean “empty set” in an informal explanation? If you put apples between two bags, and the right bag is empty, it’s reasonable to say “nothing” accurately describes the contents of the right bag.
I’m slightly stretching the metaphor here because the number lines gives me enough structure (order expressed visually) and I only deal with at most one set item at a time (since we construct in the order of simplicity and can use the already constructed numbers), so it (IMO) unnecessarily complicates things to even talk about sets when we’re sort of just making cuts on the line. But in either case I don’t see the problem with colloquially saying “nothing” here.
xg15 1 days ago [-]
I think he just wrote it in a confusing way. The quote before says:
> (crucially, “to the left of all” and “to the right of all” also count as “gaps”)
So there are two "nothings" here, left of "all" - i.e. the zero - and right of it.
Though I'm not quite sure how you'd get infinite or irrational numbers by this procedure. Wouldn't you simply get the rational numbers by this?
(unless the "put a number" step is doing more work here than it seems. He doesn't really say which number to put there. In the examples, he mostly did "new number = (left number + right number) / 2", with special cases if any number is "nothing" - but he never actually wrote what the rules are here.)
gf000 1 days ago [-]
You have infinite steps. Pi is just taking the correct turn an infinite number of times.
skeledrew 1 days ago [-]
Yeah I'm curious about the difference between "nothing" and "0", but I just decided to roll with it. Until the Greek letters made my head start to spin as they usually do.
danabramov 1 days ago [-]
Hope the newly added picture helps see each step.
svachalek 1 days ago [-]
The image helped for me at least.
Smaug123 1 days ago [-]
(Apparently I was extremely unclear with this text. For clarity: if you want to actually understand surreal numbers, go and read On Numbers and Games, by Conway, which is a delightful book; or get an LLM to talk you through Wikipedia. Original text follows.)
It’s a terrible explanation. A surreal number is defined as a pair of sets of surreal numbers (where you fiddle around the recursion in that definition by defining them in waves, so strictly speaking you’re defining “the surreal numbers born at time T” for each individual T given access to the surreal numbers born at all earlier times, and then you “take the union across all times”, scare quotes because there are too many times for this to result in a set). Zero is a surreal number but the LLM is using the word “zero” to mean “the set containing just the surreal number 0”; “nothing” here is the LLM’s obtuse word for the empty set. Wikipedia may actually be easier to follow.
danabramov 1 days ago [-]
LLM didn't write anything in my post; these are all my words and my choices. Conway himself described surreal generation in short like this in ONAG:
> We may say that Cantor was only interested in moving ever rightwards, whereas Dedekind stopped to fill in the gaps, so that R was always empty for Cantor, never empty for Dedekind. It is remarkable that by dropping these restrictions we obtain a theory that is both more general and more easy to work with.
This is precisely the intuition I present to the reader of the article. I am relying on visual aid (concretely, the ordered number line) to imply the machinery explicit in the actual recursive definition. The intended reader of this article is not a mathematician, and I think intuition is vastly more important here.
And I don't think I'm conflating 0 with {0} as you claim. When I say zero is "between nothing and nothing", I mean 0 := {|}. When I say one is "between zero and nothing", I mean 1 := {0|}. When I say 1/2 is "between 0 and 1", I mean 1/2 := {0|1}. And so on. I elide "the simplest number" because I am already going in the order of simplicity. I do not need to explain that alternative spellings like 1/2 = {0.2 | 1} are valid because it is not relevant to establishing the mental model of birthdays.
For the finite cases in my explanation, I do not need to explain that the left and the right parts form sets because I only ever need at most one surreal on either side to define the next generation. I also do not need to state the left/right order condition because it is already visually implied by the picture. For the same reason, I do not need to explicitly quantify over the set of earlier-born surreals, since in these finite cases, if we go birthday by birthday, each next day's surreals are definable via the numbers already constructed by the previous day.
I agree that these finite examples don't spell out how to handle infinitely many bounds at the omega-th day, which is where I believe the illustration embedded below is more helpful. I still think "a gap beyond 0, 1, 2, 3, ... with nothing on the right" is a useful intuition when we get there.
If you can’t explain it better in the same amount of characters (or fewer), then I don’t think you’re qualified to “nuh-uh!!!” anyone. Sorry buddy.
Smaug123 1 days ago [-]
I mean, I was intending to supply the words that would link the LLM’s explanation to a more normal one, not to explain it; apparently that was extremely unclear. An actual explanation is much longer, as indeed I attempted to indicate by pointing to Wikipedia and saying that it might be more clear.
“Doing better than a totally useless explanation in fewer characters” is in general impossible, of course, eg if the first explanation has only one character.
skeledrew 1 days ago [-]
... wut? :/
echelon 1 days ago [-]
> I think I'm officially too dumb for math.
I'm hoping someone develops an interactive tutor that can teach any subject to any depth.
The tutor should optimize its pedagogy. It should use online RL to adapt to a learner's ideal learning style, model what the student understands and to what degree, and understand what the gaps and next steps are.
I'd subscribe in a heartbeat.
skeledrew 1 days ago [-]
There's a skill for that (haven't tried myself but intend to; other of author's skills I've used have been a game changer).
Ask for analogy in terms of a thing you are an expert in.
ie. i am an expert at zig, explain this c++ in terms of zig
mrguyorama 1 days ago [-]
This is called college.
echelon 1 days ago [-]
Broadly, universities are too expensive, inequitable, suboptimal, not portable, and slow. There's a huge amount of room for improvement.
Universities are great for networking, starting projects with other students (not the ones professors mandate), and learning lab sciences. In research, they're great for institutional knowledge, having a community of peers, getting guidance from research advisors, having real equipment and funding, etc. But there's a great need for AI tools to accelerate learning outside of that setting.
Anecdotally, I'm a working adult. I'm not going to waste time in college again. I need this for me.
jdw64 1 days ago [-]
I feel the same way. Want to become a dumb and dumber duo? I believe you could be my friend.
howunfortunate 1 days ago [-]
Pitch: a mini-series where a Dumb and Dumber duo get access to unlimited tokens via a roommates's account (who is an intern at a frontier lab).
In each episode they make a major science-fiction style breakthrough and grapple with the consequences without revealing themselves.
DrewADesign 1 days ago [-]
Plot twist: they were both sold on early investment in companies that survived the .com bust. Now they’re VCs that everybody worships as business geniuses even though they’re just lucky idiots, and the sycophantic chatbots finally let them feel as smart as everyone says they are, and a whole bunch of hype-drunk fans are feeding into it.
bonoboTP 1 days ago [-]
My main thought is that he was performing something general here that is actually valuable and hard for a large proportion of humanity. It's like when Google search was a difficult thing, or troubleshooting a PC. what he is able to do here is actually a rare skill, called intelligence and he may think it's nothing, but it's actually very rare and hard for most people. The kind of judgment and interpretation of output without deep expertise is actually a very rare ability.
GPerson 14 hours ago [-]
No this guy demonstrated zero intelligence unless the word has already lost all meaning.
bonoboTP 10 hours ago [-]
I think most people could not get this conjecture proven if given infinite time at the keyboard with the same chatbot. He demonstrated much more initiative and problem solving instinct than what a large majority of people are capable of. Have you seen people trying to trouble shoot a printer or a software issue? Often the answer is just to search Google and read some forum post and execute the steps and deal with any arising minimal obstacles. Still insurmountable to most.
GPerson 5 hours ago [-]
I disagree. I see nothing creative about this guy’s efforts. The models will continue improving and soon his creative leaps of asking it to double check won’t be necessary.
Is the author suggesting that understanding has no value? because this is what I could gather from reading this.
Edit: Another depressing fact is that he generously paid to LLM megacorps while piggy backing on human help for free and in the end calls the proof his or LLM's.
danabramov 8 hours ago [-]
Is a YouTuber speedrunning a game by exploiting an integer overflow bug suggesting that gaming has no value? No, they're just having fun with the medium.
To answer directly: I think understanding is the primary value. And I have reasons to hope that my "work" here can ultimately contribute to a better understanding: https://news.ycombinator.com/item?id=49761718. However, there are also other reasons to do things than producing value. You can do things for fun, to show they're possible, to ask meta questions about the field.
Following your analogy it is not a very fair game if it allows winning for those having access to latest exploits. Playing against skiddies who bought the hack kind of ruins it.
I read your comment and my edit was inspired by it. Your article only mentions nameless mathematicians, without really giving credit.
nphardon 1 days ago [-]
My experience has been similar; I find ChatGPT to be much stronger and more precise at math and in communication. I also can not do better with a multiple agent flow than I can with a single agent.
rlue 1 days ago [-]
> Take all the numbers you have so far. Then, “spawn” a new number in every gap between the numbers you already have (crucially, “to the left of all” and “to the right of all” also count as “gaps”). Apply this step forevermore, and you’ll get surreal numbers.
I'm not a mathematician. Can someone explain to me how this approach gets you beyond the rational numbers?
Also, this was formatted as a blockquote, but as far as I can see, this blog post is the only instance of this formulation online.
sebzim4500 21 hours ago [-]
You can have an infinite number of numbers on both sides. E.g. you can ask for a number between 0 and 1,1/2,1/3,... (which gives you an infinitesimal)
doctoboggan 24 hours ago [-]
> a sort of epistemic performance art project.
Agreed, and it's a wonderful piece of art. I look forward to seeing the actual publication and reaction from the math community.
renyicircle 1 days ago [-]
The Claude output in the first one-shot counterexample attempt is hilarious. I hate its writing most of the time but this stuff is next level deep-fried slop.
> And the control column confirms the resonance-necessity conjecture empirically: break the skeleton alignment and the joint kernel dies at the constrained window, exactly as the transversality heuristic predicted.
> The den has air in it.
> Drift fuel exists.
creamyhorror 1 days ago [-]
Absolute bad-metaphor-laden slop, in a dramatic writerly voice. These LLMs are trained on too much pretentious writing.
doctoboggan 23 hours ago [-]
Nit, I know, but I don't think they are trained on too much pretentious writing so much as they are RLHF'd into it because that writing impresses people.
1attice 1 days ago [-]
Discovering that poetry was cognitive compression is alone one of the latent findings LLMs unlocked.
You can so easily imagine this shit being read in a 90s slam poetry coffeehouse. Trust me I was there
djeastm 24 hours ago [-]
>one of the latent findings LLMs unlocked
Either I really misunderstood high school English classes or this is what poems have always been known to do.
fwip 1 days ago [-]
This sort of thing is why STEM students need more humanities classes. The idea that poetry is semantically denser than prose is, like, obvious, to anybody who cares about poetry.
1attice 23 hours ago [-]
Not what I was actually pointing to, but thanks; I'll pass the critique along to my old MA thesis advisor. ;)
simianwords 1 days ago [-]
I never got the point of poetry, even after giving it an honest shot many times. It felt far too compressed to express any thing. I could not separate it from how I would appreciate music. Maybe its because I don't start from the same base as others.
You are right, maybe some humanities classes would've helped or I'm just wired this way.
fwip 24 hours ago [-]
It's pretty similar to music, yeah. Or perhaps music can be similar to poetry. Just this morning I was re-reading the lyrics for "Pale Green Things" & "Woke Up New" by The Mountain Goats. There's a sort of specificity in intention that the form tends to end up dictating, that can really carry a rich emotional context or metaphor more directly into the listener's brain.
(Of course, prose is sort of like a superset of poetry - it's not a hard dividing line. You may find especially poetic lines in your favorite longform works, especially in good fiction books.)
I do think that many English teachers do a poor job of encouraging a love, or even understanding, of poetry. One of my highschool teachers was in love almost with categorizing the forms of poetry - sonnets and sestets and villanelles, and seemed to think that if we learned the rules about what lines had to rhyme or which syllables to stress, we'd enjoy poetry. I hated it! But then the next year, even against this backdrop of learned avoidance, my teacher was so in love with the ability of poetry to communicate - the unfolding of a few words into meaning like a flower in bloom - something clicked. He focused almost not at all on the structure of poetry, but rather on finding works that spoke to you. I'm still not a lover of poetry, even now, but I have poems I appreciate.
nonameiguess 1 days ago [-]
I had a manic friend in college who later became a Emmy-winning television writer. One night over a quarter century ago, he stayed up about 24 hours straight doing nothing but writing, completely free form, some of it prose, some structured rhyme, some of it dialogue with stage directions. He pinned it to his dorm walls like it was wallpaper for a week, then we took it out to the field and burned it. The surviving paper that didn't burn made up bizarre strings of words that sounded much like this, which he retyped and called it poetry. It totally worked.
renyicircle 1 days ago [-]
That makes perfect sense. Using language in unorthodox ways to convey very specific concepts that only make sense to you and sound like bullshit to others.
1attice 1 days ago [-]
I never expected a machine to be better at articulating complex thoughts with a spare number of semantic coordinates, but then I never expected to find out that my poetry was golfing in the latent space
bastawhiz 17 hours ago [-]
These are indisputably good results all things considered. But I have to wonder whether a more scientific and hands-on approach to working on the material would have been better. When I vibe code, I don't just hype the LLM up and tell it to keep going. I interrogate it, I ask it to back up and replace its jargon, and I force it to be accountable. It smells to me like a lot of the circling could have been avoided (even without domain expertise) by just enforcing processes. Even just keeping the Lean more up to date would have likely saved tokens: it doesn't matter if it took longer each week, since the total runtime mostly wasn't the bottleneck.
danabramov 8 hours ago [-]
>I interrogate it, I ask it to back up and replace its jargon, and I force it to be accountable
I do that too when working on software. Here, I did a little bit of that, but it is much harder when I have almost no domain knowledge (aside from understanding the statement of the conjecture), since at each point the LLM might trick me anyway, and it would probably take me a year to understand the concepts enough to tell when what it's saying doesn't make sense.
>Even just keeping the Lean more up to date would have likely saved tokens
That's the conclusion I came to by the final week! But don't underestimate how much Lean needed to be written in the first place to even "catch up" with the reference papers. I had no option to keep it up to date when I started.
23 hours ago [-]
tuesdaynight 1 days ago [-]
Damn, I was not expecting Dan Abramov when I read the title.
FiatLuxDave 1 days ago [-]
This year, LLMs have been involved in a number of interesting proofs of conjectures. But that is not even half of mathematics. Has anyone tried to use an LLM to generate a mathematically interesting conjecture, on the level of Conway's refinement conjecture? If so, what happened?
With all the talk of mathematicians possibly being obsolete, I'm wondering where the future conjectures that future LLMs would prove might come from.
nphardon 24 hours ago [-]
I would not be surprised if an LLM could have generated this conjecture, since it's such a natural extension of the integers, does it hold for surreal integers is not a big leap. It's totally predictable / probabilistic and thats what they do at a high level.
My read on the current sitch is that the drama is more around the humans behind the LLMs, like the OpenAi *people*, who are dishonest, thieving, technocrats, stealing human mathematicians work and claiming credit for it. Since the LLMs are only trained on finite amounts of data on the internet I don't think anyone expects them to replace mathematicians in a serious way. But surely the field is changing dramatically, people will adapt.
cyclopeanutopia 1 days ago [-]
Someone please vibe-prove that ZFC is inconsistent.
vatsachak 1 days ago [-]
That's awesome! Congratulations!
I'd imagine that in three months when we all have access to communicating agent swarms this should be easier
alikatyc 1 days ago [-]
free time spent talking to llm, what an achievement!
j2kun 1 days ago [-]
Perhaps one thing you should devote effort to is ensuring this has not already been proved in the literature.
danabramov 1 days ago [-]
I've confirmed with the mathematicians working in that field that this is a new result.
ianjbutler 1 days ago [-]
Regardless of whether the target result(s) are ultimately correct, isn't it almost guaranteed that supporting infrastructure for surreals-in-lean is a real contribution? Is it a goal to make those polished/reusable, or more like throw-away harness, and just a stepping stone to the proof?
danabramov 1 days ago [-]
I'm a little tired from the project so not eager to jump back into it right away. But yes, I'd love for useful pieces to make their way into https://github.com/vihdzp/combinatorial-games. Violeta, who maintains CG, expressed interest in ultimately integrating the proof in some shape into the repo, but I think more work needs to be done to understand what makes it work.
> Instead, I have a more complicated view, which I actually expressed in my essay The Two Cultures of Mathematics a quarter of a century ago, and which can be summarized by saying that there is a spectrum of attitudes in mathematics to the relationship between problem-solving and conceptual understanding. At one end of the spectrum you have mathematicians who are primarily motivated by the wish to solve problems, who see conceptual understanding as a very important means to that end. At the other you have mathematicians who are primarily motivated by the wish to attain conceptual understanding, who see problem-solving as a very important means to that end.
Before, understanding and problem-solving-ability were so interdependent that distinguishing between the two was practically very difficult and probably wouldn’t have changed anyone’s research agenda. Now, they’re not connected, and this guy just did the ultimate meta-experiment of seriously undertaking a project that is intentionally 100% problem-solving and 0% understanding to prove it (maybe 99% and 1% but pretty close. In his transcripts, he never asks ChatGPT about the math, only about its opinions of the math).
As we (as a society) sit around asking ourselves what mathematicians (and software engineers, and anyone in deep technical fields) should be doing all day, we now have this case study to show us how wide our range of options has become.
omnicognate 1 days ago [-]
> I genuinely invite a refutation.
> So, assuming my proof doesn’t rely on a Lean kernel bug, it’s likely to be legit too.
He lacks the understanding to verify his solution properly, and has to lean on those who do have the understanding to verify it, only being able to say himself that it's "likely" to be correct. (And what do those mathematicians get for laboriously checking the generated proof? 40 grand?)
Seems to me problem solving is as dependent on understanding as ever.
Moverover, the version I linked above is intentionally paranoid so it doesn't use any third-party code except Mathlib. If you allow usage of CombinatorialGames and trust its definitions, the part that needs to be checked narrows down to exactly 20 lines of code: https://github.com/gaearon/conway-refinement/blob/264445c93b...
1 days ago [-]
omnicognate 1 days ago [-]
> If this file is correct and Lean kernel is correct, the proof is correct
There are two ifs in this sentence.
danabramov 1 days ago [-]
What is your point, exactly? Increasing number of people working in and around mathematics are relying on Lean kernel's correctness. That's kind of the point of tools like Lean. Why is it a problem for me to publish a result that relies on it? How do you think other Lean proofs work?
omnicognate 1 days ago [-]
My point is what I said. Without understanding you are only able to say your proof is "likely" to be correct. It's clear from your writing that you understand that your proof will only be accepted once thoroughly reviewed by human mathematicians, who will certainly not be just verifying the definition. Bugs in Lean exist (you're a programmer and it's a program, why would you assume they don't?) and reward hacking and finding bugs are both well established LLM behaviours.
> Why is it a problem for me to publish a result that relies on it?
Bit over-sensitive here. I never said it was a problem for you to publish a result. You can do what you like on your blog and spend your tokens however you choose, just as I'm free to have my own opinions on the value of such an effort. I was responding to, and disputing, a commenter's assertion that understanding and problem-solving ability are "now ... not connected".
danabramov 1 days ago [-]
I see, we don't seem to disagree much.
While Lean is tightening things up after the recent LLM-driven hacks, I agree that bugs are possible. Although usually code that exploits them is obviously aggressive and is deliberately using the more obscure features related to metaprogramming. Also note that my solution has passed the nanoda kernel as well (https://palomar-registry.org/entry?id=PALOMAR-2026-09-03-000...).
That said, again, I never implied that I'm asking mathematicians to "laboriously [check] the generated proof" which is what your parent comment says. The value to mathematicians is knowing that the conjecture is probably right, and knowing the rough path the LLM has taken to it. Instead of checking the Lean proof line by line, what mathematicians are interested in doing (at least, the ones I've been in contact with) is finding a shorter and more direct proof now that they're aware of the outline and main intermediate claims. As for how much value they find in that, I presume they would be able to speak to that when/if they would like to make their research public.
msteffen 1 days ago [-]
> The value to mathematicians is knowing that the conjecture is probably right, and knowing the rough path the LLM has taken to it.
Ah, or is the value to mathematicians that their LLMs can build results on top of this? (In which case, did this do more than save them some tokens?) Or is the value the deep mathematical insight that this result incidentally gives a few mathematicians the confidence to develop for their own personal satisfaction (e.g. if they decide to go and prove it for themselves, and come to the same result after a lot of work)?
(IMO, the deep, scary question: what if it’s soon impossible to make anything at all that anyone who doesn’t know you personally would bother to look at or use? https://www.smbc-comics.com/comic/crack)
GPerson 16 hours ago [-]
I disagree with the other guy. I think you’re a bad person flippantly participating in the destruction of a culture, wasting people’s time.
GPerson 3 hours ago [-]
I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
sethaurus 15 hours ago [-]
Be more specific. Is he a bad person for having an amateur interest in mathematics, for exploring that interest through language models, or for writing a blog post about his experience?
GPerson 14 hours ago [-]
> flippantly participating in the destruction of a culture
dev_dan_2 1 days ago [-]
What is your point? Please don't be obtuse, it is more constructive to make your points clearly.
simianwords 1 days ago [-]
Perhaps mathematicians will undergo the same split as what happened to philosophy and natural sciences.
31276ahq 1 days ago [-]
Yes, the timing of this post just after Gowers' post is fascinating. It is almost as if the marketing machine is well oiled.
danabramov 1 days ago [-]
What marketing machine? You think someone's paying me to do this?
pfdietz 1 days ago [-]
When you descend into conspiracy theorizing to defend your prejudices, it's time to stop and reconsider.
3agha 1 days ago [-]
You can see by who entered the discussion (not you) and immediately sank certain comments that this is a protected submission. One wonders why.
pfdietz 1 days ago [-]
Or perhaps the criticisms are objectively unhinged and are so down voted without having to invent a conspiracy.
GPerson 16 hours ago [-]
This is just an immoral thing to do. If you don’t understand why you should read Terence Tao’s posts about stripmining.
This guy isn’t committed to understanding anything. He’s just screwing around and hoping other people who are turn this into something beneficial to others. He’s just extracting value built up by others over a long period, depleting the finite resource of motivation to work on this topic.
danabramov 8 hours ago [-]
Imagine we discover an alien spaceship, and inside it, a codex. Decyphering the codex tells us a bunch of solutions to alien mathematics, which is remarkably similar to ours and has compatible foundations, but way more convoluted in the actual thought process. The codex would contain proofs of some statements equivalent to open statements today.
Would you, in this situation, be mad at the aliens? Would you say the aliens have "extracted value"? This is kind of how I see this project.
GPerson 5 hours ago [-]
You are not the alien. You are the other human being extracting value from labor and efforts of other humans, with nothing but contempt from them. I am not mad at the computer. What a ridiculous analogy.
danabramov 5 hours ago [-]
What is the value that I am extracting?
The alien in the analogy is the corpus of knowledge that’s newly reachable via LLMs.
GPerson 4 hours ago [-]
No, you’re confused about your own analogy. That corpus of knowledge is the extra mathematics the aliens have. The LLM based AI system and the out of control AI corporations are the aliens. You are some guy inviting the aliens to degrade the possibility of a meaningful existence.
danabramov 4 hours ago [-]
I agree, that's a clearer way to apply it. I guess I see myself more as a guy who, with some effort, managed to scan a few pages of the codex and posted it on the internet. Now, to be consistent, would you say that you'd be equally frustrated at someone posting pages from the alien codex on the internet?
GPerson 3 hours ago [-]
I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
fukaiall 21 hours ago [-]
If this proof is actually valid, this could be a pretty shocking news to the entire academic fields. A software engineer who has never been trained as a professional mathematician, not even having his college degree in numerical field, with pure interest in math, now can solve problems that not even those Fields medalists cannot.
Now I feel like all the intellectual hierarchies and reward systems are broken. Who’s gonna waste his or her fucking time and money in degrees and papers when you just mess around Claude?
makerofthings 1 days ago [-]
Here's my conjecture. Large Language Models are the great filter. They represent a local maximum in the technological advancement of a species from which we will not escape.
zerotolerance 1 days ago [-]
On the other hand here we have an amateur that could accelerate their learning and experimentation faster than ever possible before.
Feathercrown 1 days ago [-]
I don't know if this necessarily qualifies as "accelerating their learning". The user appears to know what the proof is doing, but not how or why.
GPerson 14 hours ago [-]
They accelerated zero learning. Read the blog.
1 days ago [-]
Joel_Mckay 1 days ago [-]
People will just limit publishing valid works to avoid becoming a hapless plagiarism victim class. Same thing happened to tech bloggers ripped off by low-effort you-tube content makers.
Isomorphic plagiarism makes people feel 23% smarter, but it also provably degrades core skills by 17%.
LLM are great at context search, but are also trivially proven degenerative under recursive self improvement scenarios. We look forwards to stripping their assets at a heavy discount.
Also, we shouldn't kink shame peoples cognitive dildo choices. =3
spongebobstoes 1 days ago [-]
this is a great blog post, documenting a very real process of what it's like to create large results with fallible models
though I am an expert at coding, the author's process sounds very similar. constantly double checking, asking for explanations, having AI adversarially check its own work, trying to detect bullshit
nialv7 1 days ago [-]
I don't know why the author could claim this is "their" proof, and they kept saying "they" did this, "they" built that. but in reality everything is done by the LLM and the author is merely asking it to do things. i guess they did contribute money at least...
> Me: btw how’s your mood overall?
LOL. mood??
danabramov 1 days ago [-]
Author here! My impression is that it's customary in the mathematical community to take responsibility for the result with your name, regardless of whether it came from LLM etc (as long as you disclose LLM usage). I am perfectly fine calling it "LLM's proof" or somehow else, but it's "my" in the sense that "if there is a mistake in it, it is my mistake".
GPerson 14 hours ago [-]
Dude you’re so far removed from understanding anything about what the math community thinks. Just stop this nonsense. This is not “your” result.
GPerson 3 hours ago [-]
I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
danabramov 8 hours ago [-]
What do you want me to call it? I'm fine calling it your result if you want.
GPerson 5 hours ago [-]
It’s nobody’s result what is wrong with you?
vends 1 days ago [-]
Someone had to choose the problem, steer the model, and check the output - it's clearly taken a lot of time. That's authorship with a powerful tool, same as it's always been.
nozzlegear 1 days ago [-]
I don't understand your comment. At first it seems like you think the LLM should get credit for the work. But then you mock the author for asking about the LLM's mood, which makes me think you believe the LLM is just a tool and not capable of receiving credit (FWIW I would agree.)
GPerson 14 hours ago [-]
You’re conflating two ways of understanding “credit”. One version is what the tool gets, an honest description that the tool solved it. The other is a concept used in the math community to divvy up job opportunities.
unified101 12 hours ago [-]
Jobs will come and go. I'ts their nature. What makes you treat job over pursuit of knowledge. It is lowly and unbecoming.
GPerson 5 hours ago [-]
There’s not going to be anyone to understand the pursued knowledge when this job gets killed off, which means you’re supporting the reduction of knowledge. That is what’s “unbecoming”.
unified101 4 hours ago [-]
This job based pursuite of knowledge is impure. Drop it.
GPerson 4 hours ago [-]
Do you understand anything about how the career works, or are you just cheering on the death of things you don’t understand?
unified101 3 hours ago [-]
I cheer for a path that opens up the future wider for knowledge.
GPerson 3 hours ago [-]
That knowledge won’t be accesible to people who haven’t trained in understanding mathematics. This is just a fact about the human mind. By killing off this profession you will be making it almost impossible for a large number of people to attain that.
nozzlegear 7 hours ago [-]
I don't think I am.
GPerson 5 hours ago [-]
Probably because you don’t actually know how the culture of mathematics maintains the capacity for humans to learn mathematics.
nialv7 19 hours ago [-]
not mutually exclusive.
author provided nearly no intellectual input into solving the problem, so they IMO don't deserve credit. and it doesn't make sense to anthropomorphize LLMs and talking about their "mood". these are two unrelated statements.
as to if you want to give the credit to the LLM, or if you believe nobody gets the credit, is another separate question.
danabramov 8 hours ago [-]
If there is consensus on how to attribute credit for LLM-solved human-steered (with no mathematical human input) proofs, I am happy to follow that consensus. Do you have a concrete alternative recommendation? What should I change?
danabramov 1 days ago [-]
Just saying (as an author) I don't believe that LLMs have conscious experiences, but the word "mood" was a good languagespace anchor for the kind of information I wanted to get out of the LLM at the time.
1 days ago [-]
jjordan 1 days ago [-]
When you use a drill to put a hole in the wall, do you take credit for it, or do you credit the drill? Without intent, a tool, whether it be a drill or an LLM, is just an inert object.
mattm 1 days ago [-]
They still needed to invest time and other resources into this. It's listed clearly in the 2nd paragraph. Mathematicians, or anyone for that matter, don't figure out everything from scratch. They lean on the work that others have done before them to save time. How is this any different?
GPerson 14 hours ago [-]
[dead]
empath75 1 days ago [-]
LLMs need a lot of help to get to any kind of complicated proof, really. And yes, they get in moods. I spent 3 weeks trying to prove something and frequently had to try and convince Claude that it wasn't impossible and that it could really do it.
cubefox 20 hours ago [-]
It's quite the irony that in the end he says
> Although the current generation of models is trained to complete tasks rather than to enrich our understanding, and today’s AI companies are misaligned with the goals of the mathematical community, I hope that with time we’ll find ways to use these tools in harmony with human research.
while citing "A Severe Misalignment of AI in Mathematics" [1], which condemns exactly the thing he is doing himself: Mindlessly producing theorems without a corresponding human understanding of the underlying proofs.
Perhaps you and I read this document differently. This is not a “famous problem”, I am not a “major AI company”, and I am genuinely interested in increasing mathematical understanding. To the last point, here is a comment from a mathematician who co-authored the paper that my proof is largely built upon: https://news.ycombinator.com/item?id=49761718
The happy case here is that my obtuse proof leads to a concise and illuminating mathematical proof, which is exactly my hope for the endeavour. I think this could then be a positive example of AI/human and amateur/professional collaboration.
What would a positive example look like to you? What do you think the manifesto argues for?
GPerson 15 hours ago [-]
It is a famous problem!!!!!!!!!! You literally understand nothing about the culture you’re stomping all over.
danabramov 10 hours ago [-]
I may be wrong here but my impression is that surreal numbers in general are kind of a niche area that hasn’t enjoyed a ton of interest. Additionally, this specific conjecture did not have any “prizes” attached to it and was not on any list of famous problems I could find. It’s even difficult to Google. What precisely do you mean by it being a famous problem?
GPerson 5 hours ago [-]
It’s a problem that someone could describe in a talk and attribute to a prominent mathematician and which would have resulted in significant career opportunities for having solved in 2021. You don’t understand this because you don’t care that you’re degrading the tiny chance people had to actually learn mathematics.
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
cubefox 3 hours ago [-]
Okay, at least the comment from the mathematician sounds positive.
Positive example: Someone is 1) independently interested in a particular conjecture, 2) he lets an LLM prove and explain it, 3) he ends up fully understanding the LLM proof and is then able to phrase it in his own words.
math_dandy 1 days ago [-]
[dead]
tonetheman 1 days ago [-]
[dead]
31276ahq 1 days ago [-]
[flagged]
manwe150 1 days ago [-]
Several of my coworkers — it’s not that unusual that if you can max out a couple accounts, the companies will obviously notice you (as a high cost customer), and sometimes offer more
1 days ago [-]
dcre 1 days ago [-]
$200 a month is what people pay for morning coffee in the Bay Area.
Retr0id 1 days ago [-]
I don't see why that's weird
memonkey 1 days ago [-]
ah, if it's anyone it'd be dan abramov
1 days ago [-]
GPerson 1 days ago [-]
[flagged]
nbulka 1 days ago [-]
He seems genuinely interested in Surreal numbers. Seems like he found value in exploring them and the Lean process, devoting time and interest to it and understanding what it's like to be a mathematician. I think a lot of people who go into Computer Science may have been mathematicians in the 30s before they became separate majors at the university level.
GPerson 17 hours ago [-]
The blog literally says he had no idea what was going on for the entire time.
dcre 1 days ago [-]
What is stopping anyone from finding the value they would have found before?
AIiscoming 1 days ago [-]
Every human has to go through this in modern times.
I got quite frustrated and disappointed when taking pictures because everyone was doing it and my picture of x was similiar to others taking picture of x.
Either you learn from it and accept that and still do it, or you don't.
But its not new
addlatt 1 days ago [-]
Nihilistic view
GPerson 17 hours ago [-]
It’s not nihilism to see the death of my career.
idjeicjejdjej 1 days ago [-]
Idiotic view, more like.
Let’s not infuse good faith into what is clearly meant as a derogatory comment.
GPerson 17 hours ago [-]
Being derogatory is not incompatible with good faith, idiot.
Computing has historically been a field of wizardry. It's... interesting (?) to see so many people pushing so hard in the direction of sorcery, and in fact applying that sorcery to other fields, in which they themselves aren't quite able to validate whether the spell worked or not.
The usual format that fun mathematics is presented (being talked at by someone who is very well versed in the subject) comes with a heavy cognitive burden - and often I just can't really make it through.
When the author is not an expert the writing is just so much more accessible - it's easier to understand and making it through feels more of an adventure and less of a lecture.
I've never thought previously how much I would enjoy this format though. I'm here to see more amateurs stumbling through mathematics.
Also, wasn't expecting this sort of side-quest from the guy who got me into React.
I love this approachable prose.
If anyone is aware of any other "mathematics for people who don't know mathematics" resources I'd greatly appreciate any links.
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
Knowledge is of 2 kinds: know-that and know-how. Know-that is what LLMs are enabling such as the proof here, while know-how is more useful as that constitutes understanding and puts that knowledge to use.
Primarily I thought of this as a sort of "epistemic performance art project", maybe similar to playing Elden Ring blindfolded having never played it before, or speedrunning a game by opening a box a thousand times and overflowing some counter. It's funny and absurd to do knowledge work without the knowledge.
I think it's also a stress test of meta skills. Like, how much can we do without knowing? What kind of processes can we set up around these demons that would constrain them into our requirements? How can we know when things are going wrong? In some sense, this isn't too different from engineering management.
Naturally, I'm also interested in how much of my role in this could've been automated away. Can there be a skill for that? Then "do a breakthrough" is an irrelevant implementation detail of that skill.
Note that "do a breakthrough" actually produced the worst results over the runs. The best results were from more directed runs like searching for first obstacle towards the next milestone.
The wizards don’t become sourcerers themselves - they become enthusiastic users of someone else’s sourcery. Their years of learning don’t protect them from mistaking access to power for mastery of it.
Sorcerers are the vibe coders of D&D magic, which perfectly describes how wizards in-universe feel about them.
The net output of math will increase, and mathematicians have more work now to unravel all this, and make it useful. AI plays the role of a monkey in the infinite monkey theorem [1]. We now need an LLM corollary - Something like: A finite number of LLM agents will almost surely find all theorems given an infinite token budget.
[1] https://en.wikipedia.org/wiki/Infinite_monkey_theorem
For any finite program (eg some LLMs), there is a true math theorem which they cannot prove or disprove (given fixed input of the statement with no other information sources). If that weren’t true, BB would be computable.
Math is beyond computation. Since AI is just bits in bits out, it has this fundamental limitation.
Any magic of AI systems comes from the transformed meaning of its input data. With fixed weights any LLM is just an artifact. For example a human prompting an LLM constitutes an extra information source, which removes the above limitations. In theory any input from the natural world would remove the limitations too. The natural world is a black box and we don't know what kind of meaning or intelligence could underly it.
We are talking about the same thing, but I would actually put this the other way around.
Computation and computability is "the final frontier". Math is a "subset" of that. Doesn't matter if we choose ZFC or in the future discover some "better" subset of core axioms, we will always hit limits where BB will trivially skip over whatever we could prove (let alone Gödel's theorems).
> given fixed input of the statement with no other information sources
Also, this is just trivially avoidable, so not sure if we really should be concerned about this limitation. An LLM in a loop where it can write on a tape can be Turing complete, ergo it can compute anything computable and is "bigger" than math at that point.
Maybe I'm misunderstanding you point, but I don't know how widely this would be held as true. Are you defining "math" as _only_ what can be proven under some particular formal system?
For math I don't have a fix definition, but it's surely a bit more specific than that (e.g. I wouldn't consider the computation that prints a 0 at the same place for infinity math) - but of course I do see the circularity in my argument: a Turing machine is a mathematical object in and of itself. Though being able to talk about something doesn't necessarily change which is "bigger".
As for the other direction, this gets a bit more into the philosophy behind math itself. Constructive math's territory is "easy" - but I am on the opinion that if humans (or any intelligent physical entity) are at most Turing-complete [1], then any non-constructive math "steps" or thoughts must also be at most computable. Well, unfortunately I can't prove whether math done by transcendent entities are also computable, though.
In any case, I am no mathematician, so whatever I think regarding this topic may not have much relevance to anyone, only done CS course with quite a bit of math, but that's obviously not the same.
[1] I believe religion is an escape hatch here from an argument perspective
This is a bit of a strange assumption to make. I do agree that a human, if it had infinite memory, would be an universal machine, i.e. capable of computing any given Turing machine [0]. But would that be the limits of its capabilities? It's far from certain.
You'll get into the philosophy of free will (funnily enough, a sort of inverted Turing test), i.e. for a given human with infinite memory, is there a Turing machine that exactly replicates the behavior of that human? Is our behavior governed entirely by rules? Would that imply that a human themselves is a kind of Chinese room [1]?
> any non-constructive math "steps" or thoughts must also be at most computable.
What does it mean for a "thought" to be computable? Compare to Gödel's incompleteness theorem. Clearly the act of stating the thought, or writing down the theorem, is computable. But proving it to be true or false may very well be impossible.
[0] https://en.wikipedia.org/wiki/Universal_Turing_machine [1] https://en.wikipedia.org/wiki/Chinese_room
Well, given our scientific knowledge it's a molecule-level (only important to disregard quantum physics to make the case easier) physical/chemical process, that we should in principle be able to simulate on any other medium, including a Turing machine.
Nonetheless, I can accept the definition of math where it's about "truths" and truths can obviously exist without being computable.
Do you agree that humans are physical systems?
My understanding is that any physical system can be evaluated to any degree of accuracy by a computer, no?
That's an interesting hypothesis, but I don't know why you'd assume it to be true at face value. It's a bit unclear how you would even define "evaluated", given that we don't yet have a mathematical model of all of physics as we know it. [0] And then consider unknown unknowns.
> Do you agree that humans are physical systems?
Do you consider humans _with infinite memory_ as physical systems? Do you consider computers _with infinite memory_ as physical systems?
[0] https://en.wikipedia.org/wiki/Physics_beyond_the_Standard_Mo...
As for the infinite memory, Turing machines have this nice property that they can only visit a finite amount of memory after finite steps, no matter what. A Turing machine running for a finite time (we got this) will surely use a finite space, so being "a bit short" on infinite space is not a problem, I believe.
Yes... but that's in the area of the big bang and black holes. My understanding is that the chemistry of the brain is very well modeled.
So, unless we find unknown physics, and unless that physics behaves differently than every other known physics, humans are computable?
Do I have that right?
Now consider throwing a ball in the air. Can you even write down the rules that each of the ball's subatomic particles obeys? How can you be certain there exists a computer program that in finite time can predict where any of the particles, for any ball, ends up?
> the chemistry of the brain is very well modeled
There are models, but the fact of those models is that they do not apply to "any degree of accuracy", as you claim.
Consider the ball thrown in the air again. Is the ball affected by what happened 100 years ago, inside of a black hole 100 light years away? Why would it not be affected by that? Or if you grant that it is affected by that, do we then need a model to predict those effects before we can "evaluate" them?
EDIT regarding the below linked blog post: Did you read the rest of my comment? Did you even read the blog post you linked to?
> We certainly don’t have anything close to a complete understanding of how the basic laws actually play out in the real world — we don’t understand high-temperature superconductivity, or for that matter human consciousness
Can you try to consider my central point before replying: Are the rules governing physical reality simpler or more complex than the Collatz conjecture? Does there exist a (theoretical) computer that can "evaluate the Collatz conjecture to any degree of accuracy"?
EDIT 2: I'm not the one moving goalposts. On what grounds are you classifying the question whether a given number ends at 1 or not for the Collaz conjecture as an "inifite" computation? It's a simple boolean question, yes or no. All you have to do is build a computer that can answer yes or no for each integer. Isn't that simpler than answering the position of each atom in the ball after the throw? Each is just a function, what makes one more infinite than the other?
Also, regarding determinism, just read this article by the same guy you linked: https://preposterousuniverse.com/blog/2011/12/05/on-determin...
> For everyday-life purposes, we can’t get around the fact that quantum mechanics makes it impossible to predict the future robustly.
Are you sure?
https://preposterousuniverse.com/blog/2010/09/23/the-laws-un...
You've moved the goalposts again. I said simulate, not predict. It is possible to simulate the entire Schrödinger wavefunction.
And PLEASE - just use the reply button. It is impossible to track every time you edit your comment.
> ANY physical system can be evaluated to ANY DEGREE of accuracy by a computer
It's completely SENSELESS to claim that they are distinct, because in order to EVALUATE or SIMULATE the physical system you will need a FUNCTION which COMPUTES the STATE of the system at a given point in time. The only POSSIBLE distinction between SIMULATING and PREDICTING would be the time taken for the computation, but that is COMPLETELY IRRELEVANT as long as it is finite.
Again, your own source says:
> We CERTAINLY don’t have ANYTHING CLOSE to a complete UNDERSTANDING of how the basic laws actually play out in the real world
How does that square with your claim above?
You are entirely correct. I was sloppy in my first comment. I should have said simulate. My sincere apologies if that's been the crux of our dispute.
> The only POSSIBLE distinction between SIMULATING and PREDICTING would be the time taken for the computation
No. The distinction is in determining which "you" is you. When simulating the wavefunction, every you is simulated.
> We CERTAINLY don’t have ANYTHING CLOSE to a complete UNDERSTANDING of how the basic laws actually play out in the real world
It's very understandable if you include his following sentence:
> But these are manifestations of the underlying laws, not signs that our understanding of the laws are incomplete
He's saying we don't understand emergent behavior produced by the laws - not that the laws themselves are incomplete. E.g. we don't know how/why a bag of neurons turns into a person.
And - you should reply to my comments rather than edit your old ones.
In what sense? BB(n) is a prime example of an object that can be mathematically defined, yet is not computable. Or see BBB(n) for an "even more" uncomputable function. [0]
> An LLM in a loop where it can write on a tape can be Turing complete
What does this mean? A given LLM, like a given C program, can't really be Turing complete or not in a meaningful sense. The C programming language, or the concept of LLMs in general can be said to be Turning complete or not. Do you mean to state that LLMs in general are not Turing complete, but being "in a loop" somehow makes a difference?
> it can compute anything computable and is "bigger" than math at that point
Again, in what sense is it "bigger" than math? Lots of things are Turing complete, I wouldn't classify lambda calculus as "bigger" than math.
[0] https://wiki.bbchallenge.org/wiki/Beeping_Busy_Beaver
LLMs use RNG for sampling, so they are not pure computers.
IIUC, Godel's incompleteness is less about theorems and more about axiomatic systems. Given an axiomatic system, there are statements within it which cannot be proven or disproven. It's relatively unrelated to the platonic ideal of the theorem itself. The statements it considers are axiomatic-system-specific.
Another way to view it is, who cares if we can't prove or disprove "This statement is false". Ok, the axiomatic system is incomplete; fine. What's important is can the system prove a real theorem that I care about.
The busy beaver computability argument addresses these issues. The problem format is always "For Turing machine T with no input, does T halt?". This format can encode many math problems. And we know already that BB(432) is independent of ZF, aka, there is a 432-state TMs which ZF can't prove or disprove the halting behaviour of.
So BB looks at real theorems, ranks them, and we can ask what axiomatic systems can solve them or not. Godel looks at 1 axiomatic system and produces a toy theorem which the system can't solve. That's an extremely important difference!
The core issue is that any fixed LLM can only encode so many axiomatic systems in its states, and the fixed systems implies an upper bound in terms of the BB number which it can solve. Godel is only looking at one system at a time, while BB is a way to use a common problem format to rank every axiomatic system on an infinite number line.
I'm no mathematician, but this is also the crux of Gödel's theorem, he just showed it in a more "hacky" and clever way - but BB(432)'s relation to ZF is also a consequence of Gödel's more general idea, is it not?
-----
Oh, and: All proofs are conditional on axioms. If those axioms are computably enumerable, then all of their consequences are computably enumerable too.
Most propositions may be ugly and irrelevant, but how do you know how many are not so and we just can't prove it? Also, what about stuff like Continuum Hypothesis, would you add it or not?
"Given infinite thinking time a finite number of humans will solve all theorems"
I also love the angle that this was not intelligence just brute force. As if the mathematicians didn't reeaaally want to solve this they were just too lazy to give it a good try.
What does AI have to actually do before you realize these things are actually smart?
Today it cost massive effort but it's possible 10-20yrs from now an AI could solve a problem like this in under an hour with a single thread on a free subscription paid for by serving an ad.
These arguments are so weak because you'll then have to make the same one a few years from now when it does something else impossible. The argument only stands if we assume no progress will occur.
I just think you and the other guy have different definitions of "smart". There's no denying that LLMs are useful, but I don't know if I'd classify them as "smart". There were probably people in the 80s saying computers were "smart" because they could compute 78971 * 12341 faster than a human.
In what sense is a LLM "undeniably smart" but a CPU from the 80s isn't? Or would you define such a CPU as "smart"?
You're right though they largely are "smart" in the 80's computer sense. This is largely due to continual learning being unsolved.
BUT the more you look at them, research, and try experiments there's something there not in a 80s computer. If I had to guess maybe 1-5% of a humans ability but it's there. They are able to do novel things but ever step outside of their distribution takes exponential effort for every small addition. There is a true ability to adapt and learn new things on the fly, things never seen before. That is the the smart part. There something hidden in these things we don't understand that allows novel insights built from in context learning.
It's actually measurable in experimental settings but even there it's hard to tease out. I saw it mostly while doing CL training experiments. But I also see it while working with them for coding novel things.
But the more power we provide and farther down the road of this we go those 1-5% are things like solving unsolved math problems. No human solved these things. You say brute force, I say it needed massive effort to break out of it's distribution and get those small insights. It's very human like when taken at scale. The scary thing is that scale is getting smaller every day.
And our brains must something like that at some physical level. You can't have a "turtles all the way down" of reasoning - the building blocks must be simpler. It must reduce to something like pathfinding and brute force at some point, weighted by factors in the system and maybe some randomness.
For undecidable theories, there is no computable function bounding this blowup from theorem length to proof length (otherwise, the theory would be decidable.)
This is a cool blog post and I think you're going the right way, and beginning to get an understanding of the proof as you go.
I'd recommend continuing on the simplification and understanding route, until you yourself can follow the proof. Some suggestions, as I did something similar:
1. See if (or ask the AIs) if individual parts of the proof can be found elsewhere, i.e., is an argument just a copy of something else? If so, it's important to attribute this, but also this usually allows simplification ("by Theorem X", etc.)
2. Look for redundant patterns and try to combine them.
3. Ask the AI to be a critical reviewer from some journal, and try to fix its criticisms.
4. Continue simplifying! Assume that the final result may actually be relatively short.
Good luck!
As with most interesting proofs, the number of core ideas is actually small, I'd say two for the real exponents, and presumably a third idea for lifting up to omnific integers. I have been redoing the real exponents part of the proof going on the ideas only, and with a few smarter choices, I am converging on something very short. And I mean very short, which is amazing. I didn't think the answer would be this close: it 'just' needs looking at the problem from the right angle, and also make a fairly bold guess at the outcome.
Dan's current proof is of course much longer. Between the fossilized ideas that Dan mentions in the post and the formalisation of previous results, there's a lot of cruft that inflates the proof but does not really help understanding what is going on. Luckily the word 'derivation' pops up early, otherwise it would have been very challenging to wade through the lemmas to find the important points.
The problem is rather how quickly we can change our ways of working to make sure the training and hiring pipeline does not collapse. That's the disastrous scenario, for both the individuals affected and the discipline, that we must avert somehow. I wish we had an easy answer to that. I certainly don't. But I like to think that at least engaging with the public in a constructive way will have a net positive effect.
My hope is also that Mantova and very possibly L'Innocente will get a substantial share of credit for their role in the resolution of this conjecture by Conway: the AI would not have embarked on this were it not for their prior work. So even human mathematicians with an inclination for more exploratory "problem solving" will have plenty to do in the future. (The story is actually not that different for the recent Navier-Stokes forced blowup result, which also built on key conceptual work from 2023 by Córdoba and Martinez-Zoroa.)
You have nothing more to make of my comment because you don’t care to consider the actual problem here.
The math career is going to killed off 2-3 years before law and medicine and Wall Street banker. How does this make me more secure in any way?
This is indeed also coming for law, medicine, and banking, though licensed professions will hold out for longer because you need someone to put in jail when things go wrong. The problem is that all this is extremely over politicized and nobody is able to think clearly. They want to simultaneously say all this is just hype and a bubble and will go away like NFTs did, and also are starting to worry about economic replacement. Some more coherent political narrative will have to be formed.
Regarding being jobless for the 3 intervening years, it is certainly a personal concern but in this temporary phase there are still some other jobs for smart people. Once there aren't any, we are entering the part that I was talking about where you will be far from alone and you can join together to exert some kind of political pressure but it will not be about math PhDs, but employment as a whole. And it may not be very effective if AI is on the other side, not on yours. Yeah, it sounds like scifi, and people want to dismiss scifi concerns and instead focus just one inch ahead of their toes, instead of seeing the writing on the wall.
> people spending their lives trying to promote humanity
What does this even mean? Become a monk? They are credited with helping humanity.
Ahaha. Flowers are the way to go.
You still prefer to get cheaper options yourself. I know this argument gets caricatured in the "yet you participate in society" meme, but the point is that this is the aggregate result of individual humans making decisions on where to allocate their resources. You can attack this using various ideological and religious frameworks, but if it's just some stoner college freshman's communism, I'm not interested (neither if it's the more potent version that dispossessed my ancestors in Eastern Europe).
Air your LLM greviences someplace else.
And of course, going by their own admission, they couldn't "have done this themselves": the most you can argue wrt. this is that Mantova and L'Innocente, or some other narrow domain experts, might have done this themselves and that AI "scooped" this result from them.
Yes, I was not doing any mathematical work in curating the output, but the article makes it quite clear that pivots and constraints the LLM would not impose on itself were critical to actually making progress.
Also:
>He doesn’t even know any mathematics and has never cared to learn
While I don't know enough mathematics to work on this problem, claiming something like this is preposterous. As I link in the first paragraph of the article, I've been learning mathematics on my own by going through Terence Tao's Analysis book and solving exercises. I'm familiar with the concepts of mathematical definitions, proofs, etc. I've gotten about halfway through the book solving them on paper before abandoning it (and later got through the first few chapters in Lean, also solving every exercise — by hand, mind you). Sure, this doesn't make me a mathematician, but I'm closer to a dropout first-year student than to someone who has "never cared to learn".
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week. I won’t hide this comment though it is shameful.
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
It might be possible to plainly continue-shot it with more powerful models in the future. I agree that whatever I did is probably automatable.
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
I am not saying that my work constitutes a mathematical contribution on its own. Not any more than stumbling upon an anonymous manuscript with the solution would constitute a mathematical contribution. I do, however, think that it can lead to a mathematical contribution if any mathematicians consider it worthwhile to do something with it. Whether or not they consider it worthwhile is not up to me.
It is a distinction without a difference because I want to live in a world where people get to fill their lives with meaningful things, and are not forced into Uber delivery driving jobs just because rich people like you think it’s fun to put their name next to something other people made prestigious.
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
Re: "rich", I've essentially spent $400 on this (in subsidized subscriptions), plus my free time being a mindless drone. Given that you assume my role is automatable, it sounds like this is relatively accessible to anyone with $400 (as long as AI companies continue subsidizing the frontier models). I don't think I've had some kind of an unfair advantage beyond that. If anything, a proper mathematician would probably be able to derive the result much faster with the same tools.
I don't know how the broad availability of these tools (to mathematicians and non-mathematicians alike) will change the field, what is considered prestigious, what work gets funding, how it affects the pipeline, etc. You seem to be implying that even testing the limits of these tools, or at least publishing the results obtained with them, is unethical in itself, even though it is broadly accessible now. I can understand this point of view.
[0]https://www.google.com/search?q=video+introduction+to+surrea...
I think this project is really neat, but is it appropriate to cold email specialists before you've put in enough hours of effort to describe yourself as more than an "amateur"? OP's emails may have been helpful, but billions of people use these LLMs to wade into new areas and email is already low signal-to-noise.
Upshot, if there are real errors in published work, I think most mathematicians want to know about them.
It doesn't matter who found the error nor how it was found. An error is an error.
Now here there was no fraud just genuine error, but it will annoy people nonetheless and scrape their ego that someone uninitiated can just type some stuff in a magic box and conclude that they, the established published, tenured mathematician with awards and medals can be wrong.
> However, I didn’t just want any result; I wanted something that pulls me.
> Initially, I asked Claude:
> Me: which unsolved problems in the Surreal Numbers research program pull you the most and why?
Note the switch from "pulls me" to "pull[s] you". What is the author's perception of the relationship/boundary between them and the LLM here?
1. Are they using it to find things it flags as interesting in hopes they might also find it interesting?
2. Do they consider "interesting" to be a universal (observer-independent) trait and are using the LLM to find things that are interesting?
3. Have they delegated their desire to find something interesting to the LLM so that it can instead find something that it flags as interesting, regardless of how the author feels?
4. Do they see it as a part of their thought process, and so do not distinguish "you" from "me"?
5. Do they see it as part of them, and are referring to the combined entity in the second person?
I would love clarification on this.
Let me first clarify my relationship with mathematics. I think of myself as "an awestruck observer from a distance". I find some parts that I understand beautiful, and I have also tried to understand some of the basics rigorously. However, I generally just can't make my way through any serious paper, as I both lack the prerequisites and struggle with the amount of inference mathematics tends to place on the reader. That's the "from a distance" part.
Now, about picking the problem. I am genuinely "pulled by" surreal numbers themselves. I find them irresistibly beautiful. There is also a bit of bitterness around how they haven't fulfilled their promise (yet?) as Conway hoped they would be able to become a better foundation for some mathematics. But they are a bit too difficult to prove things about so far, and we know too little about them. So what "pulls me" also is a possibility of making enough dents in this that we would be able to use them more broadly, and learn even more things about them.
However, I do not know the details of the latest research. I don't know which problems have actually been solved, which pursue Conway's original vision vs narrower approaches, and which are elegant enough to feel "awestruck" enough about. So this is an invitation from me to LLM to share what it "feels pulled by" (for whatever definition; I think of it as just navigating the languagespace) , and then sifting through that list to see if something it lists makes me feel something. I would assume that with the field currently being so small (serious mathematicians mostly don't care about surreals), it's easy to get the LLM "excited" (again, just a vector in the languagespace) enough that it would give me genuinely interesting candidates. Then it's up to me to sift through them and see if they "speak" to me.
It's like asking a mathrock nerd to share their favorite mathrock albums. Niche enough that you'd likely get good results. Then you can listen and form an opinion.
In this particular example, the "ONAG birthday" and "maybe last Conway's unsolved conjecture about surreals" part spoke to me emotionally, the statement itself amazed me with its simplicity, and I felt "blood in the water" related to the recent results bringing the conjecture closer. So I felt the pull myself and went with it.
Record scratch
Got lost here. I think I'm officially too dumb for math.
You can think of the "surreal numbers" as being built up step by step. We start out with no numbers at all, and then we repeatedly do a construction that makes some new numbers.
A surreal number is made from two sets of (pre-existing) surreal numbers. We typically call them L and R, for "left" and "right", and sometimes write it as L|R or {L|R} or something like that. The "left" numbers have to be smaller than the "right" numbers. The resulting number will turn out to be, in a certain sense, the "simplest" number in between all the left numbers and all the right numbers.
Now, as I said, we start out with no numbers at all. It might seem like that gives us no way to proceed, but it does: even given no numbers at all, we can still make a set of numbers, namely the empty set! So we can use that for both L and R, getting ∅|∅. Empty sets on both sides. We call this 0, and it will turn out to behave in the way you'd expect the number 0 to behave.
Now we suddenly have another set available, namely {0}, the set containing only zero. Which means that instead of being able to make one number, maybe we can make four: ∅|∅, ∅|{0}, {0}|∅, {0}|{0}. The first of these we already knew about. The last isn't actually admissible -- remember that the "left" numbers have to be smaller than the "right" numbers, which is "vacuously" true when one of those sets is empty (it means "if you have a number x in the left set, and a number y in the right set, then x<y", and if there are no numbers in the left set or no numbers in the right set then that's trivially true) but isn't true when both sets contain 0 because 0<0 is false.
So actually we get two new numbers: ∅|{0} and {0}|∅. The first fits into what OP calls the gap "between nothing and zero". The second first into what OP calls "the gap between zero and nothing". In both cases, "zero" means a number and "nothing" means a space where we don't yet have any numbers.
The number ∅|{0} is called -1 (it has to lie to the left of 0, and there's no constraint on its left, and -1 is "the simplest number less than 0") and the number {0}|∅ is called +1 (it has to lie to the right of 0, and there's no constraint on its right, and +1 is "the simplest number greater than 0").
I should explicitly acknowledge that I haven't defined what "less than" and "greater than" actually mean for these numbers, nor anything else about how they relate to one another that could possibly justify giving these things the specific names 0, -1, and +1. But there are definitions for "less than" and "greater than" and "plus" and "minus" and so forth, and the whole thing does turn out to work very nicely.
Anyway, once we've got these numbers we have eight possible sets that can go on the left or on the right. The requirement for left-things to be smaller than right-things reduces the possibilities somewhat, and the actual new numbers we get next time around are: ∅|{-1}, which turns out to be -2; {-1}|{0} which turns out to be -1/2; {0}|{+1} which turns out to be +1/2; {+1}|∅ which turns out to be +2. We also get some already-existing numbers in new ways; for instance, {-1}|{+1} is actually equal to 0 ("0 is the simplest number between -1 and +1"). Again, I should explicitly acknowlege that I haven't said anything about how you determine when two of these things are actually equal; again, it does all turn out to work properly.
If you keep going with this construction, you produce all the integers, two at a time, and also all the "dyadic rationals", meaning fractions where the denominator is a power of 2. And then, once you've got all those, at the next stage of construction you abruptly get all the real numbers -- e.g., the square root of 2 is L|R where L = {dyadic rational numbers that are negative or have a square smaller than 2} and R = {dyadic rational numbers that are positive and have a square larger than 2} -- and you also get {0,1,2,3,4,...}|∅, conventionally written as a lower-case Greek letter omega, which is an infinite number, larger than all the integers. (And its negation.) And {0}|{1,1/2,1/3,1/4,...} which is an infinitesimal number, positive but smaller than any ratio of positive integers. And you can then proceed further and construct a vast infinitude of numbers, including all the real numbers (which we've already made) and all of the so-called infinite ordinals (which you can kinda think of as being a sort of "infinite positive integer", though there's more to them than that) and much more, all in a system that lets you do arithmetic and suchlike. It's very elegant, if your brain has been twisted into the mathematician-y shape that finds such things elegant.
For the infinitesimal number, I think it makes more sense to use {0}|{1,1/2,1/4,1/8,...} since it gets born at the same day as say 1/3. So it is easier to understand how it arises without "waiting" for all reals.
First: the construction of the real numbers from (traditionally) the rational numbers by means of "Dedekind cuts" (sometimes called "Dedekind sections"). The idea is that if you're trying to build up the machinery of mathematics from scratch, it's not too hard to go step by step from (say) sets to nonnegative integers to integers to rational numbers, but it's harder to get from there to the real numbers, and Dedekind's idea is to say that e.g. the square root of 2 is the way of chopping the rational numbers into "things less than the square root of 2" and "things greater than the square root of 2".
Second: the construction of the ordinal numbers (a sort of generalization of the notion of "nonnegative integer" that allows the numbers to get very infinite) due to von Neumann: you start off saying that zero "is" the empty set, and then you repeatedly say: the next ordinal "is" the set of all the ordinals you've constructed so far. So, e.g., 1 = {0}, and then 2 = {0,1}, etc. -- but once you've constructed all the nonnegative integers you can then look at {0,1,2,...} and that's a new ordinal typically called ω, and then you can take {0,1,2,...,ω} and call it ω+1, and so on and so forth.
Both of these are special cases of what Conway does: Dedekind's is the case where all the numbers are rational numbers and you don't allow either set to be empty, and von Neumann's is where you _require_ the right-hand set to be empty.
There's a further connection, which I believe is how Conway found these things in the first place: if in the definition of surreal numbers you delete the requirement that everything in L has to be less than everything in R, then what you've got is (more or less) the definition of a position in a two-player game. L is the set of positions one player can move to, R is the set of positions the other player can move to. (I say "more or less" because e.g. in many games you're allowed to repeat positions, and games may have complicated winning conditions or involve chance or whatever.) And there's a whole rather nice thing called "combinatorial game theory" that's all about these, and from that perspective numbers are just one particular kind of (position in a) game. (Specifically, a number is a game in which at no point in the subsequent gameplay can it ever make your position better for you to make a move: you'd always rather pass if you could.)
https://en.wikipedia.org/wiki/Winning_Ways_for_Your_Mathemat...
https://web.archive.org/web/20230307045844/https://www.cs.st...
The numbers don't matter and you could replace -1 and 1 with anything. It's just easier to begin your new fake number at - 1 and 1. Because position does matter.
Basically what I take away is that we're inventing a new number system from scratch. So we're not "proving" that 1 is a number between 0 and the empty set. We're defining it as such, and it just so happens that a number system defined this way works out in convergent ways with other mathematics.
Is that roughly right?
Only as a mental abstraction that's based on our experience/concept of space+time.
Sorry it was confusing.
Edit: the picture is now edited into the article.
Some other comments clarify that "nothing" is more accurately "the empty set". This is helpful because at first I wrongly synonomized "nothing" with zero. But now I get tripped up on the "between" language. Maybe it's a lack of background in sets, but I don't know what "between" implies for an integer (zero) and a set (the empty set).
My favorite intro to surreal numbers is https://www.infinitelymore.xyz/p/surreal-numbers, but it is behind a registration wall.
That's a mistake. They should've written "empty set of surreal numbers" and not "nothing".
It is a constructive theory, like sets/ordinals. For ordinals you can use ∅, { }, ∪ and you construct
For surreal numbers you use the form { A | B } where A and B are sets of surreal numbers. Some restrictions apply so not all of these forms will be surreal numbers.You build up surreal numbers as
and so on.---
Edit: A happy accident: I denoted the "empty set of surreal numbers" with the symbol ∅. It works, and it gives you the surreal numbers. But if you think of ∅ as the empty set in set theory, then the same construction (using an ordered-pair construction) gives you the surreal numbers as sets!
I’m slightly stretching the metaphor here because the number lines gives me enough structure (order expressed visually) and I only deal with at most one set item at a time (since we construct in the order of simplicity and can use the already constructed numbers), so it (IMO) unnecessarily complicates things to even talk about sets when we’re sort of just making cuts on the line. But in either case I don’t see the problem with colloquially saying “nothing” here.
> (crucially, “to the left of all” and “to the right of all” also count as “gaps”)
So there are two "nothings" here, left of "all" - i.e. the zero - and right of it.
Though I'm not quite sure how you'd get infinite or irrational numbers by this procedure. Wouldn't you simply get the rational numbers by this?
(unless the "put a number" step is doing more work here than it seems. He doesn't really say which number to put there. In the examples, he mostly did "new number = (left number + right number) / 2", with special cases if any number is "nothing" - but he never actually wrote what the rules are here.)
It’s a terrible explanation. A surreal number is defined as a pair of sets of surreal numbers (where you fiddle around the recursion in that definition by defining them in waves, so strictly speaking you’re defining “the surreal numbers born at time T” for each individual T given access to the surreal numbers born at all earlier times, and then you “take the union across all times”, scare quotes because there are too many times for this to result in a set). Zero is a surreal number but the LLM is using the word “zero” to mean “the set containing just the surreal number 0”; “nothing” here is the LLM’s obtuse word for the empty set. Wikipedia may actually be easier to follow.
> We may say that Cantor was only interested in moving ever rightwards, whereas Dedekind stopped to fill in the gaps, so that R was always empty for Cantor, never empty for Dedekind. It is remarkable that by dropping these restrictions we obtain a theory that is both more general and more easy to work with.
This is precisely the intuition I present to the reader of the article. I am relying on visual aid (concretely, the ordered number line) to imply the machinery explicit in the actual recursive definition. The intended reader of this article is not a mathematician, and I think intuition is vastly more important here.
And I don't think I'm conflating 0 with {0} as you claim. When I say zero is "between nothing and nothing", I mean 0 := {|}. When I say one is "between zero and nothing", I mean 1 := {0|}. When I say 1/2 is "between 0 and 1", I mean 1/2 := {0|1}. And so on. I elide "the simplest number" because I am already going in the order of simplicity. I do not need to explain that alternative spellings like 1/2 = {0.2 | 1} are valid because it is not relevant to establishing the mental model of birthdays.
For the finite cases in my explanation, I do not need to explain that the left and the right parts form sets because I only ever need at most one surreal on either side to define the next generation. I also do not need to state the left/right order condition because it is already visually implied by the picture. For the same reason, I do not need to explicitly quantify over the set of earlier-born surreals, since in these finite cases, if we go birthday by birthday, each next day's surreals are definable via the numbers already constructed by the previous day.
I agree that these finite examples don't spell out how to handle infinitely many bounds at the omega-th day, which is where I believe the illustration embedded below is more helpful. I still think "a gap beyond 0, 1, 2, 3, ... with nothing on the right" is a useful intuition when we get there.
For a more precise but accessible treatment, I think https://www.infinitelymore.xyz/p/surreal-numbers is much clearer than Wikipedia.
“Doing better than a totally useless explanation in fewer characters” is in general impossible, of course, eg if the first explanation has only one character.
I'm hoping someone develops an interactive tutor that can teach any subject to any depth.
The tutor should optimize its pedagogy. It should use online RL to adapt to a learner's ideal learning style, model what the student understands and to what degree, and understand what the gaps and next steps are.
I'd subscribe in a heartbeat.
- https://github.com/mattpocock/skills/blob/main/skills/produc...
- https://www.youtube.com/watch?v=s5T5oQJcJ6U
ie. i am an expert at zig, explain this c++ in terms of zig
Universities are great for networking, starting projects with other students (not the ones professors mandate), and learning lab sciences. In research, they're great for institutional knowledge, having a community of peers, getting guidance from research advisors, having real equipment and funding, etc. But there's a great need for AI tools to accelerate learning outside of that setting.
Anecdotally, I'm a working adult. I'm not going to waste time in college again. I need this for me.
In each episode they make a major science-fiction style breakthrough and grapple with the consequences without revealing themselves.
and "proof map": https://gaearon.github.io/conway-refinement/#/map/conway-ref...
[1] https://eps.leeds.ac.uk/maths/staff/4058/dr-vincenzo-l-manto...
Edit: Another depressing fact is that he generously paid to LLM megacorps while piggy backing on human help for free and in the end calls the proof his or LLM's.
To answer directly: I think understanding is the primary value. And I have reasons to hope that my "work" here can ultimately contribute to a better understanding: https://news.ycombinator.com/item?id=49761718. However, there are also other reasons to do things than producing value. You can do things for fun, to show they're possible, to ask meta questions about the field.
Regarding calling the proof "mine", I've replied to that here: https://news.ycombinator.com/item?id=49755885.
I read your comment and my edit was inspired by it. Your article only mentions nameless mathematicians, without really giving credit.
I'm not a mathematician. Can someone explain to me how this approach gets you beyond the rational numbers?
Also, this was formatted as a blockquote, but as far as I can see, this blog post is the only instance of this formulation online.
Agreed, and it's a wonderful piece of art. I look forward to seeing the actual publication and reaction from the math community.
> And the control column confirms the resonance-necessity conjecture empirically: break the skeleton alignment and the joint kernel dies at the constrained window, exactly as the transversality heuristic predicted.
> The den has air in it.
> Drift fuel exists.
You can so easily imagine this shit being read in a 90s slam poetry coffeehouse. Trust me I was there
Either I really misunderstood high school English classes or this is what poems have always been known to do.
You are right, maybe some humanities classes would've helped or I'm just wired this way.
(Of course, prose is sort of like a superset of poetry - it's not a hard dividing line. You may find especially poetic lines in your favorite longform works, especially in good fiction books.)
I do think that many English teachers do a poor job of encouraging a love, or even understanding, of poetry. One of my highschool teachers was in love almost with categorizing the forms of poetry - sonnets and sestets and villanelles, and seemed to think that if we learned the rules about what lines had to rhyme or which syllables to stress, we'd enjoy poetry. I hated it! But then the next year, even against this backdrop of learned avoidance, my teacher was so in love with the ability of poetry to communicate - the unfolding of a few words into meaning like a flower in bloom - something clicked. He focused almost not at all on the structure of poetry, but rather on finding works that spoke to you. I'm still not a lover of poetry, even now, but I have poems I appreciate.
I do that too when working on software. Here, I did a little bit of that, but it is much harder when I have almost no domain knowledge (aside from understanding the statement of the conjecture), since at each point the LLM might trick me anyway, and it would probably take me a year to understand the concepts enough to tell when what it's saying doesn't make sense.
>Even just keeping the Lean more up to date would have likely saved tokens
That's the conclusion I came to by the final week! But don't underestimate how much Lean needed to be written in the first place to even "catch up" with the reference papers. I had no option to keep it up to date when I started.
With all the talk of mathematicians possibly being obsolete, I'm wondering where the future conjectures that future LLMs would prove might come from.
My read on the current sitch is that the drama is more around the humans behind the LLMs, like the OpenAi *people*, who are dishonest, thieving, technocrats, stealing human mathematicians work and claiming credit for it. Since the LLMs are only trained on finite amounts of data on the internet I don't think anyone expects them to replace mathematicians in a serious way. But surely the field is changing dramatically, people will adapt.
I'd imagine that in three months when we all have access to communicating agent swarms this should be easier
> Instead, I have a more complicated view, which I actually expressed in my essay The Two Cultures of Mathematics a quarter of a century ago, and which can be summarized by saying that there is a spectrum of attitudes in mathematics to the relationship between problem-solving and conceptual understanding. At one end of the spectrum you have mathematicians who are primarily motivated by the wish to solve problems, who see conceptual understanding as a very important means to that end. At the other you have mathematicians who are primarily motivated by the wish to attain conceptual understanding, who see problem-solving as a very important means to that end.
Before, understanding and problem-solving-ability were so interdependent that distinguishing between the two was practically very difficult and probably wouldn’t have changed anyone’s research agenda. Now, they’re not connected, and this guy just did the ultimate meta-experiment of seriously undertaking a project that is intentionally 100% problem-solving and 0% understanding to prove it (maybe 99% and 1% but pretty close. In his transcripts, he never asks ChatGPT about the math, only about its opinions of the math).
As we (as a society) sit around asking ourselves what mathematicians (and software engineers, and anyone in deep technical fields) should be doing all day, we now have this case study to show us how wide our range of options has become.
> So, assuming my proof doesn’t rely on a Lean kernel bug, it’s likely to be legit too.
He lacks the understanding to verify his solution properly, and has to lean on those who do have the understanding to verify it, only being able to say himself that it's "likely" to be correct. (And what do those mathematicians get for laboriously checking the generated proof? 40 grand?)
Seems to me problem solving is as dependent on understanding as ever.
The only thing that needs a check is this 500-line file: https://github.com/gaearon/conway-refinement/blob/264445c93b.... If this file is correct and Lean kernel is correct, the proof is correct.
Moverover, the version I linked above is intentionally paranoid so it doesn't use any third-party code except Mathlib. If you allow usage of CombinatorialGames and trust its definitions, the part that needs to be checked narrows down to exactly 20 lines of code: https://github.com/gaearon/conway-refinement/blob/264445c93b...
There are two ifs in this sentence.
> Why is it a problem for me to publish a result that relies on it?
Bit over-sensitive here. I never said it was a problem for you to publish a result. You can do what you like on your blog and spend your tokens however you choose, just as I'm free to have my own opinions on the value of such an effort. I was responding to, and disputing, a commenter's assertion that understanding and problem-solving ability are "now ... not connected".
While Lean is tightening things up after the recent LLM-driven hacks, I agree that bugs are possible. Although usually code that exploits them is obviously aggressive and is deliberately using the more obscure features related to metaprogramming. Also note that my solution has passed the nanoda kernel as well (https://palomar-registry.org/entry?id=PALOMAR-2026-09-03-000...).
That said, again, I never implied that I'm asking mathematicians to "laboriously [check] the generated proof" which is what your parent comment says. The value to mathematicians is knowing that the conjecture is probably right, and knowing the rough path the LLM has taken to it. Instead of checking the Lean proof line by line, what mathematicians are interested in doing (at least, the ones I've been in contact with) is finding a shorter and more direct proof now that they're aware of the outline and main intermediate claims. As for how much value they find in that, I presume they would be able to speak to that when/if they would like to make their research public.
Ah, or is the value to mathematicians that their LLMs can build results on top of this? (In which case, did this do more than save them some tokens?) Or is the value the deep mathematical insight that this result incidentally gives a few mathematicians the confidence to develop for their own personal satisfaction (e.g. if they decide to go and prove it for themselves, and come to the same result after a lot of work)?
(IMO, the deep, scary question: what if it’s soon impossible to make anything at all that anyone who doesn’t know you personally would bother to look at or use? https://www.smbc-comics.com/comic/crack)
This guy isn’t committed to understanding anything. He’s just screwing around and hoping other people who are turn this into something beneficial to others. He’s just extracting value built up by others over a long period, depleting the finite resource of motivation to work on this topic.
Would you, in this situation, be mad at the aliens? Would you say the aliens have "extracted value"? This is kind of how I see this project.
The alien in the analogy is the corpus of knowledge that’s newly reachable via LLMs.
Now I feel like all the intellectual hierarchies and reward systems are broken. Who’s gonna waste his or her fucking time and money in degrees and papers when you just mess around Claude?
Isomorphic plagiarism makes people feel 23% smarter, but it also provably degrades core skills by 17%.
LLM are great at context search, but are also trivially proven degenerative under recursive self improvement scenarios. We look forwards to stripping their assets at a heavy discount.
Also, we shouldn't kink shame peoples cognitive dildo choices. =3
though I am an expert at coding, the author's process sounds very similar. constantly double checking, asking for explanations, having AI adversarially check its own work, trying to detect bullshit
> Me: btw how’s your mood overall?
LOL. mood??
author provided nearly no intellectual input into solving the problem, so they IMO don't deserve credit. and it doesn't make sense to anthropomorphize LLMs and talking about their "mood". these are two unrelated statements.
as to if you want to give the credit to the LLM, or if you believe nobody gets the credit, is another separate question.
> Although the current generation of models is trained to complete tasks rather than to enrich our understanding, and today’s AI companies are misaligned with the goals of the mathematical community, I hope that with time we’ll find ways to use these tools in harmony with human research.
while citing "A Severe Misalignment of AI in Mathematics" [1], which condemns exactly the thing he is doing himself: Mindlessly producing theorems without a corresponding human understanding of the underlying proofs.
1: https://mathandai.org/
The happy case here is that my obtuse proof leads to a concise and illuminating mathematical proof, which is exactly my hope for the endeavour. I think this could then be a positive example of AI/human and amateur/professional collaboration.
What would a positive example look like to you? What do you think the manifesto argues for?
Edit I apologize to Dan Abramov for venting my frustrations about things outside of either of our control and unfairly using him as a punching bag. He seems to be an intelligent person and I hope he continues learning mathematics using whatever tools he sees fit, including AI. It was wrong of me to do this and I will take a break from this website for one week.
Positive example: Someone is 1) independently interested in a particular conjecture, 2) he lets an LLM prove and explain it, 3) he ends up fully understanding the LLM proof and is then able to phrase it in his own words.
I got quite frustrated and disappointed when taking pictures because everyone was doing it and my picture of x was similiar to others taking picture of x.
Either you learn from it and accept that and still do it, or you don't.
But its not new
Let’s not infuse good faith into what is clearly meant as a derogatory comment.