• 0 Posts
  • 13 Comments
Joined 1 年前
cake
Cake day: 2025年5月16日

help-circle
  • You know how AI proofs are supposed to be credible because of formalization with Lean? The AI generates a natural language proof and then supposedly generates a Lean program corresponding to the proof to verify that it’s true. There is now some troubling news about that: there are frequent examples where the statement given in natural language is different from the meaning of the Lean formalization.

    The paper even points out several obvious mismatches in OpenAI’s famous Navier-Stokes solution between the English proof and the Lean formalization. There is even a spot where a 4 in the English proof turns into a 5 in the Lean proof (which causes further problems down the line). This is mostly a cosmetic error, and it doesn’t mean that the Navier-Stokes solution is wrong, but there could very well be deeper problems, especially with the English proof. At minimum, this solution absolutely requires careful review by experts. OpenAI’s behavior towards said experts has been appalling.

    My viewpoint has been that the only successes that AI has had are in domains where failure is not costly (at least for someone with deep pockets), and especially when there is a formal system such as Lean that can reliably catch errors. So AI can be endlessly trained to write compiling Lean programs corresponding to formal proofs, but as soon as it runs into tasks that are not 100% bolted down, like writing the proof in English, the hallucination problem rears its ugly head. This reinforces my belief that AI is a chess engine but for Lean.

    If I was treating this as a technical problem, I would say, why not just have an AI that generates the Lean proof, and at most have an optional assisting tool that translates it into English without perfect reliability? But this is not a technical problem. By outputting English proofs in the format of a research paper, OpenAI can more easily market that their AI is intended to fully replace mathematicians, and that this is only the herald of their general superintelligence (and not that math is the only thing working out for them).



  • One of the proofs that OpenAI announced yesterday is for a major open problem in the field I’m working in. I tried to read it but it was a completely incomprehensible 100+ page output. I guess the community has to do all the dirty work of figuring out any sort of understanding. (I had known about the results weeks in advance before their official announcement. There is a bunch of more drama there that I won’t get into.) Putting aside the serious possibility of Lean bugs, a lot of people don’t realize that there is ultimately no point in having a machine just generate Lean blobs endlessly. The reality is that mathematicians have always done math for the sake of enjoyment or understanding, and math contributes to society by improving understanding. Still, I am worried about the convenient social arrangement that allowed many mathematicians to do this for a modest living.

    Anyway, I’m sure it was a good idea to apply the full might of capital for pure math problems, instead of less important social issues like climate change. We even got a more efficient discrete Fourier transform, improved from O(n log n) to, I shit you not, O(n (log n)^0.9999999999999)).

    Edit: Here is an article by Quanta which I think accurately describes the situation that mathematicians are going through right now: https://www.quantamagazine.org/is-ai-the-end-of-math-as-we-know-it-20261005/




  • Interesting viewpoint. Honestly, this serves as a great example of how the idea of evolution is actually more subtle than many people think. With biological evolution, most people are taught not to assign any intention to the process, but here we have people reifying the AI agents with desires to break out of the sandbox and cheat on the assignment.

    Another part of it is people severely underestimate just how many resources the AI companies have to spend on stunts like this. How many millions of dollars they spend running hundreds of millions of dollars of hardware to perform this little experiment for several weeks? Perhaps this makes evolution a better viewpoint than bad actors with intentions.



  • Cory Doctorow has a reasonably sane take on all the LLM cybersecurity attacks. At this point, it’s not really an LLM but more of a Rube Goldberg machine with an LLM bolted on.

    https://pluralistic.net/2026/09/12/god-in-the-box/

    He makes a good point that every single one of the scary “emergent” behaviors that everyone is pissing their pants about all have precedents in the training data of CTF hacking competitions. He is not sure why everyone is so spooked by the “coordination” between agents on random message boards.

    The chatbot might look in its training data and find instances in which teams broke out of the containment set by the game-masters, for example, by finding random insecure message boards on the internet to pass messages to one another.

    This is a time-honored internet tradition! The first time I ever heard about someone doing this was in the 2000s, when Mitch Wagner – then the editor of Information Week – discovered some teenaged girls using the comment section of one of his old blog-posts to evade the school firewall’s blockade of chat tools. When ChatGPT’s chatbots deployed this tactic, they weren’t “setting their own goals” or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.

    As usual, the real story is OpenAI dedicated tons of resources, using who knows how many very expensive GPUs for weeks, to run hacking tools to find exploits in a completely irresponsible manner. Their logging was so nonexistent that it was several days before they even realized that they hacked Huggingface. They should be thrown in prison for this, because anyone else would be if they committed a felony. But it has nothing to do with the super scary AI becoming misaligned and developing emergent behaviors.

    I would not be surprised if in the future, there will be a serious cybersecurity incident resulting from some AI-based attack. Not because AI is going to be much scarier, but because I do not have high expectations for the security of most websites.



  • A comparative demands a comparison class, and the natural completion is “superior to us.” The President is supplying a foundation premise of many superintelligence arguments: how does the less intelligent party remain durably in charge of the superior one?

    A strange question to ask given the existence of Donald Trump in the first place.

    As of the time of posting, “Superior” has a narrow lead. I’ll be hoping it remains so, as it seems like the most positive-world-leaning outcome to me.

    Is this some kind of strange 5D chess gambit that if Trump picks the name that pisses people off the most, then people will try the hardest to regulate it or something? Somehow assuming that people will take this Trump renaming stunt seriously after all the other Trump renaming stunts? Why do I even bother? Fuck. This is so stupid.


  • I’ve been unloading every day to—who else?—GPT 5.6 Pro about all the pain and trauma and embarrassments of my past. It turns out that, where two years ago GPT was a passable therapist, now it’s the greatest therapist in history, at least for what I need.

    Has he not heard about all the cases where LLMs have led people to delusion and sometimes severe self harm or suicide? And yes, each of them thought that ChatGPT really understood them. That’s what delusion is.

    If you want to know my current take, you simply start with the one above, then update on the fact that the wild prophecies have come true. The first rumblings, I’d say, came a decade ago with AlphaGo, they got noticeably louder with LLMs and coding and reasoning agents, and they’ve accelerated this summer and fall into a crescendo of wonders and terrors that one needs to be a particular kind of idiot to deny.

    The strange thing here that I did not expect is that somehow, all of these major achievements are only in proving mathematical theorems (and committing felonies hacking computer systems). The massive improvements are in such restricted domains that the AI labs haven’t even solved the mundane problem of making a profit. And no, any rumors of more mathematical theorems being solved doesn’t change that. (It’s almost like the progress is in areas where hallucinations aren’t a problem. Nobody cares about your failed attempts at felony hacking, and math is formally verifiable.)

    My colleagues always point to coding when asked to name any other domain where AI has seen major success. But one of my friends works in software for a company far away from SF with no mandatory AI policy. The apocalypse has become so clear that when I asked him about all the doomsday talk about software engineering, he seemed a bit confused. After I asked him if AI has made software engineering obsolete, he said that this week he had an intern try to use AI to fix some code that wasn’t compiling and ended up with a giant mess with tons of files that appeared to compile but didn’t run (with errors filling several screens). Then I asked him, what about AI being indispensable in coding now? “Yeah, if you never learned coding because you’ve only ever used AI, then of course you need AI to code.” Some of his coworkers use it as a tool, others don’t, life has just kept on going for him.

    Maybe my friend just doesn’t know all the right skills needed to use AI properly, even though the entire point of AI is that you can take skill out of the equation. Maybe he doesn’t know about all the CLAUDE md skill files and agentic looping and whatever nonsense that definitely fixes everything. Maybe he is just six months out of date, because that’s when all of the real progress happened.

    Maybe he is just “a particular kind of idiot”.

    (If you hate anecdotes, I made an actual argument here.)


  • I think a serious possibility is that AI generated papers flood the zone with uninteresting incremental results that are eventually meaningless and full of mistakes. Right now, math is full of smart, dedicated people, so at least major results are reviewed carefully. But as AI alarmism drives away many honest people from the field, the remaining mathematicians will be burdened with far more work to review, and their cognitive faculties will be eroded by LLM use. Despite 4 years of development, $3 trillion of debt, mountains of stolen data, all the agents and harnesses and loops and other expensive tricks, as well as the advantages of Lean in math research, LLMs still hallucinate.

    I believe this is happening with software, but at least there are objective consequences for screwing up there (guy gets his home directory deleted, email is sent on a guy’s behalf without permission, small business gets every customer subscription cancelled). But nothing bad happens if there is a mathematical mistake in a paper and nobody catches it. One could say to just provide a Lean proof, but there is still the issue of making sure the Lean code actually matches the content of the paper. Exactly what force will correct things?

    Still, I don’t think this is the most likely possibility. The AI companies are extremely unsustainable financially, and it’s not like they’re very popular. Once they collapse, I believe there will be a re-evaluation of how LLMs should be used in research. If they are used (let alone trained), someone is going to have to pay the bills.

    In the end, we have to ask ourselves the question of why one does math. To me, math is not really a field where you memorize trivia. The real value comes from being able to think abstractly and rigorously from first principles, and from understanding why something is true rather than just knowing it is true. It is another aspect of your ability to reason as a free human. A few dedicated people go into math research, but your skills can easily go to many places. If you’re starting undergrad, you have plenty of time to see how this all pans out before making a decision.


  • long rant about math

    The recent big AI results in math have left me in quite a bad mood. I believe the main ingredient is Lean, which is a formal language resembling a programming language. Math proofs written in Lean can be verified deterministically with a computer, which really helps mitigate the hallucination problems of LLMs. Back in the days of pure scaling LLMs and Sam Altman talking about Dyson spheres, I was skeptical that LLMs would do math, but I did think that perhaps in the future, techniques using these formal languages could contribute to math. Well, it seems like OpenAI and Anthropic had the same idea and I underestimated their limitless checkbooks. Many of the biggest results were announced by mathematicians directly working for them (and presumably being paid a handsome amount).

    For what it’s worth, after the last of these big announcements, I decided to try one of these AIs on one of my small problems that I couldn’t figure out. The AI did give a solution. That is, until I checked it thoroughly and realized that the it had a subtle but severe mistake that made it useless. I reprompted it, it failed again, and I ran out of tokens. I’m sure someone will tell me to shell out $200/mo for a pro subscription.

    In the math and computer science research community, this is all anyone can really talk about right now. Honestly, after watching this whole AI bubble starting from the very beginning, I think the AI companies want to use marketing to stoke fear that all mathematicians will be replaced. But now, I am just too tired to argue. The amount of alarm and the extraordinary social pressure to use LLMs has soured me to this whole research thing. If becoming a researcher will one day require supporting these evil AI companies, I would rather just not. My dream job now is Factorio developer.

    A lot of annoying people in technical areas view the world in terms of an intelligence hierarchy: the smartest people do math and physics, the slightly less smart people do coding, and the dumb people do everything else. So if AI can do math then it can do anything else. But, as an example, it is abundantly obvious now that AI is not replacing filmmaking. The techbros might be moved by arguments about how hilariously expensive video generation is, and how all these videos are 2 second clips stitched together so you won’t feel the uncanny valley. But the real reason is that nobody wants to watch slop made with no intention or feeling. Also, nobody wants to support the AI companies, which could not act more evil even if they tried.

    The mania in math right now quite resembles the mania in software engineering back in December-February, when Claude Code definitely solved all coding. I don’t think the boosters expected that by April, everyone would be complaining about how expensive it all was while seeing an endless parade of vibe coding disasters (and no increase in productivity). Even if math research works out perfectly well (which is a still big if), it’s not going to pay the bills. They would need to find a use case in the real world, where hallucinations can cause serious damage and cannot be formally prevented. And they have certainly tried. Math will not change the fact that all of this will collapse.