You might want to check up on the current controversy surrounding OpenAI’s math “achievements”.
I’m aware of the controversies, this benchmark isn’t about making new proofs on previously unsolved problems, it’s whether it can answer complex math problems, which it’s getting better at.
By my reckoning, the difference between Mythos and Opus is smaller than the difference between Opus and Sonnet. Same with the difference between GPT 5.5 to 5.6 is smaller than the difference between GPT 4 to GPT 5.
Do you have any benchmarks or data to back this “reckoning”
Do you have any benchmarks or data to back this “reckoning”
I work with LLMs daily. I read papers as they hit arxiv. Also daily. You clearly don’t.
I’m not interested in convincing anyone, which is why I’m speaking non-technically.
The benchmarks being cited aren’t as interesting as you appear to believe they are. You’ve not fully grasped the fact that solving pre-made problems where the solutions are known or knowable isn’t anywhere close to the same thing as asking truly novel research questions independent of a human prompt. For OpenAI to also be embroiled in allegations of plagiarism only serves to underscore the gap between the two concepts.
fact that solving pre-made problems where the solutions are known or knowable isn’t anywhere close to the same thing as asking truly novel research questions independent of a human prompt.
I understand that, but the original statement was about the models stalling out in progress.
Would you say a student stalled out in progress if they could barely do 2 + 2 a couple years ago and is now able to consistently solve some of the most complex math problems known just because that student isn’t creating novel research?
You’re trying to analogize your way into a subject you clearly haven’t studied.
There’s pre-existing research here. Godel’s Incompleteness Theorem holds, plus others.
There’s already a known upper bound here that you’re clearly unaware of.
There’s as yet been zero LLM-based architectures that have created new information. Everything they produce is somewhere within the training data. LLMs are a very specialized data compression algorithm, in a fashion.
The stall is around whether Recursive Self-Improvement is achievable. Recent papers out of China are trying to chart a course to it. But, until someone succeeds, The current pace of improvement is already slowing signs of slowing. It’s not about where the finish line is placed, it’s about how fast they get there.
I’m aware of the controversies, this benchmark isn’t about making new proofs on previously unsolved problems, it’s whether it can answer complex math problems, which it’s getting better at.
Do you have any benchmarks or data to back this “reckoning”
I work with LLMs daily. I read papers as they hit arxiv. Also daily. You clearly don’t.
I’m not interested in convincing anyone, which is why I’m speaking non-technically.
The benchmarks being cited aren’t as interesting as you appear to believe they are. You’ve not fully grasped the fact that solving pre-made problems where the solutions are known or knowable isn’t anywhere close to the same thing as asking truly novel research questions independent of a human prompt. For OpenAI to also be embroiled in allegations of plagiarism only serves to underscore the gap between the two concepts.
I understand that, but the original statement was about the models stalling out in progress.
Would you say a student stalled out in progress if they could barely do 2 + 2 a couple years ago and is now able to consistently solve some of the most complex math problems known just because that student isn’t creating novel research?
You’re trying to analogize your way into a subject you clearly haven’t studied.
There’s pre-existing research here. Godel’s Incompleteness Theorem holds, plus others.
There’s already a known upper bound here that you’re clearly unaware of.
There’s as yet been zero LLM-based architectures that have created new information. Everything they produce is somewhere within the training data. LLMs are a very specialized data compression algorithm, in a fashion.
The stall is around whether Recursive Self-Improvement is achievable. Recent papers out of China are trying to chart a course to it. But, until someone succeeds, The current pace of improvement is already slowing signs of slowing. It’s not about where the finish line is placed, it’s about how fast they get there.