> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
There is just a lot of private data and even proprietary code (like SupGen) in the commit story, so I just squashed it. I didn't think that'd be an issue? Why?
Or AMD’s putting sram on a separate chip. I’m surprised they didn’t go there yet except for extra L3 instead of all of it. At some point the costs will shift the decisions.
Nothing about this depends on the provider, you could spin up a local Qwen and give it a search tool like SearXNG or something. At this point local models are more than good enough for simple tasks like that. Using ChatGPT is just (usually) faster and simpler
You know how there's a router mode to use the cheapest provider? That only takes into account uncached rates, last I checked. Make another one that takes into account effective rates (the ones that include cache).
So then they DIDN'T "learn the secret to cracking the problem". They simply knew that part of the problem was solved. Knowing a problem can be solved and knowing the solution are not the same thing.
The claim that OpenAI somehow used the mathematicians' ideas to leapfrog them seems unsupported at this time and IMHO it was irresponsible to bring it up because credulous people will immediately believe that narrative.
And from my perspective, if some math folks typing in a few questions to OpenAI provides sufficient training data for OpenAI to solve a big problem... that's amazing! A few conversations/prompts out of the billions that OpenAI trains on lead to this result- that means there is an awful lot of low-hanging fruit that could be exploited cheaply.
I'm curious how many other 300 billion output tokens OpenAI has "paid for" that have resulted in no breakthroughs.
Either they had a pretty good idea that investing this type of money in that compute on a model in training would lead to these specific results, or they gambled with other people's money.
I want to hear about the gambles and expenditures they don't brag about. In America's energy economy, there's finite resources to expend.
The (unprovable, yes, without OpenAI being willingly transparent) argument is that openAI constructed a prompt to scoop them using some inside knowledge about the approach, which they allude to in the announcement.
In the transcripts, Brubeck is very cagey and evasive about the prompt, when it was supplied, and its contents.
> learning the answer might be in model X’s training data made them believe that model X specifically might be able to solve the question, and they were able to very quickly find enough certainty about the former to commit millions of dollars to the latter.
They don’t need to know, because their IP stealing machine knows for them. They just have to buy enough compute, and someone else’s work is theirs.
I said “very quickly find enough certainty” to suggest hypothetical situations like “someone searches the conversation logs, confirms for themselves the solution is present, then shares the confidence gained from this knowledge without explicitly sharing the knowledge itself”. That person could recuse themselves from the project so the project can still legally make claims like “conversation data was not used” in the announcement, while also knowing that they are guaranteed to get there if they just pull the lever enough.
(Naturally, I have far too much respect for OpenAI’s legal team to suggest this is what happened in their project.)
reply