> Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results.
Looks like they listened to Terry Tao’s request for CoT in his talk on LLM use in mathematics?
Since it's only discounted on the standard "OpenAI," non-ZDR route (old pricing on Azure), I'm guessing a lot of users won't see this benefit? Since a lot of users enable a global "ZDR-only" toggle on OR
Yea the "When the fact lives outside the model, a wrong answer has an address" sentence seems aggressively AI written. Saw that and my senses went off.
Ah Ah Ah – how we did't like this. Downvoting me will surely help!
Can't you see how much slope is already around?
Don't we – you and me – contribute to that, especially at work?
Isn't "dull your senses" standard answers of most expensive shrinks?
So what did you disagree with?
Or you simply didn't like the truth? Ok, I got it. No problem.
Neat. Curious to see if RL pans out.
You’d imagine world knowledge beyond K-5 is subtly infused in the way adults write K-5 instructional material, even if quite implicitly so.
Yeah, the filtering process probably wasn't particularly robust. The Schrodinger's cat example ("It's a cat that has been misbehavin'!") sounds like... a joke?
Looking at the paper, it looks like they started with FineWeb-Edu, then filtered it based on an "age of word acquisition" dataset, with word frequency used as a proxy for values not in the dataset. They "only discard
samples in which more than 5% of the words exceed the target age of 12." Maybe 5% was too high? They also filtered out beyond K-5 math symbols, like sigma. Then they trained a classifier to do more filtering.
And they tested it on two grade-level benchmarks, and it only got 0-3% correct on the beyond k-5 boundary, while also decreasing in performance on the k-5 boundary (which they say is an acceptable tradeoff, since they were trying to get a sharp cutoff). So presumably since they got good results from the benchmark they stopped.
> The modes are the load-bearing piece:
lol
reply