It's not possible to keep shrinking down parameters and keep "frontier" performance, it's like saying it's possible to take a 3 hour movie and compress it down to 3 megabytes, there are information theoretic limits on the amount of bits of information that can be compressed.
What I'm saying is, if you're expecting a model that can be run on a 16GB or 32GB machine with the intelligence/knowledge of Mythos or Sol, it will never happen. It cannot happen, just like you cannot watch the Odyssey saved as a 16MB file.
Smaller models can get faster and smarter, but by definition they can never compress all of the knowledge of a frontier model and they will approach a limit by which they cannot get better.
The current models are not close to approaching the limit of compression for intelligence. They aren’t even focused on it like Chinese labs are. The training of Qwen’s 27B parameter model showed that by structuring model training from fundamentals to more difficult topics they were able to drastically reduce the number of parameters needed.
The ‘frontier’ models rely on scale to achieve their results but that’s not the only approach. Eventually we will hit up against the fundamental limits but we are not close with Sol and Mythos.
Yes they are approaching the limits, try asking smaller models niche questions about almost anything, they hallucinate massively because you cannot simply pack in all the raw knowledge from a massive frontier model into something that’s quantified down to 20GB etc.
It breaks fundamental laws of information theory. It’s like saying you can extract 100 joules of energy from 10 joules of energy source. Not possible.
It doesn't really matter though. Hardware performance is still growing. The new Mac Studio could just about run this model locally (rather slowly) - something that sits on your desk, that you as a consumer can buy.
Imagine prosumer desktop hardware 10 years from now. The 2036 DGX Spark. For a few thousand dollars you will be able to buy something with hundreds of GB (maybe TB if manufacturers step up) of unified RAM, memory bandwidth in the 10-20TB/s range. Overall AI "compute" will increase 10-20x, while at the same time AI model capability per byte will increase 5-10x.
The hardware would fit today's models, something like Kimi K3, quite comfortably and give performance of maybe 100 tokens/second. So what needs data center hardware today will run on your desk.
But if we also assume the models become more efficient, a 2036 Fable-class model (in terms of intelligence/capabilities, not size) will easily run on this thing at hundreds of tokens per second.
Unfortunately it'll still slow to a crawl with 5 Chrome tabs open, and every Electron app will need at least 200GB of RAM.
I think the assumption here that might not hold is simply that increases in efficiency and smaller size will be achieved by linearly just training smaller models better.
You are absolutely right that there is a physical limit about these things, but very often I find that the solution is a clever way to work around the problem. Maybe the problem with knowledge of the models will be improved by them looking the information up in a better way - so smaller models will not have to have the knowledge trained in but will default to checking. Maybe Models will, I dunno, focus on training in assembler and start to only ever check the compiled output so they only ever need to learn assembler and will then compile the solution to reason about the assembler code.
Obviously that last part is a ridiculous example because I'm not gonna be able to come up with a solution myself - I'm not nearly smart enough for that. But I h ope you get what I mean. Not going the direct route but instead finding solutions people didn't think of before.
Sounds like you are describing a quantized model which is a naive form of compression, not a model that is trained more efficiently.
Additionally the information theory angle is for information storage, but a model can access resources and tools to gain information and what we are really seeking to train is reasoning not information retrieval. We reduce the needs to the right capabilities and we don’t get upset if it does not know the lyrics to every song ever written.
You heard of JEPA? LLM's have all sorts of garbage they have memorized. Reasoning in latent space instead of in text significantly reduces the number of needed parameters.
You should look at some literature around it. I don't have time to pull it up now but it's been shown that much smaller small million parameters JEPA model outperforms much bigger LLMS in some applications. Keep in mind JEPA is area of active research.
Let's see.... Unemployment, AI Lunatics, forced to use AI, AI Slopping...
Much of joy we got from coding was taken way, I can not even imagine anyone wanting to get back to a coding role to vibe slop all day long.
I even heard from colleagues that presented algorithms better than the AI slop and heard verbatim from their boss: I prefer the AI solution instead of your.
Eh, this is just the latest rendition. Before that we were still complaining about tech bros ruining tech.
With all this said, the demand for ever increasing profits squeezing everything that exists till it bleeds out will eventually ruin anything so even without AI tech would have reached that point.
I don't buy it, this is essentially vibe-consulting.
Deal with one customer can be quite a nightmare, and honestly, sometimes using off the shelf solutions can result in saving time and money.
Even today I'm creating SaaS, not every problem is simple to solve. Take a look at CRMs, there are thousands of CRMs due to enormous egos wanting to build "my way", "the way", "the superior way", "our company is not like others (LOL)", etc. At the end of the day it's just a database with CRUDish UI.
My customers pay me to solve their problem well because they can not solve it themselves, when they do they just don't have the time.
Reality check: I have been coding many products with all frontiers models: Still takes a lot of time and, surprise surprise!, money for tokens. Ironically software development became more expensive for serious software that won't vibe-break/vibe-delete-prod/vibe-delete-database.
They're already feeling the effects of competition.
They try to push the narrative of having the best models but people finally discovered that Chinese models are actually great and a lot cheaper, I don't even think this is a reaction to GPT 5.6
I was really curious about it - looks like the extra drive slots are NOT wired to the GPU's VRAM, they're just using PCIE bifurcation to free up some extra lanes for people who plug an x8 GPU into an x16 slot.
So, it's just a matter of time until they destroy this project in favour of their cloud interests. Such a shame, it is (was) a nice open source project.
Nice in theory, in practice I remember having to support Internet Explorer about 4 years ago. Hard to justify the investment sometimes, at least polyfills gave use some sanity back. The only reason to do it was: Rich old enterprise customer who can't install chrome due to policies created by Dinosaurs.
Websites are surprisingly hard to maintain long term, specially for a broad audience of devices. Developer Experience can lead to better UX, the easier it is to build/maintain, the more likely we're to do it.
Given how bad AI is at design plus all the unstoppable slop train, I expect websites to become much, much worse.
Now game devs can optimised their game for it and Steam Machine will get royal treatment. It's not unusual for a PS5 game to run slow on a PC with much better hardware due to not being optimised. The Nintendo switch is a great example, pretty old hardware but the games run well (for their intended experience).
Many people are complaining about the price but you can bring you entire steam game collection and even use as a PC if you want, I sold my PS5 once it became a useless brick cause Sony prevents you from running Linux.
reply