Hacker Newsnew | past | comments | ask | show | jobs | submit | easygenes's commentslogin

In the lead-up to the Grok 4.5 release, the CEO of HuggingFace was petitioning Elon Musk to open the weights. It’s odd to me that he didn’t accept, as the model is at a similar competitive level to Muse 1.2. It would have been good PR and essentially zero risk, while getting ahead of Meta who seemed likely to go back to shipping weights soon.


Oof, even the post-mortem promising none of it is AI slop is obvious AI slop.


I was expecting them to state that they were doing a bit with the post-mortem, but the punchline never came. It is even more consistently AI coded than the full text. Give us a /s, give us a wink a nod, give us a markdown sygil of LOSS.png, anything that says "I am actually a human", because all that reads as is BEEP BOOP KILL ALL HUMANS.


That’s not what this indicates. This is the biggest and most expensive to serve, and most capable open weights model yet. They’re just pricing it in line with capabilities.

Kimi also offers generous subscriptions. Subs aren’t going anywhere. Think of subs like running an insurance business. There might be some users you lose money on (ones who max out their weekly quota without fail), but they’re managed such that the average subscription turns a healthy profit. There’s never been subsidies in model serving, inference is just cheaper in terms of ops TCO than people assume, and API margins are very high.


> They’re just pricing it in line with capabilities.

So... convergence?

> but they’re managed such that the average subscription turns a healthy profit.

It didn't work like that, or at least that's not how it played out. People max-out their subs all the time which is why strict and multiple limits were implemented by all providers. Also, I subscribe to z.ai and recently they dropped the quota significantly that now their sub offers less than Claude and OpenAI. It's still x5-6 what it would cost on API costs though.

> inference is just cheaper in terms of ops TCO than people assume, and API margins are very high.

API margins (at least american ones) are probably healthy. But I don't think that inference is that cheap. It would cost 300-500k to just run GLM 5.2. There are lots of other factors too: reliability (can you keep the GPUs running all time), electricity cost, sys. admin costs, location costs, etc.. I wouldn't be surprised if the API margins are quite close to operational costs.


It's as good as gpt 5.6 sol and _half_ the cost..


Was fun to see their developers make nods to Le Chaton Fat in the announcements for this on Twitter.

I suspect a true "big new general-purpose" model is around the corner from them, whether or not they were in on Le Chaton Fat for real. They've mentioned it after the media circus. Hopefully more creatively named than just "Large 4".


I'm a heavy enough user that I have both the OAI and Anth $200 plans. I always use at least 50% of my weekly Opus quota at Extra setting (meaning I use double the limit of the $100 plan, at minimum). Max I rarely touch because it is twice as slow and the incremental capability gain is minimal. Usually if Opus can't sort something well at Extra, the answer isn't to use Max but to hand the issue off to GPT-5.5 at XHigh.


I too have settled into a kind of dual Claude/GPT model setup. I will often use one to review the other's work, or critique the other's plan in some way. Sometimes I'll have Claude implement a feature one way, then have GPT do it the other way, then have them both review each other's implementation. Then synthesize a final plan from the previous implementations+reviews.

I might just be having fun with models, but I have actually noticed their capabilities vary somewhat, and so my (perhaps vain) hope is that by using both, one can catch each the other's blindspots. It's still unclear to me if that's consistently happening, but I am making substantial progress in my personal and professional projects, so something seems to be working.


> Sometimes I'll have Claude implement a feature one way, then have GPT do it the other way, then have them both review each other's implementation. Then synthesize a final plan from the previous implementations+reviews.

I've done variants of this a number of times, but feel like it was a generally waste of my time to then have to compare them and write up which parts I liked or disliked: if the output is something substantial, each will have its pros and cons. Clear-cut wins aren't very common. Of course it could work well if we automated the whole thing with an orchestrator; you just need a model with actual good taste (according to your own preferences) ... so we'll have to compare all the models to find that one


Yes, same, between the two of them I feel like results are just better because they have different priorities.

At the same time, I’ve invested in tooling that prints and lints architecture I want, so which model is less of an interesting decision, because the results tend to be very close.


There are. If the kernels are nondeterministic (e.g. timing issues) there are minor changes between runs, on a single system, even with eager decode enabled (typically what temperature=0 achieves).


This is a strange one. We know the hardware capabilities of Cerebras force them to do aggressive REAP pruning to serve Kimi K2.6. Meaning that about 750B parameters is the upper limit of what they can serve economically. Not sure if this means Sol is smaller than anyone thinks or that they're just going to charge so much that a very inefficient serving regime is feasible.


M5 Ultra will ship before end of year, likely. Though with current RAM shortage, likely max spec will be 256GB and in short supply.

In late 2027 or early 2028, Nvidia will release Vera Rubin DGX Spark, likely with double or better the performance of current Blackwell, though unclear if memory capacity will go up much from current 128GB. Two to four of those will run models like this decently.

In 2028 we should expect Vera Rubin RTX discrete lineup, including the replacement to the RTX PRO 6000. Likely memory spec will be minimum 128GB. Good chance of up to 200GB. Two to four of those will run NVFP4 models in this class very well.


It might be M6 Ultra and I think the real reason for stopping selling top-tier units was to avoid mid-generation price hikes and increasing demand for the more expensive next-gen systems that I assume will come with 512gb (maybe 1TB) of RAM and a massive markup to match.


I hope all this speculation comes true. Right now this ram crunch is ridiculous.


Article reads as though written by someone who doesn't have much experience with deployments like this. Underestimates the memory needed to run with a reasonable amount of context. Misses two other obvious targets:

  1) 4x DGX Spark (or equivalent other GB10 boxes) with a switch (MikroTik CRS504 or CRS804) and TP=4.
  2) 4x RTX PRO 6000 box. Probably the most practical for cost/perf if you want on-prem as an individual.
Both would be best to run a 2-bit quant so everything can stay resident (article claims you could run a 4-bit quant with 4x RTX 6000 Ada, and while technically true it would mean a lot of the weights are streaming from DRAM, so it would be slow and impractical. You would need 8x RTX PRO 6000 to run 4 bit at a good speed).

This model quantizes unusually well: https://unsloth.ai/docs/models/glm-5.2#quantization-analysis


Can you really say you're running GLM 5.2 if its a 2 bit quant? It might be usable but the capabilities will definitely not be the same.


The Wired headline reframes the issue in a way that’s misleading. SK Telecom was a previously resolved issue (as in prior to Fable launch).

It may have been a contributing factor, but the crux of the shutdown was the industry reporting of Fable jailbreaks (reportedly spearheaded by Amazon CEO Andy Jassy). The more interesting and honest angle is that the industry which has taken the seriousness of Glasswing at face value felt blindsided by Fable release and totally exposed by the residual risk, when they know they still have a months-long bugfixing backlog exposed by Glasswing and are desperate to buy more time.

This misleading looks deliberate on Wired’s part, to appear as though they’re getting a scoop when they’re really just being dishonest. Shameful.


> honest angle is that the industry [felt] exposed by the residual risk [and] have a months-long bugfixing backlog exposed by Glasswing

Two problems with this theory.

1. Amazon complaining to the White House wouldn't have been the opening salvo. Amazon and Anthropic would find it much easier to talk to each other than go through the White House. We'd need evidence that Amazon (and probably others) already asked Anthropic to not release a Mythos-class model but Anthropic released it anyway. Are they on record saying this?

2. The jailbreak Amazon found needs to be real. Maybe the White House staffers are not AI experts and they don't really understand what a jailbreak is... but it's much harder to make that claim about Andy Jassy. For the jailbreak to be the real reason for the export control order, the jailbreak would need to be significant and cause material harm to Amazon. Then Jassy might pass it along to the White House assuming he already was refused by Dario.

But there is no evidence the jailbreak was real. There is one story that it amounted to a request, "fix this code." In any case, Anthropic is on record saying the so-called jailbreak didn't enable any vulnerability work that couldn't already be done by other models.


It’s been reported that Amazon has said in the last few days that the gov reached out to Amazon first for their opinion.


Wired has been NY Post-tier (of the inverse polarity) for several years, now.


Is there any serious journalistic magazines/papers left nowadays? Felt like ultimately every single one of them succumbed to the chase of the clickbait in the end.


Ars Technica still fights a good fight, in spite of Conde buyout.


[flagged]


If you check my post history, I'm very left-leaning. I'm just establishment-left+woke (i.e., liberal) rather than populist-left+socialist, and Wired is populist-left. I'm the kind of person who actually likes Obama, for example.

It's actually unfortunate, because a not very credible publication writing about the vile behavior of the Trump administration (including Elon Musk personally literally killing hundreds of thousands of children, and possibly millions) only serves to help them. Those particular stories may or may not be fine, but I've seen enough truly abysmal stories to not have any trust in them as an organization.

Like, within one second of opening the article I already saw the headline is deeply misleading and then checked the comments here to see if anyone else had noticed (and they did). It's a joke. The article doesn't even make sense. It's like an AI-generated high school essay forced to hamfist a response to a contrived prompt.


“Elon Musk personally literally killing hundreds of thousands of children”

What’s that high quality news source you got that one from?



Yep - the "company at the center" of this is Amazon. But they're not alone, I was able to jailbreak Fable accidentally last week: https://news.ycombinator.com/item?id=48576628


Is that called jailbreak?


Yes. The biology/infosec fallback to Opus is the jail. Promoting to get around it (in my case accidentally) and have Fabre find exploits is the jailbreak.


Only if you were jailed correctly. Getting out of jail when you shouldn't've been jailed in the first place is not a jailbreak


True. But in the fictional scenario where I had NOT produced the Rust code myself, I could have lied and said I made it (jailbreaking was literally asking it to check the program I had made for program vulnerabilities or readability issues - perhaps the nosie from 'readability' also contributed).


Yeah ! You should close your eyes on dangerous source code nowadays. Think about a kind of DRM for programming. Perhaps the future is a prompt terminal with no access to the source code. Of course the code should be deployed only on a “secure” anthropic platform.


Is there any interpretation where Amazon is involved but for a reason besides trying to screw over Anthropic? I don't see why Amazon would act against Anthropic when they have a deal to offer Anthropic models on AWS.



Most headlines seem to be misleading these days. Social media broke journalism.


actually been like that for ages but people notice it a lot now and have a way to talk about it openly. mid 2000s CNN was really bad about this kind of thing. headline sounds shocking but once you got to paragraph 4 or so the story starts to change. then when you get to the bottom paragraphs the truth starts to come out - in stark contrast to the headline. not sure how they are these days.


Things really did get worse as online gave visibility into what drives clocks and undermines the value of reputation.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: