That's what they're always going to be, so not sure what would be "insane" about it. They literally feed on and emit natural language, and are put to work on informally defined, arbitrary tasks.
When people figure out any reliable strategies to test and benchmark them, that's insane, and in the positive sense. This very same issue has been a thing for humans as well forever, and remains only very questionably solved (IQ, academic tests). This is not easy.
Doesn't just sound like it, it is. That's life for you. That's what I'm pointing out to you. [0]
Just consider your own example. Do you think a less or more "rebellious look" is not something designers can actually ellicit? Less so in software design, sure, but in character design for example? Or general product design? Do you think e.g. Monster energy drinks are branded the way they are completely due to happenstance or something?
Except people don't usually put numbers to it, because they understand that that's hard to defend. You're the one who's describing such an idea, and wants such a thing to happen, classifying anything else as just vibes (that's the point!) and unhingedness. You're handwaving the difficulty and fundamentally limited nature of that, assuming that it is some laziness or mental delusion that's preventing it instead. You're also pretending as if it was somehow not real as a result. What I'm telling you is that you're wrong about that. Any kind of qualitative analysis that's actually defensible with these is genuinely difficult and limited in nature. See also all the opining about benchmaxxing. It also doesn't mean they're useless though, see also benchmarking.
The guy above didn't put numbers to his vibe assessment, they just drew a comparison, exactly because they know that there's not much else they can earnestly offer. You're sulking at them not lying to you by overstating their rigor, and you're flipping the arrow as if this limitation was some sort of mistake, not a necessary and intrinsic property, which it absolutely is. Natural language is an inherently subjective medium.
I think we’re saying the same thing?? It is literally just vibes and all this faffing over which model has better “taste” is pointless
Where you’re wrong is pretending there is any intellectual rigor to the discussion which justifies promoting from the domain of vibes to actual reasoned debate
We do agree that this is a vibes discussion, but disagree on whether that makes it a pointless discussion to have.
To give you a practical example, small, self-hostable models are very popular on HN. But my personal experience with them has been absolute dogwater, so this difference informs me that this is not the place where I should shop for a signal on whether a specific model like that is worth trying, or on whether it's game over for large models and remote models yet. I can "take the temperature" and make use of that without it having to be any rigorous, high assurance or mechanistic thing. It's suboptimal, but not useless.
Conversely, it also inspires people to try and substantiate these issues, so that it can eventually be more rigorous, higher assurance, and mechanistic. This is why I brought up benchmarks and benchmaxxing. Hard to know what points of consideration are salient when there is no abstract grievance to investigate, but Goodhart's law does also keep looming.
If people start moving away from Claude to GPT because e.g. Claude's output is too hard to work with and parse, that's relevant for the respective model providers, because it's a revenue share shift.
If I really like using Claude and strongly prefer its output stylistically, then claims otherwise will infuriate me, and will drive me to substantiate. If for no other reason, then because it's likely that Claude will have its language tuned in response to people's feedbacks and the revenue shifting, which may not be to my liking, and so I better prepare to call it and contend it.
It's literally like any other real world thing ever. Think tuning video codecs. Or tuning user journeys in frontend development.
It's easy to dismiss "taste" when you either have none or just fail to appreciate it, but nothing gives you an appreciation for the importance of taste like LLMs. There is no benchmark for taste, so while many things improve taste does not. Bad taste is, in fact, a huge component of what makes AI slop so sloppy.
But human coders can have bad taste too. There is code where there is nothing obviously objectively wrong, yet the choices feel like they were made by someone who just doesn't value or put emphasis on the right things, yet spends a lot of effort on trivialities. It comes in many forms.
You are confusing personification, which is purely a rhetorical device, with anthropomorphization. I am not anthropomorphizing GPUs. There is nothing particularly weird about personifying model weights, as we do with computer programs, cars, and all other manner of inanimate objects every day.
To be honest, in this case, I wasn't even personifying them, because I in fact didn't say that an array of GPUs has "taste", or in fact even that model weights did. I was saying I preferred GPT 5.6 Sol's outputs as a matter of taste. I can see why someone would confuse the two statements since in this case they're pretty much the same thing, but if you re-read what I said I was actually more careful than you're giving me credit.
It's just vibes