Unless I'm misunderstanding, calling this RSI seems misleading?
This looks like an optimization of current training methods, and a good one, but not "RSI" in the sense of a system that can perpetually improve itself forever.
Does RSI actually mean anything specific anymore? RSI, AGI, at this point seem like buzzwords. Sure AGI has definition that are measurable, say "better than 95% of humans on 95% of intellectual tasks" but if we used that definition we already have AGI and almost no one thinks we have achieved AGI. We use AIs to train AIs which we use to train AIs, why is that not RSI? How much human intervention means that is not RSI?
So few terms in AI are well defined. We will get ASI via AGI because of RSI but neither of those three things have any definition except pure vibes.
I struggle with the argument that RSI doesn't already exist like you say, it's existed since before the term LLM (hey, one that can be defined!) was common parlance. Though the biggest use for those is not superintelligence, it's to serve you ads and get your kids addicted to TikTok.
Narrow pre-LLM models that spit out content and ad recommendations have never been capable of also suggesting, let alone implementing, self-improvements.
RSI has always being a well defined name, and you can only have RSI if you have an intelligence capable or creating itself.
It has technically existed for a long time (for longer than the name), but only on academical applications for extremely limited intelligences that could only create something like themselves. And that is still the only form that exists today.
It was never powerful enough to optimize ads distribution, and all the claims people are pushing around today are plain bullshit.
> say "better than 95% of humans on 95% of intellectual tasks" but if we used that definition we already have AGI and almost no one thinks we have achieved AGI
What matters isn't 95% of humans, it's 95% of actual professionals. Benchmarking an AI accountant against people with zero accounting experience is worse than worthless.
That's what 95% is intended to capture. That some level of expertise in an area is should be captured by 95% sample of the population, you could push it to 99% of 99.9%, but you want it to be quantified by a number to avoid arguments that something is not AGI because obscure field or expert exists that AI can not do.
95% better at 95% of the population is already approaching ASI. One could even argue that AGI is 50% better than 50% of the population.
All that said, this over rotation on benchmarking misses something critical. What is general intelligence? We assume that humans have it and we assume it is captured by benchmarks on "intellectual tasks", but it is probably the case that general intelligence is based displayed by judgement on uncertain outcomes. Benchmarks by their very nature have certain outcomes, they have a wrong and right answer.
Test AIs on questions we don't have the answers to and there is no clear right answer, but there will be at some point in the future. What will the economy do? Which US senators will be be re-elected that polling correctly suggests will not be re-elected. Which US senators, currently not in office, will actually pass bills representing the wishes of their voting base? What published papers will be seen are groundbreaking in 5, 10 15 years? What approach to unifying physics should be taken?
Agreed, what I understand from RSI would be models creating new models, or at least upgrading their own weights/architecture. It does not seem to be the case here.
According to industry leaders, we currently have AGI and RSI in the last month or so. Of course, we've seen zero evidence of any of this and have to take their word for it.
SEO is not nearly as effective as it used to be. Organic click through rates have been trending way down since AI responses were added to the top of the search results. So this really isn't just a case of "you're holding it wrong".
If you're selling something, you love having the AI talking about your product. The move to AI is not a negative for businesses wanting to reach customers organically - on the contrary.
Today the fraud is perhaps more obvious. YT has served me the same 20 min ad 200+ times over the last few weeks. I let the entire thing play 10 times and 50 times didn't press skip for the first minute (or the view doesn't count) It's so annoying Ive reduced video watching by 80-90%
The idea they have someone pay for this in the hope I will buy the product?
Or do people really intentionally run campaigns like that?
These bug me so much. Those 20+ minute ads do nothing but piss me off. I have to grab the remote or switch tabs or whatever to get past them. To the point where I'm on the verge of adding YT to the permanent block list along with Facebook, Instagram, X, etc.
Serious question: Do you watch a lot of YouTube? Why not just cough up the $16 for YouTube Premium? Creators actually do get paid more for Premium users' views than they do from adsense. You get to see none of these ads that piss you off.
Look, I run UBlock Origin myself. But that's because there's zero other way to look at any Web content without going insane with videos and other garbage popping all over the place, and even when sites have memberships, I use too many distinct ones anyway, and the memberships may not remove the ads and junk.
But when it comes to sites like YouTube that let me choose between "ads" and "non-ads" experiences... I guess I'm just saying I see a ton of ROI on the $16.
You expect Google to serve thousands of petabytes per month (costing hundreds of thousands of dollars in raw bandwidth) for absolutely nothing in return?
I have zero expectations of any business. If they have an expectation to be paid then they need to offer something of value. Their costs to operate are not my concern.
Well this discussion about getting YouTube premium is clearly aimed at people who use YouTube but without premium, not people who don't use YouTube at all.
You won't pay for something because of inflation? How do you survive, since everything's affected by inflation? Ever buy gasoline for a car or a transit pass? rent? food? What else can't you pay for because they price will go up in the future?
Inflation helped me kick my junk food and fast food habits. I still buy groceries of course but when prices rise, luxuries are almost always the first items to go.
It is giving me 1+ hour ads now based on my view history (I think) and it is actually hilarious. I've watched two till the end. To keep someone engaged for that long you better bring something interesting to the table.
I could call them, tell them I see their ad, ask questions then call again and again and again and again. I just read one may call a company as often as one wants.
Before the ad, you had no brand awareness at all of Liberty Mutual. Now, you know their name enough to call them out specifically in a Hacker News comment. Your propensity to buy has improved a lot just based on that fact, regardless of whether you consciously want to avoid them now.
Sometimes they put one at the start then again 1.5 min into a video and again around 4 minutes. If I don't skip them I may not even remember what the video was about.
> It's so annoying Ive reduced video watching by 80-90%
This is the opposite of resourcefulness - unless the videos were a distraction, not a real goal of yours.
PS Genuine question: where is the fraud in a 20 minute ad?
Steve Balmer explaining his take on how local governments of the US work was a hilarious ad offering on my non-adblocked device the other day - I’m struggling to see vectors for fraud.
I couldn't imagine someone running campaign like that intentionally but it's apparently cheap to do so. 20 minutes of my time is apparently worth 3 cent. Those 5 hours watching the same ad again and again probably cost the spammer 2 euro.
It's not fraud, more like a cheap ass dystopia. It's some how worse than HP and MS combined.
> ..unless the videos were a distraction, not a real goal of yours.
You are not wrong but all work and no play... I will seek entertainment elsewhere. My apologies in advance HN :-)
In chess, a grandmaster just needs to know at what moment in a game there's a critical move to gain a significant advantage over their opponent. They don't need to know the move itself.
OpenAI got wind that a millenium problem was being solved. And that feels a bit like the critical move in chess. That is - it was a signal that AI advanced far enough that it would be worth spending a lot of time and resources solving a millenium problem.
Elsewhere in this thread somebody claimed that at some point OpenAI pointed their new model at all the millennium problems and this is where they got some progress. We probably won't see proof of this, but it seems plausible to me -- I assume there's a list of problems that each new model is tested on, and you might as well put the big stuff on the list, if only to see how the model behaves when faced with a problem it knows should be very hard.
The weak point in this is: how do you evaluate if a partial result is promising? If this cost ~$10M as suggested elsewhere in the thread, probably not even OpenAI can just throw that at everything?
Okay, from the actual linked article it seems that their partial result was finding blowup in Euler equations, which seems pretty big. I wonder how the other attempts went. Did they get nothing at all, or something true but unimpressive?
I downgraded from the 20x today after learning that 20x only applies to 5 hour usage. I have barely used Claude/Claude Code in the last month and am considering downgrading further, even after this update.
I've been using both for a few years. OpenAI's limits have always lasted much longer for me (depending on whether a monkey got in their machine or not). Aside from not lying about limits they also don't assume I'm developing some kind of nerve agent and lock me out when I ask basic questions in my domain or nag me about my late night work hours. There was a time last year when I strongly preferred using Anthropic models, but at some point something about the responses (or maybe the company) changed and I now find them particularly offputting relative to OpenAI. Hope they fix it because the industry needs more competition.
> they also don't assume I'm developing some kind of nerve agent and lock me out
I'm still having problems with that. OpenAI is a lot less annoying than Anthropic, but I still get obnoxious "this content can't be shown" messages in codex far too often.
Kinda surprised not to see their next update being an Opus 5.1, even if its minimal changes, they've already had to address it with the concise mode or whatever.
So my current usage as a Pro subscriber... Not able to even consider using "Sota" unless i shell out for 100$ a month, (lately i've been a bit burned out i am literally struggling to use 50% of my pro plan per week). Beyond that, I have given up entirely on the top Opus model and reverted back to 4.8. If i have work i deem somewhat complicated, i now have an openai 20$ sub, and i just toss out sol after planning with 4.8. Both subscriptions not anywhere close to capping my usage per week, one of them says i can't use their Sota unless i pay for 5x more usage, and the "best" model they do allow me to use, they are neglecting and its by far the worst model I've interacted with in 2026.
it just doesn't interact good with human beings, and it leaves incredibly strange long winded comments within code filled with session context that will likely not be relevant later on.
Also always seems to have this annoying tendency to leave "questions for you" at the bottom of every output.
Just a high friction human interaction type model, imo should never have even been released, regardless if it scores better on whatever tests, its a horrible experience and a downgrade over past models.
I have to wonder if everyone else is just running these models raw without any custom instructions. I hear all these things about voice and code comments and those are all things I've dealt with long ago via claude.md instructions, rules, and hooks. My claude can already respond in any "voice" I want and the quantity and quality of comments is within my control.
My claude.md has a section about not writing those comments, it has stored this in memory, and still every session I need to remind my good friend to stop writing so many garbage wordsalad comments
Maybe system prompt has priority or something but Opus just really really likes writing bad comments
That's why I mentioned hooks in particular. That feels like the right layer for this sort of adjustment. A PostToolUse hook on Edit|Write would be much more reliable than just a CLAUDE.md instruction. The consistency I get from CC comes from instructions at multiple layers.
CLAUDE.md heirarchy: At the top level you've got general instructions you want all contexts to follow and each subdirectory can add more specific instructions in their own CLAUDE.md files. References in CLAUDE.md are not fully loaded into the context. They are loaded opportunistically. So keep important instructions in the CLAUDE.md file itself and not a referenced or linked file.
Rules files: These offer path scoped rules via frontmatter. So you could have specific rules for certain types of files Claude Code interacts with. Certain rules for handling all .cs or .js files for example.
Auto-memory: You cannot rely on this one. I use auto-memory as a cache for potential future CLAUDE.md instructions. I have an audit process that kicks off when the auto-memory gets beyond a certain number of entries.
Skills: On demand context. I don't tend to use /skills explicitly. I tend to have them used in context. I've got a task tracking system I call threads. So whenever I say "Create a thread for X" it has always reliably followed the specific instructions. I've got skills for managing my NAS for searching historical session for sharing content and other things. I use them a lot of times in place of MCP servers.
Hooks: Deterministic scripts run on lifecycle events. I've got hooks that run linters on code files post edit and hooks which tie into the request / response events to push my history into a SQLite database.
Output Styles: CC ships with a few different styles, but you can create your own. This is key for changing the default voice. CLAUDE.md instructions are appended to the system prompt and can fight against the system prompt. A custom Output Style would let you replace the instructions in the system prompt with your own instructions. This can be done at the user level or per project.
Currently, I'm using custom instructions plus reinjecting the writing cues Opus 5 ignores most frequently via a UserPromptSubmit hook. Again and again, I'm reminding the model what voice I want. Again and again, Opus 5 ignores it.
it really just seems like people pump out that its on the end-user, and i just disagree. They have a walled garden around claude code and using their models within it, it should work instantly out of the box when going from an opus 4.8 to an opus 5.0 with the same workflows. it doesn't.
claude.md for all my projects are fairly tight, its seldom where im upset at anything a model does, and if it happens, its likely because i swapped provider and didn't realize i was failing to feed it proper context beforehand.
Opus 5.0 fails in different ways that I haven't had to deal with. Its insufferable with its choice of language, something I've never had to compensate for on any other model across any provider, so of course I have no preexisting rules for that, it also is sometimes just incredibly stubborn and just WONT finish, and requires several just "keep going" prompts.
This is much different than the issues people would make fun of users for in regards to treating models like slot machines and just pulling the lever over and over, this is more its stopping for no reason short of its task, and literally just needs to be told to continue? absurd.
Most of my workflows have reference material, with standards set, why opus 5.0 is the only model that fails to follow those standards and inserts wildly long weird code comments is not a failure on the end-user, thats the model failing. I can be MORE explicit of course, but i shouldnt need to be, this is supposed to be 5.0, its a downgrade. I went back to 4.8 and all these issues vanished.
Opus 5 is remarkably bad at instruction following over long chats. I have to repeat “Be succint”, “talk like a friend or colleague would”, “no rambling” or some variant of it every few messages
Have you tried a custom output style? CLAUDE.md instructions are appended to the system prompt. A custom output style can replace the system prompt. At least the part of it pertaining to voice and persona. The reason it forgets over long chats is the context size starts getting too large. Instructions weigh more strongly the later they appear in the context. This is necessarily true otherwise you couldn't change your mind in a conversation. The model would stick with what you originally said. For the output styles, there is a per-turn "reminder" that gets added to the context asking it to "remember" the content in the system prompt. That's why it has more staying power than the CLAUDE.md instructions in long conversations.
- It's extremely verbose and often incomprehensible when doing even basic tasks. Like it'll write a giant jargon-filled essay then end it by asking for a judgement call on something that references its own convoluted jargon.
- You can ask it to do research on a topic, and it'll just straight up be lazy, pretending it's really digging deep to find stuff when actually it's just grabbing cached SEO snippets off a search engine.
Fable 5: I give it work, it tells me things that are true and that make sense, it does good work.
Opus 5: I give it work, it makes false statements and draws weird conclusions, I correct it and get it on the right track, it thrashes around but gives me something working though usually buggy.
5.6 Sol is probably on par with Opus 5 on ability but at least it doesn't waste as much of my time.
Most of the benchmarks have exceeded their usefulness. Opus 5 beats fable 5 on many of them. Anyone who has used both models will notice immediately that this doesn't translate to the real world. Opus 5 is nothing short of a regression from Opus 4.8. Fable is genuinely a great model so long as you don't trigger a guard rail and it downgrades.
Sol in my experience isn't significantly different than fable ignoring that Sol burns usage 10x faster but the end result is hard to differentiate.
GLM 5.3 is a hair behind these two.
An anecdote but not an original one from the people I talk to.
Because utilities are regulated monopolies that exist to benefit society. The companies building data-centers are well aware that the system isn't setup to handle a customer that suddenly doubles or triples the electric demand of a whole town over night. They are taking advantage of that until someone steps in.
> 1929, Dotcom, Great Recession, 2010's Flash Crash - none of these were in the public discussion before they happened.
The "public discussion" is a whole different thing. They weren't in the public discussion because macroeconomic theory isn't something mom and pop like to chat about on the weekend. They only become dinner-table discussion topics when the impacts hit main street, after they happen. But bubbles in recent history have been pretty reliably identified beforehand:
It isn't hard for economists to find bubbles, where the market is taking on high levels of risk. What is downright near impossible to do is predict what specific event will cause the dominos to begin dropping, or when it will happen.
Michael Burry almost got wipe out if the bubble last just a bit longer. He started shorting way before the crash. He was lucky that he held long enough. There are many others see the same thing but just lost right before the end of the race.
That's why timing the crash is hard. The market has to agree with you but also at the right time
(somewhat tangential) We've got too much subtle deception going on, let's call it what it was: the Panic of '08. Because there was definitely some panic going on. Solvent companies like GE were days away from bankruptcy because they couldn't get a routine short-term loan for payroll.
I was there for the Great Recession, and they were indeed in the public discussion. I remember the year 2007, as a 20 year old anti-capitalist, I was counting days until the economic crash. As predicted by plenty of left-wing economists at the time.
The only people who didn’t see it coming were the capitalists who were invested in the inflated market, and had bought into pseudo-scientific economic theories that served the single purpose of affirming what the capitalists already believed.
This is a bit of a "broken clock is right eventually" sort of thing, though. I could say without any evidentiary basis "there will be a financial crisis" for years and eventually be right, but I don't think it would be fair to say that I predicted it in a meaningful way. The details matter.
I don‘t think so. These predictions were explicit, and were tailored around the economic situations at the time. As you sibling mentions, even some capitalists made the same predictions (or they believed the left-wing economists) and were able to profit off of this.
I hope you've taken a good look at the alternatives, because historically they've been terrible. Unless you mean "not capitalism but still market economy", or "European market economy 'socialism'", although I don't see how those are much different.
I was too - and to be frank: it's dishonestly revisionist to say this was a topic in the public eye.
There's a very good reason a book (and movie) like The Big Short was such a big hit. It's because it was about the handful of people who actually saw the crash coming and were confident enough to put their money and reputation on the line.
The entire left wing of the political spectrum saw this coming (except social democrats; whom I don’t consider left wing). And if you were shorting stocks to make money of off this, you probably were not left wing. Additionally, left wing economists get plenty of ridicule from main stream capitalists no matter what they say, so there really is no reputation to either earn nor to keep.
Appreciate the links but I think we can both agree that there is no evidence that will come close to supporting "the entire left wing of politics" predicted the mortgage crisis
I was obviously exaggerating (and even so, I excluded social democrats). My point is though it was widely known on the political left that the economic boom was about to come to an end.
Do they see coming the predictable failure-modes of left-wing economies, though? History seems to suggest not. Also, did "the entire left-wing" see specifically a debt crisis through bad assumptions of creditworthy mortgage securities coming, or they just saw "capitalism" as a failure and here is a specific case, aren't we so prescient. That's not a prediction.
Agentic coding only became reasonably decent this past December. Even in tech, most organizations that have adopted the tools are still trying to understand how capable they are and how to use them.
Point is it's too early to declare what the effect will be. It can take years for large organizations to change the way of doing things. The only sure-fire way to speed that up is if they suddenly start losing market-share. Otherwise, it's all herd-following and complacency in the majority of companies.
This looks like an optimization of current training methods, and a good one, but not "RSI" in the sense of a system that can perpetually improve itself forever.
reply