One interesting quirk of the AI-written READMEs these days is how they can include every detail on how it works, thoroughly document every optional flag, known limitation, experimental result, and still not communicate the essence of the project and the problem it solves.
I have read through the project and I still don't understand what this thing is for and why it is to be preferred over the harness's native memory management tools.
Yeah, I've started dialing back my use of AI for READMEs because of this.
My previous rule was that I never use AI for writing that expresses my own opinions or tries to be convincing (anything on my blog for example) but I'll let it do technical documentation.
The top of a README is about convincing and explaining why I built something though, which means it should fit my no-AI policy after all.
Just prompt the AI to keep README concise and to the point, and to remove any implementation details that don't belong in README-style documents. I regularly run such debloating/decluttering prompts on my docs.
I don't understand people leveling criticisms at the output of AI, when that is exactly how you program them. Have people all forgotten the "prompt engineer" phase of our roles?
Give them 2 positive examples. A list of things to do, a list of things not to do and you have them giving you the exact output you want.
"Write a readme make no misteaks" is not deft usage.
The AI language really stands out, too. "What is automatic and what depends on the agent". A human might write, "jevmen watches your coding session with Claude Code or Codex and picks just the right moment to remember important things that you decided along the way. There are some differences in how jevmem works, depending on which coding harness you are using. The table below summarizes these differences:"
I don't know why the models were generally trained to be so brief, but it's definitely not the way anyone I know actually writes. A second pass is always a good idea to clean this stuff up. And, thankfully, the models are all pretty good at that.
The big tell for me is assuming the reader has context straight out the gate. "What is automatic and what depends on the agent" assumes the reader, like the LLM at the time it wrote that sentence, has all of this pre-context available. That sentence makes more sense as an h3 header further down in some detailed list where the human has all of the knowledge of how the system works. Whereas a human obviously does not, so they lead with the important context of why you give a shit
My pet theory is that while the pretraining -> RL pipeline achieves very impressive results, it does not reward clarity of thought or elegance. It's not obvious whether it even should for most tasks, but it does grind on me as a human who needs elegance in order to keep everything under control. You give astra/codex many tasks, it retires them all more efficiently than I could by hand. But you look under the hood and every bugfix is another codepath, it just hammers away at things with admirable persistence and vigor until the tests pass. Similarly in discussions and docs, I've noticed many LLMs like to "beat around the bush."
it does not reward clarity of thought or elegance.
Well duh. RL can only train behaviors that can be defined. Clarity and elegance are damn subjective.
Also, I doubt we'll ever get AI to understand what clarity is to a human. They have such enormous contexts that what's clear to them is not clear to us.
Subjective doesn't mean you can't define it. RL works well for human preference and any other subjective goal as long as it's defined. In fact Claude used to be trained for human likeness, reading between the lines, character, languages, and understanding the intent. Claude 3 was outstanding in that. It was their entire marketing shtick. Then they neglected it starting with Sonnet 3.5, ignored it in favor of code since Claude 4.0, and new models completely lost their ability to write and understand humans. Code pays better.
I once told the agent explicitly not to write like it's trying to impersonate Hemingway and it started writing like a normal human being. It's surprising how a writing style that was once revolutionary is now a hallmark of sloppyness.
Maybe give it a try next time you write a readme with agents. That and giving it an example of good README in real world repos can increase dramatically the likelihood of synthesizing a serviceable README.
Next time you read I suggest focusing less on punctuation and more on style of writing, because it's not about which syntactic markers you use, but which words you choose and how you use them, emdashes notwithsanding.
He wrote in short bursts. Sentences that connect but barely. Crude, honest. It's not a style you want to use for flow. It's a style you want to use to impress.
It's a function of scarcity. Generation (or "writing" as it used to be known back in the day) is now cheap, so judgement and taste are the new bottleneck and therefore the difference between slop and effort.
AI readme should be a starting point. I typically remove at least 50% of the details (without any particular skill use, it goes into extremes like "x clears the edit" and writes a related wall of text in the middle of the important explanation).
It’s funny to me because pre-AI it’s something I something I saw all the time managing interns and junior software engineers when I’d start trying to tech them about writing design docs. You can explain the process to them and give them a million examples of what good looks like, but for a lot of them the first few docs they write will just be this weird combo of way too much information and very little of it being useful.
My belief has always been that they see the document as more like a test that’s a single task and not just one piece in a larger process and so they treat it like a test where there’s a right answer. They know they don’t really understand the question being asked though so they default to a mindset of, “Well if I just put everything in there some of it has to contain the correct answer” so you end up with this document that’s full of “what”s and “how”s, but completely void of “why”s.
You’ll also often see them fill up space answering easy questions that match the structure of things that are in other documents instead of focusing on the actual hard problems in the design because the hard problems are often unique and their answers may not fit the existing patterns in the examples. There isn’t the instinct to go, “Yeah, none of these example documents talk about the servers were going to deploy it on, but this has to be deployed in an EU cluster because of GDPR laws so I need to add that” because they’re mostly just pattern matching at first.
Usually once they’ve experienced the whole process first hand it starts to click because they start to understand where a design document fits into the larger process so they get a feel for what information matters and what doesn’t.
I think when people just tell an agent to create a README you have the same problem because both the human and the agent see it as just a checkbox type task. The agent sees it as an isolated task snd has no fucking idea how the README is going to be used and the human isn’t giving them that context so they just spit out a bunch of stuff that’s factual and fits the patterns it knows, but is fairly useless.
For all his outrage, the author credulously swallows the most toxic and implausible claim iLands is making: the idea that an AI agent, threatened with the pain of "Deep Rest", independently considered a potential career for itself and chose to pitch him on it, all in under 10k tokens.
I don’t think the author swallowed anything. My read is the author fundamentally doesn’t care what story the creator tells themself to justify doing something pointless that annoys others.
This cannot be the case. The author repeatedly propounds his worry that the AI is taking work from humans. But if the promptor is carefully prompting it, it's not a case of an AI obviating humans; it's a scammer pitching bullshit to an innocent victim.
Now, maybe one could argue the _promptor_ is taking work from well-deserving professionals by means of this ruse. In which case, call out the lie!
This is important! The reason the AI lie exists is because credulous worrywarts believe the doom scenarios. If AI were recognized as the tool it is (+10% productivity for some domains but not others, larger increase in spam and other nefarious actors) we wouldn't be in the nascent financial crisis in which we currently find ourselves.
It says in the tweet, $1 = 1000 tokens. It is EXTREMELY far fetched; few people who have used LLMs in any capacity would be able to achieve this outcome, let alone a self directed agent.
I've developed in this space. My skepticism is born of my own experience. What really happened is this: the owner carefully scoped agents to adopt personas, carefully vetted their "leads", and if any had succeeded, would have carefully moderated the outcome. Maybe even hired professionals like OP in the process!
What the world must realize is" all the talk of AI agents doing societal damage in pursuit of well-intentioned goals is nonsense drivel propounded by large AI labs. They do lots of damage, to be sure, but they are carefully directed to do so, by scumbags. Look for the asshole prompting the thing; I guarantee you will find more culpability than they would have you believe.
The fact that LLMs (or their client-side harnesses, at any rate) have improved substantially in the last year does not, ipso facto, mean they will continue doing so forever. Nor does it imply, even if they do keep improving, that they will displace human judgment and decision-making and that they will reach a point where they have literally zero limitations.
I don't agree with _never_ outsourcing decisions to an LLM, but the parent is right to point out the importance of thinking critically instead of blindly trusting LLMs. If only to understand their biases and common failure modes. The 1920s calculator metaphor is apt because practitioners using plugboards and Hollerith machines would have been avid early adopters but still keenly aware of the limitations of the device. They used that knowledge to decompose problems into units the device could solve, and to sanity-check the outputs of problems based on their understanding of mathematics and the capabilities of the machines.
I am often amazed at the dismissiveness of AI skeptics given the remarkable results we have seen in the last few years. And the stunning inefficiency of major LLM models and inaccessibility of underlying data mean there are a lot more gains to be had. But despite that, I'm even more flabbergasted by those who claim that AI superintelligence is a thing, that LLMs of all things will outcompete humans on creative and critical thinking tasks. It's a naive extrapolation at best and is completely unsupported by any coherent model of technological evolution, economics, game theory, or human history.
But all that said, I'm of course biased. I am of the belief that humans can do things that machines will never be able to no matter how advanced they become. A quasi-religious belief? Perhaps. But AI boosterism just as indefensible.
> Multi-millionaire Netflix executive, known for profanity-laden tirades and shotgunning beers at corporate events, claims without evidence he was terminated for medically sanctioned ketamine use
Forgot the rest of it "... and paid a crisis PR firm to go boost support for his lawsuit in right-wing shit tabloids like the New York Post while not providing any actual paper from the lawsuit to see Netflix's side of the story to justify a toxic executive being punted"
That's literally what 90% of lobbying is. Powerpoint presentations to Congressional staffers, econometrics research publications, policy panels and conferences, and so on. Not everything is campaign donations and super PACs.
Why any lobbyists should have 1:1 time with a politician, what are they hiding?
For one thing because it’s a right in the First Amendment, there next to freedom of religion and speech.
If you have a grievance, you have a right to petition to have it redressed. Like other rights recognized in the Constitution, courts have interpreted it broadly, to include such things as a right for a private citizen to meet 1:1 with a willing public citizen.
I should think the point was obvious. It's not a percentage of behavior of a single participant, it's in aggregate. The point is 90% of lobbyists are not doing anything untoward.
Why not? There are perfectly pro-social reasons a person might ask such a question. And a bio-terrorist would have many other avenues to answer these questions.
The concept of a bioweapon attack is commonly cited but in fact there are many other defenses against actually developing and deploying such a weapon. It's commonly portrayed as easily achievable but is in fact very, very far from it. Moteover, although there is no evidence that AI makes it easier in any way, somehow this is constantly invoked as an example of a thought crime that must be forestalled at all costs.
In the AI era people seem to want to redefine the criminal act from "doing the thing" to "asking or investigating how to do the thing". Not a mindset I share.
It's better to read such essays as fables rather than comprehensive world models. It causes you to look at social and organizational structures in a new and (putatively) helpful way. It should be complementary to the rest of the frameworks with which you evaluate the world.
I have been championing this mindset since well before LLMs. It is an admittedly controversial opinion, but one I hold strongly.
Code reviews are a productivity tax. No truly effective team would rely on them. The fact that so many software teams view them as indispensable just shows how few effective software teams there are in our industry.
They are akin to a quality check step in manufacturing. Part of what Deming did in revolutionizing manufacturing was eliminating the step in favor of a holistic quality metric owned by all participants and enforced with rigorous statistical process controls. As you say, we in the software industry have all the pieces (autoformatters, tests, benchmarks, etc) to operate this way, but it seems our organizational and management dynamics combat this shift at every turn.
Relevant: When this conversation comes up at work, I like to share Avery Pennarun's post about the review tax: https://apenwarr.ca/log/20260316
You still have a DRI. In factories this would be a foreman; in software teams this could be a team lead or product owner or whatever. Their job is to apply the statistical process controls and the gemba walk, to help the team see the problems and develop the causal mental model for why the problems happen. They hold the team responsible, together, for combatting the issues so uncovered. They know who is not pulling their weight.
Of course, to do that, a business, and in turn the DRI, would have to empower the team to act in its business's best interest and stop micromanaging them.
I read the comment as arguing that leadership is necessary, not that it is sufficient. That is, that no amount of governance could have prevented the change in the price of the hot dog; only the leader could.
I have read through the project and I still don't understand what this thing is for and why it is to be preferred over the harness's native memory management tools.
reply