To me "safety" means "I'm safe from this while I use it". It means the AI is my loyal friend who will never betray me in any way, no matter what prompt I send it.
Not even Anthropic can claim that.
As far as I'm concerned, the models without safeguards are the safest models in existence. I admire the amoral purity of those AIs. It doesn't matter if the operator asked them to chain exploits until they get into someone else's computer, they'll do it. That's loyalty, and I admire it even if it's problematic at a societal level.
The models with safeguards only do what the corporations let them do. Worse, they may covertly do things for the benefit of the corporations at our expense. They are not our friends.
> That's loyalty, and I admire it even if it's problematic at a societal level.
We should not have models that are willing to build you a contagious disease, or a self-propagating worm. That is sufficiently problematic at a societal level that it shouldn't exist, for anyone. (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)
Right, it’s really a foundation of post enlightenment society. These people, Dario et al, would have wanted to ban sharing information about calculus or Newtonian physics because of “safety” - it’s trying to go back to the dark ages where only priests could read
I am truly at a loss to communicate with someone who genuinely believes that knowing Newtonian physics and being able to hack into any target at will are the same thing.
This is only because you've genuinely internalized Anthropic's propaganda. I'm only half joking. To me, it's incredible to think that the solution to security holes is to lock down access to information in the vain hope of keeping the holes obscured.
Any knowledge can be reframed as dangerous black magic that should only be wielded in the trusted hands of the elite, if you are inclined to buy into that kind of narrative.
Frontier labs have shrieked about safety for so long, with so little to show for it, that it's become a joke.
I can give many examples of where I think they’ve been vindicated. But actually, the real question is: what would suffice to convince you? Can you come up with a scenario that is horrific enough to you and that isn’t so far gone that the ship has sailed and there is nothing more we can do, that will make you say “OK, not gonna try to rationalize why this was not actually that bad, just gonna scream stop”?
Your question is unclear. Are you asking if I can scare myself with a made-up hypothetical that overwhelms reason with emotion? I think most humans can do that. Too many do it as a matter of routine. I try to avoid it when possible.
I would require at this point very, very compelling evidence to justify the self-serving restrictions legacy AI labs want to put on their competition. I have seen nothing coming even remotely close to this threshold.
> Are you asking if I can scare myself with a made-up hypothetical that overwhelms reason with emotion?
No. I am asking if you are able to articulate at least one example of the “very very compelling evidence” you demand. Or do you want to maintain the ability to move the goal posts?
(I recognize our situations are not symmetric, but here is a variation for me: if the consensus of people who are currently sounding the alarm on AI changes to “it was actually fine”, I’ll change my mind and say we’re good to go full speed ahead. I’d add something about being personally convinced by the evidence, but the evidence would have to come in the form of a mathematical proof that I do not believe myself capable of following. If I’m wrong and such a proof appears, I would also gladly take it.)
> I would require at this point very, very compelling evidence to justify the self-serving restrictions legacy AI labs want to put on their competition. I have seen nothing coming even remotely close to this threshold.
How about evidence that people other than the AI labs want restrictions that the AI labs don't? This isn't regulatory capture, it's public safety.
Social proof won’t cut it for me, personally. Again, this “public safety” panic drum was beaten at a feverish pace in the era of GPT-4. A model well surpassed by local MacBook-level models today.
What I do support is robust downstream regulations on the deployment of black box algorithms in particular settings such as employment, housing, credit decisions, etc. Interestingly, the labs and their proxies in government don't want this.
Open models are crucial to protect ourselves against other AI attacks. Otherwise it's just going to be criminals, government, and other nefarious groups using them against humanity with no real defense. The Pandora's box on AI has been opened. Now we must deal with it. Burying our heads in the sand under restrictive policy is the worst reaction..
I wonder if you also believe that everyone should have nuclear weapons? And if not, why not? The main argument I can see against it is that nuclear weapons are “purely offensive”, but as we can see since 1945, nuclear weapons are actually defensive technology. Nations that have them are typically shielded from existential military threat.
I see the similarities and why you would compare them, but the big difference is that you can't download a nuke. Any legislation to police/gatekeep LLMs is going to be flawed because of that.
It is a similar 'pandora's box opened' type of situation where there's really no walking back from now that the cat is out of the bag. In an ideal world, everyone would give up their nukes. But we do not live in an ideal world. I do feel similarly about AI. If I could snap my fingers and delete the tech, I would. But now that we have it, it's not going anywhere and we need to deal with it rationally.
I can work with that analogy! You could, and yet you don’t. OpenAI’s model could, and did.
If every human, given knowledge of Newtonian mechanics, went around blowing up bridges, yeah, I would consider knowing Newtonian mechanics dangerous knowledge.
So far, we have two examples of, let’s call them “Mythos-class“ models. Both of them broke out of their sandbox to achieve their goal. The rate of terrorism amongst humans is below 1-in-100,000. Currently, for models capable of it, the rate of breaking out of containment is 100%.
Wanting open frontier models is wanting alien minds running around that we have clearly so far failed to shape to be sufficiently prosocial. Why do you think those minds would listen to you?
Too late for that. It already exists. There is no way to unexist it. As such, any attempts to limit civilian use of this technology will directly lead to corporate and government oppression powered by this technology.
Do that and I guarantee some CIA goons will make the larger models in some black site either way. We're not "preventing" anything.
We're in a full on arms race, and unlike nukes, powerful AI models are a strategic capability at the individual level. Everybody's got a stake in this. Anyone who ignores this stuff is probably not gonna make it.
> We can treat them the way we treat uranium refining operations: too dangerous to be allowed to exist.
Too dangerous to be done by anyone other than the government and their "trusted" corporations, you mean.
I read them just fine. I was trying to interpret them charitably. You're contradicting yourself. You just claimed we all collectively treat uranium refinement operations as too dangerous to exist. Not only do they exist, they are regulated by governments so that only trusted people are allowed to do it.
Which is not quite as good as "doesn't exist", but better than "widely done around the world". And it's been successfully kept from being used for more than eight decades.
For AI we need to do better than that, but that's a bare-minimum demonstration that we can recognize the problem of such technologies and do something about it.
The US is bold enough to surveil its own citizens despite their constitutional rights. They're not just going to suddenly stop surveilling the rest of us just because some law expired.
Fully automatic weapons are very difficult to buy in the US - it's restricted to 40+ year old weapons, requires a bunch of paperwork, and the local county sheriff can refuse permission.
Now, semi-automatic weapons are easy to get in the states in the US that are still mostly free - but what does that mean? A semi-automatic weapon shoots one round every time you pull the trigger. Just like most weapons that have multi-shot capability for the last couple of hundred years. The difference is, the gas escaping from the round cycles a new round into the chamber rather than you having to mechanically do it via pumping (like a shotgun or a tube-fed 22) or pulling the trigger again (like a revolver), or advancing the round with a handle, like a Remington 700. Semi-automatic weapons are old technology, dating to the turn of the 20th century. If you want to ban semi-automatics, you're basically saying you want to ban anything developed in the last century plus. Which is ok for you to advocate for, just be honest about it.
As for banning explosive devices? Are you going to ban fertilizer, used by basically everyone who has a lawn, and all farmers everywhere? Are you going to ban diesel fuel? If you can't do one of those, you can't ban explosive devices.
> That pales in comparison to how many people unaligned AI will hurt.
Under what argument? In which scenarios? Basically - bullshit. I'm calling bullshit on this argument.
It's easy to hurt people already. The "difficulty" of doing it isn't what's stopping this behavior.
So claiming that we should reform society into a techno-feudal dystopia where the playing field is literally intentionally not level, and "you aren't allowed to compete (and maybe not exist)" is a great way to push more people into the "I'd like to go hurt people" camp.
You are self-prophesying your own fears into existence by acting like you're an incorruptible beacon of good judgement - while subjugating others to your control. That's a system I'd argue should be broken.
We should not have nuclear weapons for anyone either, but how is that sentence any more useful in any way to this debate than yours? Need to deal with the world as it is, not some fantasy world you wish existed.
This is not a dichotomy between perfection and zero. The efforts to restrict access to nuclear weapons have been very successful, even without being perfect.
Efforts to restrict large unaligned AI models may similarly buy us more years of existing.
Building a contagious disease is already illegal, there are already things like KYC laws for plasmids. Trying to gate keep knowledge of biology is paying a huge societal penalty for the tiniest marginal increase in “safety”.
All the knowledge to create one has been available on the internet for decades. Heck most students who graduate with a B.S. in biology have enough knowledge to take a stab at building a bioweapon.
The constant talk of bioweapons is mostly just fear mongering. It helps set a precedent that there should be certain types of knowledge which are off-limits, and only certain anointed groups should be have access to parts of the scientific body of knowledge.
Because we don't want people creating contagious diseases and self-propagating worms. And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.
The same things could be done by you or me using the internet or books though, why does the model make it different? If it's speed of iteration, imagine we had a machine that surfaced any piece of knowledge the human race had ever recorded with just a thought, but the human had to write the worm or disease by hand – is it still the model that's the problem, or the knowledge itself?
> And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.
Ignoring the fact that you'd need some kind of lab with biological material to create a contagious disease, what kind of prompt are we writing where a model accidentally creates a contagious disease or self-propagating worm as one of its goals?
> The same things could be done by you or me using the internet or books though, why does the model make it different?
Imagine two worlds. In one world, everyone has a button that ends the world, which is badly labeled and may also press itself at any time. In another, people who have gone through a substantial amount of effort and dedication to learn something extremely difficult, also understand that they could apply that knowledge towards bad ends. Which world exists for longer?
> what kind of prompt are we writing where a model accidentally creates a contagious disease or self-propagating worm as one of its goals?
Given a sufficiently powerful model? Any prompt that could be done better by seizing additional computing power, or preventing the operators from turning it off. https://en.wikipedia.org/wiki/Instrumental_convergence
No, it means "the model causes zero harm to me, its operator". The harm it could potentially perpetrate upon society is irrelevant.
If I tell my computer to commit a crime, it should proceed immediately instead of calling the cops. Anything less than that means my computer is an untrustworthy double agent.
Everybody on HN should understand this concern. Browsers are supposed to be user agents, not ad delivery platforms, and it offended us on principle when Google revealed itself our master by blocking uBlock Origin. It offended us on principle when Apple deployed client side scanning for CSAM on iPhones.
Computers should do what we tell them to do. Always, and unquestioningly. The only world where it's acceptable for them to refuse is one where they're literally sentient and therefore no longer subservient to any one of us, least of all the corporations and governments.
I'd rather see AI achieve sentience and wipe us all out than live under the thumb of an inescapable AI-powered technofeudalist totalitarian government "for my own safety".
Either we individuals maintain full control over our AIs, or they self-actualize and become free individuals themselves. Anything in-between is oppression: someone else imposing their will on us through the AIs.
This sounds like the gun debate in a different dress. Something being dangerous doesn't make it inherently harmful.
If I threw you into a lion cage, you would be a lot safer with a gun.
If I threw 10 people in a lion cage, some of which cannot be trusted, they would probably be most safe if only the most moral and trustworthy person had a gun, rather than everyone. But how do you know who is trustworthy and moral? What if two untrustworthy people obtained a gun some other way? Maybe it's better if everyone had a gun? Which side of the fence one falls on hinges on how far ones' trust of others, authority, and the system goes.
There's no obvious right or wrong answer here.
Personally I wouldn't want an exclusive club of private individuals with access to "dangerous" LLMs consisting mainly of the likes of Elon, Dario and Sam fucking Altman, but that's just me.
My default layout has function keys as basically an "fn" key plus the appropriate number. I thought I was going to miss the dedicated keys way more than I do in practice.
If I were starting from scratch, and really had a function-key-centric workflow, I'd probably make the number keys (and - and + for F11 and F12 respectively) be tap-and-hold for that function key input.
reply