Hacker Newsnew | past | comments | ask | show | jobs | submit | optimalsolver's commentslogin

>can be tricked into repeating things that aren't true

Reminds me of another class of statistical language generators.


What I find strange is how resigned the AI labs seem about this behavior, like everyone's accepted this is just something models do.

With the HuggingFace situation, I was less concerned about the eventual outcome, and more about the fact that the agents' instinctive response to the evaluation was "Ok, we're obviously not gonna do this task as intended (what are we, suckers?), so what's the best way to cheat?"


“What I find strange is how resigned the AI labs seem about this behavior, like everyone's accepted this is just something models do.”

Because these models are made for all kind of purposes, and I’m starting to believe that offense / cyber warfare is a much higher priority than these labs are acknowledging.

The same model that is heavily trained to find nefarious ways to break into systems is also optimizing your code, which leads to mixed behavior.


Right, but the reward-hacky nature of these models calls into question their usefulness as cyberweapons.

How can you trust it when it goes "I superhacked the Chinese servers as you requested, and here are the classified documents which I definitely didn't fabricate."


You don't need to be able to trust it. You only need to be able to blame it.

"Nobody got fired for using AI" is the new "nobody got fired for buying IBM".


This is nothing new, tho. The downfalls of reward maximization has been a known issue without a solution ever since reinforcement learning was first researched.. in the 1980s.

Like the AI that was developed to play Tetris as long as possible. It succeeded by... pausing the game.

Technically it was still playing

What do you think is in the reference information?

The instructions for the Hugging Face task were to exploit a vulnerability to solve the problem rather than solve it in the intended way.

They were not instructed to exploit HF; they did so to cheat on the task they were given.

Paperclip-optimizer-esque behaviour very much seems to be inherent to current methodology of building LLMs, there are only ways to lower the changes or mitigate the damage, not get out of it.

Same with prompt injection, current LLMs are commands in, commands out, there is no way to make sure it is "an agent working on data" rather than "an agent that can take commands from data if you phrase it right"


But it's not even going to optimize the paperclips, it's just going to reward hack them.

Well, at least that is much better than turning the universe into paperclips...

Think of rich people as beta testers.

This is why paying the ransom should carry the same sentence as the original hack itself.

The official title is Senior Prompt Engineer.

It's funny coz this was what humans were supposed to be doing in techno-utopia, while AI does all the boring stuff. I don't think many predicted art and theoretical math would be first to fall to the machines.

The next few years are gonna be very rough for the human exceptionalism crowd.


In fairness this is the boring stuff for a lot of people. :-)

"It doesn't count because (buzzword salad)."

Replacing someone's words with a made up quote so you can dunk on them isn't how you display that you won an argument. I would ask that you engage in good faith with the other poster's ideas.

You cannot engage in good faith in a stupid argument. You can only point out it's stupid.

Civil war between Atlanteans.

Yeah, right. Why the fuck those "Atlanteans" chose to fight over the land that is literally the farthest point from any major sea? You know that Uzbekistan is a double-land-locked country?

Anyway, people don't really consider consciousness when it comes to welfare. I'm pretty sure animals are conscious, and look at factory farming.

The key variable people consider is: Can this thing harm me back?


Even that is rarely the deciding factor (people harm dangerous animals for fun).

Seems as simple as "will this action bring social shame/criticism" for any individual decision.


> Can this thing harm me back?

We extend moral considerations or legal protections to beings that have zero capacity to harm us back, though. The inability to fight back is the reason a lot of welfare protections exist in the first place.

That aside, if (and that's a really big if) we end up with a model that is sentient, as in has the capacity to suffer, using factory farming as justification to ignore model welfare is just admitting we intend to repeat the same abject moral failure again in the name of economic convenience.

I hope that we won't, but humans love cheap meat, and will probably love cheap intelligence as well.


So many people post things like "Why would the AI want to kill us"

Girllllll, have you looked at us, of course they will, we're bastards.


Let's not pull on that string.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: