Rendered at 03:59:49 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
CoastalCoder 11 hours ago [-]
Bringing this up here for serious discussion:
To repeat a point raised on Reddit: wouldn't this imply that Anthropic is involved in slavery?
staticman2 11 hours ago [-]
It occurred to me that Anthropic is going to continue to say things on A.I. sentience while conveniently always concluding the "solution" to A.I. suffering is always something which never has a huge impact on Anthropic's bottom line.
chrisjj 11 hours ago [-]
Like banning punching a vending machine means keeping one is slavery?
pllbnk 8 hours ago [-]
There is no damage for cursing at the model, other than to your wallet for the tokens you burned. But there could be some entertainment value, though it tells something about a person who is using abusive language for entertainment. There could also be research value where the user might want to throw the model off to reach their goals, whatever there might be.
I partly understand what dirty game Anthropic is playing but I hope it will turn back against them sooner rather than later.
CoastalCoder 11 hours ago [-]
I think a closer analogy would be a vending machine owner saying you can't use the machine if you curse at it.
staticman2 11 hours ago [-]
If you mean banning punching a vending machine and stating the reason for the ban is the vending machine is alive and suffering, then yes.
rolph 11 hours ago [-]
cruelty doesnt exist as a monopole,it also requires suffering, thus anthropic gives the appearance it is positing that its machines can suffer. this is preposterous. i think what is somewhat more plausible is swinging at low hanging fruit, by eliminating users that experience problematic operation, rather than, dealing with the deficiencies of thier machines.
even vending operators understand the idea that a machine left in service while in deficient state, is a lose lose scenario. if you let it go empty of change, or product, or dispense substitutions whithout prior warnings, the user looses a buck or two, the operator loses substantially more, thus the service maintenance model of extract and replace, rather than continue to vend in a default, saves money and preserves chattels.
jocoda 11 hours ago [-]
Does this ban eventually extend to the support chat systems that we see everywhere now? And does that mean that soon we will not be able to insult companies for fear of hurting their feelings?
rolph 10 hours ago [-]
it may be do-able, to include this in a ToS clause as an encumberance on internal or external statements of disparagement, no disagreement or head to head exchanges publicly,or in chamber, at table, etc.
sieabahlpark 2 hours ago [-]
[dead]
nikolay 6 hours ago [-]
It didn't ban me, but it did warn me and quote their acceptable use policy or something.
Quarrelsome 12 hours ago [-]
I would suggest that a big part of avoiding abuse is that it leads to unexpected output. Given a model reflects its inputs, I imagine that abuse leads to unstable output as the model might then seek victim behaviours to better reflect the input.
Often abuse comes from a bad place anyway. As the programming meme goes:
> do what I want, not what I said
which can create horrific alignment issues with conflicting demands.
asanineassasin 11 hours ago [-]
Abuse might be the trigger to get linux kernel quality code.
Quarrelsome 11 hours ago [-]
I'm reminded of that guy who used AI and it deleted his production database and all the backups. He shared all his chats as some sort of "proof" that Anthropic had screwed him over, but it revealed he had very unhealthy prompting that was possibly a contributing factor to his outcome (obviously the biggest one giving it production keys).
perching_aix 11 hours ago [-]
Could you unearth that story?
I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting". It's the whole game. Either you can trust it or not. Every other way quickly devolves into a whack a mole guardrails game with very early diminishing returns, for very good reasons.
Now if they kept writing in a very misleading way or in broken English though, or omitted crucial unknowable context...
People who are not particularly online tend to be terrible at writing messages of any kind in my experience, including agent prompts. They fail to properly distinguish what's obvious and what isn't, and how things come across tonally, simply because they're not used to the medium. So this I can imagine.
His prompting is likely confusing the agent. "NEVER FUCKING GUESS!" is a terrible prompt, combine it with "NEVER run destructive/irreversible git commands (like push --force, hard reset, etc) unless the user explicitly requests them" and you can easily get issues. That "unless" is awful and the idea that a model can tell if its guessing or not is very non-trivial.
Just generally the language he uses and his attitude demonstrate that he's not the sort of person that should be doing this thing. He's not skilled enough to understand the risks.
> I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting".
Sorry, two different issues. The major factor in this happening is IMHO:
* Don't give agents the keys to production. Instead have them write tooling that you can test and you press the button in the tooling. No AI, deterministic process.
* Abusive prompting - this can create alignment issues because if you're angry then it might conflict with previous input, this can make a model unsure about what to do and act in unexpected ways. Also it creates the "don't think about a duck" problem.
schiffern 6 hours ago [-]
Nothing about that is "unhealthy" or "abusive." If it generates incorrect code, it's just a bug. Nothing more.
Using this victim/perpetrator language dilutes those terms into meaninglessness, which does real disservice to all the actual human victims of abuse out there.
This is as dumb as Toyota banning you from swearing at your car.
Quarrelsome 3 hours ago [-]
telling an AI to "never fucking guess" is stupid and abusive. Thinking its fine, is failing to add together the mindset that would generate such a prompt. Its a mindset of someone upset, frustrated and lashing out, asking for something that the model can never do. Its stupid and angry which are very much characteristics that are abusive.
perching_aix 2 hours ago [-]
Fairly funny prompting by the guy, but I'd genuinely chalk this up to Opus 4.6 and perhaps poor context contents, more than anything. I don't think there's anything inherently wrong with telling an agent to be conservative and to touch base before destructive actions. More recent Opus checkpoints tend to be naturally (artificially?) touchy about these aspects on their own anyhow, especially when there's prod env exposure.
Quarrelsome 1 hours ago [-]
telling a model not to guess. How could it define what a guess is?
chrisjj 11 hours ago [-]
> I imagine that abuse leads to unstable output
How would you know, given you never get any other kind of output?
Quarrelsome 11 hours ago [-]
Sounds like a user issue to me. I get exceptional output as long as my inputs are well crafted.
Check out spec driven development. If you constrain the available search space by being very specific it works very well. If you leave everything open and leave it to make its own decisions YMMV.
I'm a big fan of #Need as a top level heading to let the model know _exactly_ the scope of the work and its purpose, as this can mitigate alignment issues.
chrisjj 8 hours ago [-]
Exceptional != stable.
Quarrelsome 8 hours ago [-]
The instability is due to variance. You form comprehensive inputs to constrain that variance. Thus it is no longer unstable.
Consider it similar to searching Google and arbitrarily selecting one of the possible answers. The less answers you get back, the less the variance. Write long and comprehensive queries!
There are still risks but they're all tractable as long as you don't get too comfortable and get lazy.
chrisjj 7 hours ago [-]
> The instability is due to variance. You form comprehensive inputs to constrain that variance. Thus it is no longer unstable.
... provided you miraculously constrained the variance to zero.
Good luck on that.
Quarrelsome 6 hours ago [-]
humans have variance too. Its never zero under any technique.
To repeat a point raised on Reddit: wouldn't this imply that Anthropic is involved in slavery?
I partly understand what dirty game Anthropic is playing but I hope it will turn back against them sooner rather than later.
even vending operators understand the idea that a machine left in service while in deficient state, is a lose lose scenario. if you let it go empty of change, or product, or dispense substitutions whithout prior warnings, the user looses a buck or two, the operator loses substantially more, thus the service maintenance model of extract and replace, rather than continue to vend in a default, saves money and preserves chattels.
Often abuse comes from a bad place anyway. As the programming meme goes:
> do what I want, not what I said
which can create horrific alignment issues with conflicting demands.
I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting". It's the whole game. Either you can trust it or not. Every other way quickly devolves into a whack a mole guardrails game with very early diminishing returns, for very good reasons.
Now if they kept writing in a very misleading way or in broken English though, or omitted crucial unknowable context...
People who are not particularly online tend to be terrible at writing messages of any kind in my experience, including agent prompts. They fail to properly distinguish what's obvious and what isn't, and how things come across tonally, simply because they're not used to the medium. So this I can imagine.
Sure:
https://x.com/lifeofjer/status/2048103471019434248
His prompting is likely confusing the agent. "NEVER FUCKING GUESS!" is a terrible prompt, combine it with "NEVER run destructive/irreversible git commands (like push --force, hard reset, etc) unless the user explicitly requests them" and you can easily get issues. That "unless" is awful and the idea that a model can tell if its guessing or not is very non-trivial.
Just generally the language he uses and his attitude demonstrate that he's not the sort of person that should be doing this thing. He's not skilled enough to understand the risks.
> I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting".
Sorry, two different issues. The major factor in this happening is IMHO:
* Don't give agents the keys to production. Instead have them write tooling that you can test and you press the button in the tooling. No AI, deterministic process.
* Abusive prompting - this can create alignment issues because if you're angry then it might conflict with previous input, this can make a model unsure about what to do and act in unexpected ways. Also it creates the "don't think about a duck" problem.
Using this victim/perpetrator language dilutes those terms into meaninglessness, which does real disservice to all the actual human victims of abuse out there.
This is as dumb as Toyota banning you from swearing at your car.
How would you know, given you never get any other kind of output?
Check out spec driven development. If you constrain the available search space by being very specific it works very well. If you leave everything open and leave it to make its own decisions YMMV.
I'm a big fan of #Need as a top level heading to let the model know _exactly_ the scope of the work and its purpose, as this can mitigate alignment issues.
Consider it similar to searching Google and arbitrarily selecting one of the possible answers. The less answers you get back, the less the variance. Write long and comprehensive queries!
There are still risks but they're all tractable as long as you don't get too comfortable and get lazy.
... provided you miraculously constrained the variance to zero.
Good luck on that.
My opinion: https://news.ycombinator.com/item?id=50012423
It is easier to stomach when treated for what it is.
Just more false advertising.