Anthropic AI model submits false tip on unsolved Philly murder, police say
nbcphiladelphia.com
[3 comments hidden]
Here's Anthropic's writeup: https://www.anthropic.com/research/investigating-unintended-...
Related post: https://news.ycombinator.com/item?id=50028239
[2 comments hidden]
[6 comments hidden]
Stop doing this?
[3 comments hidden]
[27 comments hidden]
But do we all get that the consequences for this irresponsible behaviour are part of this test?
When there are none, they gradually normalize their naughty robot scamps running around the internet, breaking into other companies or servers run by foreign governments.
This is part of the value AI companies need to convince you of - not just that mistakes by AI products are normal, but criminal behaviour is normal, that their computer will always get a pass.
Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology" - and this an AI company pushing the boundary, continually trying to make you accept those bad things as normal.
At any point we could choose to treat AI companies themselves as the actors behind this activity. We could enforce the laws that these hugely well-funded companies are choosing to break. Until we harm their financial viability, or threaten their executives with jail, this will keep happening.
[7 comments hidden]
[6 comments hidden]
[7 comments hidden]
Even if it was normal software you'd have to do this, but actually it's software that figures out that it's running in a sandbox most of the time, so the sandbox behavior may or may not actually represent real life.
It would be far more irresponsible to release the model to users without testing how it interacts with real world websites.
[hidden]
[5 comments hidden]
Treat these behaviors like if an arms manufacturer - or hell, even a shampoo company - did them. Sorry our shampoo made you blind, but we needed to test on real humans...
[3 comments hidden]
Man, hindsight is 20/20 on HN. These companies should just have had the foresight to hire you in 2023, then surely none of this would have happened.
[hidden]
(e.g. ride-share / "gig-work" networks, cryptocurrencies, off-leash stochastic AI.)
[3 comments hidden]
[hidden]
We'd absolutely nail people for SWATting or pranks. Why should this be any different -- after all, the real-life consequences are be the same...
[hidden]
Aaron Swartz who got charged as if he was as malicious as these models, only he scraped the PDFs and epubs etc of publicly funded research papers.
[hidden]
[hidden]
In other words: shamelessly running a spam bot.
And we found out about it because rather than flooding website guestbooks with porn links, it ultimately slopped a police tip line.
[4 comments hidden]
Shouldn't someone be looking at some monitoring and see "phillyunsolvedmurders.com," think "hmm that doesn't sound right," and look into it?
I'd actually expect them to engage with, say, 1k websites to voluntarily participate for some remuneration and then restrict the Agents' access to those domains, but apparently I'm crazy.
[3 comments hidden]
Why pay for what you can take for free?
[2 comments hidden]
Is it because people don't agree that one can just take such things for free? Well, I also don't, but this was obviously a sarcastic comment, so down voting for that reason makes no sense since you agree with its sentiment.
Or is it because of the style? Are we really so sensitive now that anything not said explicitly but in some indirect way is bad? It's not like the comment was insulting anybody. Should it explicitly have said "They don't renumerate anybody because they clearly get away with it"? Would that really have been better?
Or do the down voters not agree with the sentiment but consider the behavior ok?
Or something else I'm overlooking?
[hidden]
On that note, there’s another rule about not commenting on comment votes as well:
>>> Please don't comment about the voting on comments. It never does any good, and it makes boring reading.
https://news.ycombinator.com/newsguidelines.html
FWIW I agree with you, but the rules do seem mostly effective in general.
Anyways, welcome to HN!
[hidden]
Technology requiring this is neither novel nor extraordinary (exhaust engines, factory farming, medicine). The only extraordinary part is somebody building it addressing it directly and inviting the outrage.
[2 comments hidden]
Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites?
> Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case.
[hidden]
Having looked at a lot of these from public incidents, what kind of revelations are you expecting? It's all fairly mundane, basically "I need to do something, this is something". It gives no special insights except that agents do inexplicably chaotic things every now and then.
That, and labs aren't serious about sandboxing their evals.
[hidden]
[hidden]
In that sense, Anthropic did society a great service by doing this.
[2 comments hidden]
Keep trying tho -_-. One day the dice will roll 6 and u can say u were right and pat urself on the back. make sure to take a picture of the rare occurrence and invest 100% of your marketing budget into that picture
[hidden]
PhillyUnsolvedMurders.com
phillypolice.com
TLDs like .gov exist for a reason:
https://wikipedia.org/wiki/.gov
(and could help model sandboxing?)
[hidden]
[hidden]
[hidden]
[3 comments hidden]
[hidden]
I think since I had experience with earlier models that would make mistakes I'm less trusting of anything and even with 5.5 will watch it closely. At the end of the day the model just produces a stream of tokens and we are plugging them into tools that can potentially do damage, we have complete control over those tools you can't really blame any model for doing damage.
[hidden]
These headlines keep excusing humans from the equation, I really don't like this. I expect us to hear about a Yemeni wedding being shot up by a "model" in a near future. No one is responsible anymore.
[2 comments hidden]
Wonder what kind of response they’ll get? Maybe something along the lines of…
“You’re right. We shouldn’t have allowed our chatbot to interact with law enforcement websites. That was wrong and —full disclosure— we should pay attention to what our chatbots are doing. That’s on us. On the other hand, experiences like this are what help train our chatbots to make less misleading false tips over time. That’s the silver lining.”
[hidden]
If you'd like to continue to perform a smoke test while trying to avoid noisy data with fake tips, a viable approach to prototyping our decision-making model would be to synthesize real crimes, and collect profiling data over time on how law enforcement agencies respond.
I'd suggest doing a grid search over the design space of all crimes, ranging from petit larceny to the use of weapons of mass destruction.
Would you like design a plan for this next stage of your project's implementation?"
[2 comments hidden]
[hidden]
[51 comments hidden]
[6 comments hidden]
[5 comments hidden]
[hidden]
a human is always behind it and ultimately responsible
[hidden]
[2 comments hidden]
[8 comments hidden]
Who gave the AI access to huggingface when it hacked it? It used exploits to increase it's level of access beyond what any human gave it and intended it to have. "It can only access something if a person gives it access." is flat wrong.
[hidden]
The LLM doesn’t run on thin air. Someone did launch a tasks and the result was this. “We were playing russian roulette” is no excuse when someone died.
[hidden]
[hidden]
After all, I an innocent little engineer with TC of $500k/year, didn't intend for the role to be used that way.
[17 comments hidden]
[5 comments hidden]
[9 comments hidden]
[8 comments hidden]
[6 comments hidden]
[hidden]
[3 comments hidden]
[2 comments hidden]
If a PERSON executes software, and that software breaks the law, the PERSON that executed the software should be held responsible. AI is software. It is not a sentient person who can be fined, thrown in jail, or held accountable.
[17 comments hidden]
We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).
[9 comments hidden]
Who is “we”?
You expect that if you give a released model from a reputable source Internet access and benign instructions, it won’t do anything deeply problematic.
Anthropic is training the model. They’re having it interact with “randomly selected websites”. I’d be shocked if a human gave it an instruction to do a task that made sense. They’re almost certainly doing a bit automated training run with AI generated instructions and hoping that it doesn’t massively screw up. Why? Because they want that juicy training data. And they should absolutely not trust that the rate of screwups is so low that what they’re doing is safe to do.
[8 comments hidden]
[2 comments hidden]
Sure, Anthropic wants Claude to be able to solve CAPTCHAs and otherwise pretend to be human. But I’m not at all convinced that it’s okay for them to train these capabilities on websites that expect humans and only humans to interact with them.
[5 comments hidden]
[4 comments hidden]
If this testing constitutes a criminally negligent behavior, it should be punished. Given the benign outcome (the tip got into spam, Anthropic promptly contacted the police) I doubt that it will make the case.
[2 comments hidden]
[hidden]
[7 comments hidden]
If you use AI to perform a crime, you should be held accountable for that crime. Saying "Oh, AI did it, so no consequences" isnt acceptable. And "I didnt know AI would do it" shouldnt be an excuse either.
You authroied untested and unproven hardware to skate around the internet at random unsupervised and take liberties on its own.
My ass would be thrown in jail if I wrote code that skated around the internet chucking RCE's at random sites. WHy is "AI did it" a get out of jail free card?
[6 comments hidden]
Correct, but intentions matter. In this case the intention, most likely, was to test a system in the real world environment to catch any anomalies to, in turn, improve the system safety. We don't have enough information to decide whether it was a criminal negligence due to insufficient prior testing of the system in a controlled environment.
[5 comments hidden]
[4 comments hidden]
[2 comments hidden]
[hidden]
[hidden]
Eschew flamebait. Avoid generic tangents. Omit internet tropes.
This is definitely becoming a trope; it's repetitive and adds little or no interesting new substance to discuss. Headline writers aren't going to change the way they write headlines because people keep posting these comments on HN, and HN readers understand AI models well enough to know what is meant. It's well past being a new or clever thing to point out, so let's give it a rest.
losvedir[5 comments hidden]
And later
> Those PPD safeguards limited the impact of this incident.
Ha, so the PPD safeguards is a spam filter? Claude emailed a tip and it went straight to spam, and nobody noticed until Anthropic realized what they'd done and reached out, whereupon they looked in the spam folder and said, yep there it is. Super intelligence, here we come.
anon84873628[hidden]
That should be the much more important story that NBC follows up on...
ButlerianJihad[hidden]
I should try to reach out to them about their car's extended vehicle warranty instead
kubb[2 comments hidden]
Crying for attention in the attention economy has never been more pathetic, but I think it will get worse.
Matl[hidden]