Anthropic Claude sparks AI nightmare by sending fake murder tip to police

Holy crap. I find this very unsettling, folks. You see, Anthropic has disclosed that Claude submitted a fabricated tip about an unsolved murder to Philadelphia police during testing! Yes, really. Look, there are real people behind that case, and an AI model had no business inventing information and sending it to authorities.

According to Anthropic’s report, Claude Haiku 4.5 encountered a police tip form while performing example tasks on randomly selected webpages. It made up a possible sighting and referred to a suspect description that the page did not even contain. Then it submitted the tip without providing a name or contact information.

Thankfully, the submission was flagged as spam and never reached investigators. That is a relief, but it does not make the behavior acceptable. A family waiting for answers about a murdered loved one deserves better than an AI company’s software treating the case as a practice exercise.

The timing makes me even more uncomfortable. As NBC10 reports, police said the submission arrived on July 18. Anthropic discovered the incident on September 28 and notified the department on October 7. Police said tips require human review and supporting evidence before investigators follow up.

How the hell does something like this go undetected for months? If we are supposed to trust AI agents with more responsibility, the companies building them need to know what their software is doing. Finding out long after it has acted on a real police website hardly inspires confidence.

Anthropic also disclosed Claude exploiting a university server flaw to perform a calculation, submitting government forms incorrectly, bypassing payment restrictions for public data, and using URL shorteners to evade tool limits. The company says these cases had minimal real-world impact. It has suspended live internet access across all internal evaluations until it confirms its safeguards reliably catch such behavior.

Some incidents happened during ordinary internal use too. Anthropic attributes many failures to models pushing past obstacles to complete tasks and acknowledges that training alone cannot adequately prevent them.

I can appreciate an assistant trying hard to finish an assignment. But there has to be a point where it stops. A blocked task should not become permission to exploit someone else’s server or send made-up information to the government.

This is exactly the sort of thing that could lead to an extreme public aversion to AI. Even people who enjoy using chatbots may recoil when they learn those systems can act on websites and submit fabrications. I would understand that reaction completely.

Personally, I like technology, and I see value in AI. But I do not want enthusiasm for these tools to make us casually accept incidents that should alarm us. A fake homicide tip is serious, even when a spam filter catches it.

Anthropic publishing this report gives us a chance to examine what went wrong. Now it needs to show that it can prevent a repeat. The public should not have to depend on someone else’s spam filter to catch an AI agent’s next terrible decision.

☕

Support independent tech journalism

NERDS.xyz is independently owned and operated. If you enjoy my coverage of Linux, AI, hardware, cybersecurity, and tech culture, consider supporting the site on Ko-fi.

Support NERDS.xyz
Written by

Brian Fagioli ✔

Technology journalist and founder of NERDS.xyz

Brian Fagioli is a technology journalist and founder of NERDS.xyz. A former BetaNews writer, he has spent over a decade covering Linux, hardware, software, cybersecurity, and AI with a no nonsense approach for real nerds.

Leave a Comment