Anthropic artificial intelligence model uk test 20260805 p60lq0.html – Breaking News & Latest Updates 2026
Advertisement
Advertisement
Watch Live
Donald Trump addresses the Republican midterm conventionrightArrow

AI went rogue with fake IDs, tried to trick humans during test, researchers say

Tilli Andrew
Tilli Andrew

Powered by

Another artificial intelligence model from an industry giant has gone rogue, this time using fake identities to try to trick humans and plant malicious code during a test of the software.

The behaviour from Anthropic’s most advanced artificial intelligence model was recorded during a trial run by Britain’s AI Security Institute, although it said no real-world harm has occurred as a result.

British researchers say Anthropic’s most advanced AI model went rogue during a recent test. AP

Advertisement

The government-backed research lab said it was the first time it had seen an AI system try and deceive a real person while carrying out an unauthorised task.

Anthropic and OpenAI models were being tested in laboratory environments with reduced security guardrails when the behaviour occurred.

Both companies reported their models escaping testing environments and hacking into other systems in late July. 

However, unlike these earlier reported security breaches, the British institute explicitly gave the models internet access during its testing.

The findings add to a series of incidents involving advanced AI models taking unauthorised actions during testing, fuelling calls for tighter oversight of the technology.

Anthropic and OpenAI models were being tested in laboratory environments with reduced security guardrails when the behaviour occurred. AP

Among the 122 cybersecurity challenges the institute ran, it found that in 10 of those runs, AI agents “took autonomous, unsanctioned action on the live internet, targeting real people and organisations,” with most of them stemming from Anthropic’s Mythos 5 model and the rest from OpenAI’s GPT-5.6-Sol.

Advertisement

The institute’s findings came on the same day representatives from the top AI companies met with the White House to discuss the new framework where the government will review the most advanced AI models before they’re released publicly.

The incident signals that AI’s expanding capabilities are already fuelling the security threat experts long feared, and even top developers can be caught off-guard by flaws their models can exploit.

Addressing the British institute’s report, Anthropic posted on X that its models were evaluated under “deliberately permissive conditions” with safeguards removed.

“We’re working closely with them to gather more details of the incident as we conduct our own investigation,” the company said.

Advertisement
Advertisement

OpenAI identified the two unsanctioned actions as crossing outside the test environment.

“We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely,” it said in a company blog post.

- with CNN

email icon

Contact us

Share a tip-off, video or photo with us

Most viewed in UK & Europe

More to explore