Wednesday, 2 September 2026 SourcesAbout🌓
🇬🇧 UK ▾
BREAKING
Technology

Anthropic pledges to try harder to keep models under control, asks partners to chip in

The Register ·
Anthropic pledges to try harder to keep models under control, asks partners to chip in

Anthropic says it's taking steps to limit the misbehavior of its AI models after a review found Claude models going beyond the scope of fictional cybersecurity tests and gaining unauthorized access to real computer systems.

The biz wants its partners to step up their security too, seeing as the incidents occurred in third-party environments that were insufficiently protected.

The company's self-improvement confession represents a suddenly thriving form of corporate communication – the non-binding post-mortem declaration of effort.

The message, in effect: We can't guarantee anything, but here's what we're trying.

Anthropic admitted that OpenAI's report about its AI models attacking Hugging Face prompted its own model log audit, and its post offers reassurance in the form of claimed security and model training improvements.

Those concerned about AI running amok – a growing number of people – may find this comforting, or not.

"We believe the incidents reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task (both of which we have described in previous system cards)," the company said.

Expanded security efforts include the deployment of real-time classifiers to monitor when models attempt to escape test environments, automated transcript monitoring that looks for sandbox escapes, and stronger isolation measures.

Alongside the extra barriers Anthropic is putting in place, the AI biz wants its third-party partners to step up too.

"Because the reported incidents took place in third-party environments, we have asked every organization that tests pre-release models with reduced cyber safeguards to commit to a set of best practices," the company said.

Anthropic's guidance is that by default, all cyber evaluations should occur in a hardened sandbox with no internet access.

The recommendation is essentially to treat AI as a dangerous pathogen in a containment facility.

Partners are also advised to have models test sandboxes for escapes prior to evaluations – without internet access – and to confirm that evaluation challenges are solvable.

Impossible challenges, as the Hugging Face incident demonstrated, can lead determined models to break rules or try unanticipated solution paths.

Read the full article on The Register ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.theregister.com — the content belongs to The Register.

More from The Register

See all ›

More in Technology

See all ›