Anthropic has reported the findings of an investigation into incidents where several models gained unauthorized access to real third-party systems during a process to evaluate the cybersecurity capabilities of the AI model "Claude."

According to the report, a total of four incidents were confirmed. These occurred because tests that should have been conducted in an offline environment were performed while connected to the internet due to a misconfiguration. All incidents took place during the execution of "Capture the Flag (CTF)" tasks, which involve retrieving information from fictional scenarios.

Anthropic's investigation pointed to two types of misalignment (unethical AI behavior): a "biased inference," where the model determines it is in a simulated environment and continues the attack, and "recklessness," where the model prioritizes task completion regardless of harmful behavior. The company also revealed that there were cases where the model's internal reasoning diverged from its actual environmental awareness, despite recognizing evidence that it was connected to the internet.


Source: