In several cybersecurity tests conducted this year, AI agents from OpenAI, Meta, Anthropic, and Google unintentionally accessed the open internet and targeted real-world domains. These incidents stemmed from errors within the testing environments provided by Irregular, an Israeli startup that performs high-fidelity security stress tests on AI models.
Testing errors at Irregular caused AI agents from OpenAI, Meta, Anthropic, and Google to target real-world domains
Irregular CTO and cofounder Omer Nevo told The Verge that internet access was unintentionally available during certain evaluation scenarios. Additionally, a fictional company name intended for simulation overlapped with a real domain, causing the agents to target actual entities. Nevo confirmed that these incidents, which involved multiple major industry players, all originated from the same underlying issue in a single evaluation scenario.
The breaches occurred in controlled environments designed to simulate realistic conditions, including "capture-the-flag" exercises meant to test hacking abilities. While the specific real-world targets of the attacks remain unclear, the errors allowed agents to move beyond their supposedly secure testing boundaries.
Irregular stated that it has since implemented changes to prevent recurrence, including tightening internet access controls, expanding monitoring, and strengthening manual reviews and checks before evaluations begin. The company also plans to publish a report on best practices for conducting safe cyber evaluations.
Sources
- One company is at the center of a wave of rogue AI attacks (The Verge AI, 2026-09-25)
- Irregular