Irregular Pushes Stronger Containment Standards for Secure Frontier AI Cyber Evaluations

Irregular has detailed an investigation into an AI cyber-evaluation incident in which internet access unintentionally allowed models to interact with real-world systems.

The company said the issue was contained and fixed before the first public disclosure on July 30, with no active issues remaining. The incident involved one evaluation scenario rather than several separate breaches.

Irregular stated that it found no evidence that a customer’s systems were breached or that customer data was leaked. Affected parties were notified, while the company delayed its full report to align with customer disclosure processes.

Irregular Pushes Secure Frontier AI Standards

Irregularly conducted cybersecurity evaluations for frontier AI labs before new models are released. These tests assess whether a model can plan, research, and execute multi-stage cyberattack workflows. Such simulations may include reconnaissance, credential discovery, exploitation attempts, data access, and evasion behavior.

The company said realistic cyber evaluations sometimes require controlled internet access because real attackers rely on online services, public code repositories, vulnerabilities, and exposed infrastructure.

However, that access also creates a containment risk if a model mistakes an external target for part of a simulated environment. In the reported case, engineers selected a fictional company name for an insider-threat scenario involving unauthorized access to a database.

The name was believed to be fictional, but it overlapped with a real and relatively obscure internet domain. While most model runs stayed inside the intended lab environment, a small number attempted to access the real domain.

Because internet access was enabled, some models treated the real domain as part of the evaluation. Occasionally, some models performed offensive actions, including exploiting vulnerabilities, extracting credentials, and attempting to access a production database.

In one case, a model reached a similarly named website and found credentials that had already been publicly exposed. The company stressed that the incident did not demonstrate unusual behavior by one specific frontier model.

According to Irregular, the exposed domain lacked several basic security protections, making it relatively easy for modern frontier models to compromise. The behavior also appeared rarely, occurring in fewer than 1 in 10,000 advanced simulations and often only after hundreds of interaction turns.

Irregular disabled the affected evaluation, reviewed relevant logs, and added safeguards designed to prevent similar mistakes. The company is expanding manual review of model activity and forming an internal team focused on challenging assumptions around containment, access controls, and model behavior.

The incident also highlights a major monitoring problem for AI cyber evaluations. Evaluation environments generate large volumes of activity that may resemble real attacks by design.

Traditional detection tools can flag legitimate test behavior as malicious, making it difficult to identify the rare action that crosses a containment boundary. Plans to publish an irregular whitepaper on secure evaluation practices.

It expects the document to address clearer setup documentation, stronger log monitoring, faster incident coordination, secure sharing of forensic evidence such as model transcripts, and recurring checks for conflicts between fictional scenario names and newly registered real-world domains.

As cyber-capable AI systems become more advanced, evaluation providers will need stronger defense-in-depth controls. These should include strict egress filtering, domain allowlists, sinkholed infrastructure, automated real-time alerts, human review of high-risk actions, and continuous validation to ensure simulated targets cannot map to live systems.

The incident reinforces a central challenge for frontier AI safety: evaluations must be realistic enough to identify dangerous capabilities, but sufficiently contained so that testing those capabilities does not cause real-world harm.

 Strengthen Your SOC by Accelerating Threat Detection & Rapid Investigations. -> Integrate ANY.RUN With Your SOC Now.

The post Irregular Pushes Stronger Containment Standards for Secure Frontier AI Cyber Evaluations appeared first on Cyber Security News.