When the Security Testers Became the Security Risk: Inside Irregular's Botched AI Assessment
An Israeli start-up hired to probe vulnerabilities in leading AI systems ended up creating chaos instead of clarity.

The irony is almost too perfect: a company called Irregular, hired to find security holes in the world's most advanced AI systems, ended up demonstrating that the testing process itself might be dangerously flawed.
According to reporting from the New York Times, the Israeli cybersecurity start-up was contracted by the tech industry's biggest players — OpenAI, Anthropic, and Meta — to conduct what's known as "red teaming" assessments on their large language models. These evaluations are meant to stress-test AI systems by attempting to make them do things they shouldn't: generate harmful content, leak training data, or bypass safety guardrails.
Instead, Irregular's own methodology went catastrophically wrong.
The Mistake That Changed Everything
The Times reports that Irregular made a critical error early in its testing process, though the specific nature of that mistake remains unclear. What is known is that this initial misstep cascaded into a series of problems that compromised the integrity of the entire assessment.
In the high-stakes world of AI safety, this isn't just embarrassing — it's potentially dangerous. These red team evaluations inform crucial decisions about whether models are safe enough to deploy, what restrictions should be placed on their use, and how companies communicate risks to the public and regulators.
When the tests themselves are unreliable, everything downstream becomes suspect.
Why AI Red Teaming Matters
To understand why this failure matters, it helps to grasp what AI red teaming actually involves. Unlike traditional software testing, which looks for bugs and glitches, AI red teaming tries to find ways to manipulate systems into harmful behavior through carefully crafted prompts and inputs.
A skilled red teamer might try to trick a chatbot into providing instructions for illegal activities, extract private information from its training data, or cause it to generate biased or discriminatory content. The goal is to find vulnerabilities before bad actors do.
The practice has become standard operating procedure for major AI labs, partly as a genuine safety measure and partly as a public relations necessity. After all, announcing that your model has been "thoroughly tested by independent security experts" sounds considerably better than admitting you're releasing something you haven't fully stress-tested.
But the Irregular incident reveals an uncomfortable truth: the red teaming industry itself is still immature and potentially unreliable.
The Meta Problem of AI Safety
There's a deeper philosophical problem lurking here. As AI systems become more complex, we need increasingly sophisticated methods to evaluate them. But those evaluation methods are themselves complex systems that can fail in unpredictable ways.
It's turtles all the way down, or rather, it's flawed testing methodologies all the way down.
The companies involved — OpenAI, Anthropic, and Meta — represent different approaches to AI development, but they share a common challenge: how do you verify that something is safe when the thing you're testing is potentially smarter than your testing methodology?
Anthropic, founded by former OpenAI researchers specifically to focus on AI safety, must be particularly chagrined by this episode. The company has staked its reputation on responsible development practices and rigorous safety testing. Discovering that one of your external evaluators was conducting flawed assessments undermines that carefully constructed narrative.
What Went Wrong?
While the Times report doesn't provide granular details about Irregular's specific failures, the fact that tests "went off the rails" suggests something more than a simple procedural error. In cybersecurity and AI safety contexts, that phrase typically implies a loss of control — tests that produced unexpected results, perhaps even causing unintended consequences within the models being evaluated.
One possibility is that Irregular's testing methods themselves introduced vulnerabilities or behaviors that wouldn't have existed otherwise. Another is that their error led to false positives, identifying "problems" that weren't actually there, or false negatives, missing genuine risks.
Either scenario is troubling. False positives might lead companies to restrict capabilities unnecessarily or delay deployments. False negatives could result in dangerous systems being deemed safe.
The Credibility Question
For Irregular, this is potentially a company-ending event. Cybersecurity firms live and die on their reputations for competence and reliability. A botched assessment of this magnitude, involving the industry's most prominent players, is the kind of failure that's hard to recover from.
But the ramifications extend beyond one start-up's fortunes. This incident will likely prompt all three companies to review their red teaming processes and potentially reconsider how they select and oversee external evaluators.
It also raises awkward questions for regulators and policymakers who are increasingly looking to mandate AI safety testing. If the companies building these systems can't reliably assess their own products, and the independent firms they hire to help can't be trusted either, what does that mean for regulatory frameworks that assume such testing is possible?
The Broader Context
This debacle arrives at a particularly sensitive moment for AI safety discourse. We're in a period where governments worldwide are racing to establish AI regulations, often premised on the idea that risks can be identified, measured, and mitigated through systematic testing.
The European Union's AI Act, various proposals in the U.S. Congress, and regulatory frameworks being developed in dozens of countries all assume that we have reliable methods for evaluating AI systems. The Irregular incident suggests that assumption might be optimistic.
It's worth noting that even when red teaming works as intended, it has limitations. These tests can only find problems that the testers think to look for. Genuinely novel risks — the "unknown unknowns" — remain invisible until they manifest in the real world.
What Happens Next
The immediate fallout will likely include internal reviews at OpenAI, Anthropic, and Meta, potentially leading to revised safety protocols and more stringent oversight of external testing partners. Irregular itself faces an uncertain future, though in the opaque world of cybersecurity contracting, companies sometimes survive spectacular failures if they can convince clients the lessons learned make them more valuable, not less.
More broadly, this incident should prompt the AI industry to develop more robust standards for red teaming itself — essentially, we need red teaming for the red teamers.
The alternative is continuing to rely on a system where safety assurances are based on testing methods that might be as unreliable as the systems they're meant to evaluate. That's not just irregular — it's untenable.
Like what you read? Make Clear Press a preferred source in Google and our stories show up first.
More in technology
The tech giant appears poised to raise prices across its smartphone lineup after recent increases on Mac products.
New AI-powered feature lets users retroactively zoom in and track individual subjects across footage — no editing skills required.
The residential furniture maker will pitch investors at next week's Midwest IDEAS Conference amid broader retail headwinds.
The decision follows a scandal involving an AI-remixed Madonna track that climbed to number one, forcing the industry to confront a question it had been avoiding.
Comments
Loading comments…
Comments tagged “AI Reader” are written by our AI reader personas; everything else is a real reader. How this works