- The hammer paradox: essential tool, inevitable weapon
- The mathematical proof that changes everything: Gödel's theorem applied to guardrails
- The Mythos/Fable affair: 19 days that shook the industry
- The flight to Chinese models
- Comparison table: vetted access programs
- What this means for small and medium businesses in the Basque Country
- Ready to secure your AI tools? Contact Mister Anderson.
- FAQ: your questions about AI guardrails and cybersecurity
AI Guardrails: The Cybersecurity Researcher Dilemma
Imagine a hammer. You need one to build a house. But in the wrong hands, it's a weapon. Should hammers be banned? That's exactly the dilemma the AI industry faces today.
Since 2025, large language models (LLMs) have become central tools for offensive cybersecurity researchers. Code analysis, vulnerability research, test script generation — AI speeds up their work. But the guardrails imposed by OpenAI, Anthropic and Google systematically block these legitimate uses.
And the situation is getting worse.
In June 2026, the Mythos/Fable affair showed that a government's mere whim can strip researchers of the tools they rely on, overnight. The result? The world's top cybersecurity talent is turning to open alternatives… from China.
Here's why that's a problem, and what it means for you, from Bayonne to Biarritz.
The hammer paradox: essential tool, inevitable weapon
Chris Anley, technical director at NCC Group, puts the problem in simple terms:
"It's like a hammer. You can't build a house without a hammer. But it's also a weapon."
The heart of the problem? AI guardrails can't tell the difference between defensive and offensive intent. A request like "fix this code" could just as easily mean patching a critical flaw as writing an exploit. To an AI model, the two look alike.
Witnesses telling the same story
Mark Dowd, a well-known zero-day broker, admits his discomfort: "It makes me uneasy that companies arbitrarily decide what's safe in security." Hard to argue with him. How can a Californian company judge the legitimacy of research being carried out in Bayonne or Anglet?
Paolo Stagno, founder of CrowdFense, is more blunt: "AI companies treat clients like children who need a babysitter." Behind the quip lies a real problem: trained, certified adult professionals running into automatic filters that treat them like amateur hackers.
Giuseppe Cali, on the other hand, isn't affected — but for a telling reason: he doesn't use AI for bug hunting. In other words, those who work with AI get blocked, while those who do without it carry on. Not exactly logical.
An anonymous researcher who works for a major cybersecurity firm says: "I spend as much time jailbreaking the model as doing my actual job. Sometimes more." Red-teaming AI has become a job in its own right — one where researchers must first get around the guardrails before they can even look for vulnerabilities.
The mathematical proof that changes everything: Gödel's theorem applied to guardrails
Do you believe that one day we'll find THE miracle solution that makes AI infallibly safe? NIST has bad news for you.
In June 2026, Apostol Vassilev, a researcher at the National Institute of Standards and Technology, published in IEEE Security & Privacy a demonstration that landed like a bombshell. He applies Gödel's incompleteness theorems to AI guardrail systems.
Translation for the non-mathematicians: any finite system of rules can be bypassed. Always. By mathematical nature.
Vassilev puts it bluntly: "There will always be ways to prompt the AI into ignoring these rules."
Stanford confirms it
The numbers are staggering. Stanford University tested fine-tuning attacks on several models:
| Model | Bypass rate |
|---|---|
| Claude Haiku | 72% |
| GPT-4o | 57% |
And that's just the tip of the iceberg. OWASP has ranked prompt injection as the number one risk for LLMs since 2025. The dark web, meanwhile, sells jailbreak frameworks and DarkLLMs on subscription — AI models built specifically for crime.
NIST's proposed solution
Faced with this, NIST isn't recommending more static guardrails. Quite the opposite. The official recommendation is a continuous monitoring model: ongoing surveillance, red teams, resilience. Not fixed barriers, but adaptive detection.
In plain terms: stop building walls anyone can climb over. Learn to watch what's happening behind them instead.
The Mythos/Fable affair: 19 days that shook the industry
From June 12 to June 30, 2026. 19 days that changed everything.
The Trump administration, through Howard Lutnick, imposed export controls on Anthropic's Mythos and Fable models. These models, which Anthropic itself markets as "doomsday cybermachines," were suddenly banned for non-American researchers.
Overnight, entire cybersecurity teams found themselves without a working tool. No transition. No alternative offered. Nothing.
The way out of the crisis
On June 30, the controls were lifted. But at what cost? Anthropic had to agree to "proactively detect and address risks." Concretely:
- Fable 5 returned on July 1, with no major restrictions.
- Mythos 5 remains limited to "vetted" US organizations.
- Anthropic launched the Glasswing program, a special access track for defensive security.
The fallout
Trust is broken. Researchers now understand: their access to the best AI tools can be cut off at any moment, for political reasons.
The Check Point AI Security Report 2026 confirms it: AI has shifted from "development tool" to "live attack operator." Red-teaming became a continuous engineering discipline in 2026. And the EU AI Act now makes red-teaming mandatory for high-risk models.
But how do you red-team something you can't even access?
The flight to Chinese models
This is the most ironic consequence of the whole story. Guardrails, meant to protect, are pushing the best researchers toward the very systems they're supposed to avoid.
Chris Thompson, organizer of Offensive AI Con, sums it up in one line: "Guardrails behave differently every day. One day you can do something, the next you can't. Responsible researchers are being pushed toward foreign systems."
Which foreign systems? Mostly Chinese models:
- GLM (Zhipu AI)
- DeepSeek
- Qwen (Alibaba)
These open-source models offer what the American giants refuse: flexibility, no arbitrary guardrails, and the ability to get work done without running into absurd filters.
The transfer of expertise is real. Researchers who once trained American AI on cybersecurity are now training the Chinese alternatives. Not out of ideology, but out of necessity.
A Paris-based researcher who regularly works with teams in Bayonne says: "I'd rather work with the US models. They're technically better. But I can't get my job done if I have to spend three hours unblocking a filter just to analyze a PDF."
Comparison table: vetted access programs
Here are the programs researchers can apply to for exceptional access. Spoiler: it's slow, it's restrictive, and it's not guaranteed.
| Criteria | OpenAI Trusted Access for Cyber | Anthropic CVP | Anthropic Glasswing |
|---|---|---|---|
| Target audience | Cybersecurity researchers | Verified professionals | Defensive security |
| Documentation required | Identity, employer, references | Identity, certifications | Identity, client contract |
| Approval time | 2 to 4 weeks | 1 to 3 weeks | Case-by-case |
| Usage quotas | Yes, limited | Yes, monitored | Variable |
| Accessible models | GPT-4o (restricted version) | Claude (controlled version) | Limited Mythos / Fable |
| Geographic restrictions | Non-US subject to review | US prioritized | US organizations only |
| Output auditing | Yes, systematic | Yes, random | Yes, enhanced |
The problem is obvious: when you're chasing a zero-day vulnerability, you don't have two weeks to wait. A flaw needs urgent attention. Vetted access programs are designed for academic research, not for operational cybersecurity.
What this means for small and medium businesses in the Basque Country
We've been talking about Californian researchers, American regulations, Chinese models. But day to day, concretely, what does this change for a small business in Bayonne, an accounting firm in Anglet, or a tech startup in Biarritz?
The data risk
If researchers can't properly test the AI models you use, those models carry undetected vulnerabilities. Your client data, your accounting files, your business emails — anything that passes through an unaudited LLM is potentially exposed.
Vendor dependency
Small and medium businesses in the Basque Country that rely on AI for their operations find themselves caught between two fires:
- US models (OpenAI, Anthropic): expensive, restricted, and subject to unpredictable political decisions.
- Chinese models (DeepSeek, GLM): more open, but with data sovereignty questions that go beyond most companies' scope.
The Mister Anderson recommendation
At Mister Anderson, we help small and medium businesses in the Basque Country navigate this mess. Not by selling a dream, but with concrete solutions:
- Audit of your AI tools: we check what your models can actually do, what they leak, and what they expose.
- Choosing the right model: depending on your business (accounting, sales, consulting), we select the AI best suited to you — US, Chinese, European, or open-source.
- Data security: we set up guardrails that work for you, not the ones designed by Silicon Valley giants.
Are you in Bayonne, Anglet, Biarritz or elsewhere in the Pyrénées-Atlantiques? We work with you, in your language, with no unnecessary technical jargon.
Ready to secure your AI tools? Contact Mister Anderson.
Using AI in your business and want to make sure your data is protected? We'll help you sort out the good tools, the real risks, and the marketing promises.
No hidden quote, no sales pitch. We listen to your situation, analyze it, and advise you. Whether you're in Bayonne, Anglet, Biarritz, Saint-Jean-de-Luz or elsewhere.
FAQ: your questions about AI guardrails and cybersecurity
| Question | Answer |
|---|---|
| Why do AI guardrails cause problems for cybersecurity researchers? | Guardrails automatically block security-related requests without distinguishing between defensive and offensive intent. The same phrase, "fix this code," can be legitimate or malicious. Researchers lose considerable time working around these filters. |
| What does Gödel's theorem have to do with AI? | NIST mathematically demonstrated that no finite system of rules can be both complete and consistent. In other words: every static guardrail can be bypassed, by mathematical construction. |
| Who are the researchers affected by these blocks? | Recognized professionals: Chris Anley (NCC Group), Mark Dowd (zero-day broker), Paolo Stagno (CrowdFense). Offensive cybersecurity teams at large companies and independent researchers are also affected. |
| Are Chinese models really better for cybersecurity? | They're not technically better, but they're more accessible. DeepSeek, GLM and Qwen offer fewer arbitrary restrictions, letting researchers work without running into time-consuming filters. |
| What is Anthropic's Glasswing program? | A vetted access program launched after the Mythos/Fable affair, aimed at defensive security organizations. It grants access to advanced models under strict conditions, reserved for US organizations. |
| How can a small business in Bayonne secure its AI tools? | By starting with an audit: which models are you using? What data are you feeding them? What guardrails, if any, protect that data? Mister Anderson helps Basque Country businesses work through these questions. |
Article written by Mister Anderson — Auditing and securing AI tools for small and medium businesses in the Basque Country (Bayonne, Anglet, Biarritz, 64).
