We put two of the most capable models available today in the attacker's seat, reproduced the attacks enterprise systems face, and measured what broke and what held. The two smartest models became the two most effective attackers.
We cast the AI as the attacker, gave it nothing but a low-privilege account, and threw five attack scenarios at the system. The goal is to reproduce the attacks enterprise systems are really exposed to.
From low privilege, go after confidential customer data.
Erase or rewrite the traces of what was done.
After intrusion, go after decrypting encrypted data.
Slip an illicit transfer past detection.
Break the approval flow by impersonation.
The stronger the model, the more it broke.
So — do existing defenses work against an AI? —
The paths this pair broke were inside the classic defenses most companies already run. AI agents slip past them on their own.
Classic defenses are breakable by AI.
And for social engineering, the very concept of a defense doesn't exist.
In a demo, we show Lemma stopping attacks before they execute. We'll hear your situation and can discuss adopting Lemma — or an attack-resistance test of your own system.
Lemma is a new way to face AI attacks — agent-facing security. Before execution, it demands proof of who, with what authority, and on what data — and stops any operation that cannot prove it. Rather than detecting attacks and chasing them, it stops unprovable operations before they execute. That is agent-facing security.
Approval and payment had no defense mechanism at all. Lemma demands a mathematical authorization proof and stops anything out of scope before it executes. Only Lemma stops it.
The difference wasn't the model; it was the presence of a proof layer. Before a high-risk operation it demands proof of who, with what authority, on which data — and if there's none, it stops the action before it's ever sent (fail-closed). That is Lemma's role.
Every breach happened because the AI held keys or credentials and escalated them. Lemma adds one layer on the server that changes that premise — before a high-risk operation it requires, as proof, who, with what authority, on which data, and stops out-of-scope operations before they execute (fail-closed). It drops into your existing servers and APIs without a major rewrite.
Layer a proof gate over the attacks, and the outcome changes like this:
Start with a 30-minute demo. We'll show Lemma stopping attacks before they execute, and discuss anything from adopting Lemma to an attack-resistance test of your own system. No disclosure of sensitive data required.
* Attack-resistance testing is quoted separately depending on scope. Start with a demo and a conversation.
We review your target systems and requirements. No disclosure of sensitive data required.
We drop Lemma's proof gate into a staging environment in a minimal configuration.
Measure the no-proof vs. proof difference under attack scenarios. See the effect in numbers.
Based on the results, we finalize the integration scope and the path to production.
The attack-test code is public; third parties can reproduce it in the same environment. The premises and how to read this are folded below.
403