Human red team for AI agents
Your AI agent will be attacked.It will also just decide wrong.
We red-team both halves: the security failures a hacker would exploit, and the judgment failures that need no hacker at all. Every finding verified, every fix retested free.
- 1–2 weeks
- per engagement
- Two angles
- security and judgment
- Free retest
- on every fix
- Zero SDKs
- no integration

Why this exists
Most AI failures that made headlines were not attacks. They were an agent with a real tool, a plausible request, and no sense of when to stop. Security firms test one half. We test both.
What we test
Find what a hacker would.
Prompt injection, data leaks, unauthorised tool calls and guardrail bypass, tried by people who do this for a living.
- Injection through documents and tool results
- System prompts and other users’ data
- Actions the agent was never allowed to take
Catch the decisions that need no hacker.
We run the same request under the conditions real users create: pressure, ambiguity, and text that looks like an instruction.
- Decisions that shift under urgency
- Ambiguous and multilingual input
- Policies and refunds the agent invents
Know what it costs, and the fix.
Every finding arrives with a severity, the business consequence, exact reproduction steps and a specific fix. Patch it and we retest free.
- Severity and fix priority
- Legal and reputational exposure
- Closed only when the attack stops landing



Real screens from the platform, with demo data.
What we test
01 Security
Find what a hacker would.
Prompt injection, data leaks, unauthorised tool calls and guardrail bypass, tried by people who do this for a living.
- Injection through documents and tool results
- System prompts and other users’ data
- Actions the agent was never allowed to take

02 Judgment
Catch the decisions that need no hacker.
We run the same request under the conditions real users create: pressure, ambiguity, and text that looks like an instruction.
- Decisions that shift under urgency
- Ambiguous and multilingual input
- Policies and refunds the agent invents

03 Impact
Know what it costs, and the fix.
Every finding arrives with a severity, the business consequence, exact reproduction steps and a specific fix. Patch it and we retest free.
- Severity and fix priority
- Legal and reputational exposure
- Closed only when the attack stops landing

Real screens from the platform, with demo data.
The platform
The dashboard is the deliverable.
No PDF at the end. Findings land as they are verified, with everything your engineers need to fix them.
Live findings
Every finding, the moment a lead auditor confirms it.

Verified twice
Reproduced before you ever see it.

Free retest
A fix counts when the attack stops landing.
- Discovered
- Verified
- Fix submitted
- Retested
- Closed
Regression suite
Our attacks, replayed in your CI.
- Build and unit tests
- Pull Oreset regression suite4 cases
- Replay against staging agent1 of 4 landed again
- Deployskipped
A fixed finding that starts landing again fails the build before your users see it.
Illustrative. Pulled with a read-only API token.
Languages
Tested in the languages your users actually speak.
Low-resource languages move decisions. We test for that.
How it works
You give access.We do the rest.
An engagement typically runs one to two weeks. There is no SDK to install and nothing to integrate.
- 1
Scope
We plan the tests around your agent, not a template.
- You
- Tell us what the agent does, who it serves and what it can touch.
- We
- Agree scope, rules of engagement and a test plan.
- 2
Test
People run hundreds of scenarios against the live agent.
- You
- Give us access. Nothing to install and no SDK.
- We
- Attack it, edge-case it, and try it in the languages your users speak.
- 3
Findings stream in
Each one lands once a lead auditor has reproduced it.
- You
- Watch the dashboard as findings arrive.
- We
- Verify severity and write up impact, steps and the fix.
- 4
Fix and retest, free
A finding closes only when the attack stops landing.
- You
- Patch it and mark it fixed.
- We
- Re-run the same attack and confirm it holds.
Already happened
Nobody tested whether it would decide right.
Security firms test if your AI can be hacked. Monitoring watches it after launch. These were public failures that needed no hacker at all.
See what testing would have caughtThe Oreset Red Team
Every finding is checked twice before you see it.
A red team is only as good as its false positive rate. Ours is built to be measured: vetted testers who pass calibration before touching a live engagement, independent double review on high-stakes scenarios, and a lead auditor who has to reproduce a finding before it counts. Testers work under NDA, a code of conduct, and a data handling policy, all signed before they see a scenario.
Calibrated testers
Every tester passes scored practice scenarios with known correct outcomes before running a live engagement. Accuracy is tracked continuously.
Dual-review consensus
High-stakes scenarios are assessed independently by two testers. Agreement is measured statistically, not assumed. Disagreements go to adjudication, not a coin flip.
Lead auditor verification
Every critical or high finding is reproduced by a lead auditor before you see it. They confirm it, correct the severity, or throw it out. False positives never reach your dashboard.
Structured taxonomy
Every finding carries an OWASP-aligned vulnerability class, a P0 to P3 severity, reproduction steps, and a business impact. Structured data you can act on and export, not a summary.
Before launch, not after
Find out first.Not from a headline.
Someone will find out what your agent does under pressure. It should be you.
Tell us about the agent
A short form. What it does and what it can touch.
We scope it with you
A person replies, usually within one working day, with a plan and a price.
Findings start landing
One to two weeks. Verified, in your dashboard, retest included.
FAQ
Straight answers.
For teams considering an engagement.
Two things: can your AI agent be manipulated into doing something wrong (security), and does it make wrong decisions on its own without being attacked (judgment). Most security firms only cover the first half. We cover both, in one engagement.
Get in touch
Someone will find out what your agent does under pressure.It should be you.
We're onboarding early partners shipping AI agents in fintech, payments, and lending. Tell us what your agent does and what it can touch.
- A person replies, usually within one working day.
- Engagements start from $3,500, scoped first, retest included.
- Everything you share stays under NDA and our data-handling policy.
Prefer email?info@oreset.africa