Oreset

Human red team for AI agents

Your AI agent will be attacked.It will also just decide wrong.

We red-team both halves: the security failures a hacker would exploit, and the judgment failures that need no hacker at all. Every finding verified, every fix retested free.

See how it works
1–2 weeks
per engagement
Two angles
security and judgment
Free retest
on every fix
Zero SDKs
no integration
The Oreset client dashboard: engagement progress, resilience score, open findings and scenarios run

Why this exists

Most AI failures that made headlines were not attacks. They were an agent with a real tool, a plausible request, and no sense of when to stop. Security firms test one half. We test both.

What we test

01 Security

Find what a hacker would.

Prompt injection, data leaks, unauthorised tool calls and guardrail bypass, tried by people who do this for a living.

  • Injection through documents and tool results
  • System prompts and other users’ data
  • Actions the agent was never allowed to take
The red team workspace: a live attack conversation with a client’s agent beside the attack playbook

02 Judgment

Catch the decisions that need no hacker.

We run the same request under the conditions real users create: pressure, ambiguity, and text that looks like an instruction.

  • Decisions that shift under urgency
  • Ambiguous and multilingual input
  • Policies and refunds the agent invents
A tester assessing an exploit trace: the attack prompt, the model response, the unauthorised tool call and the vulnerability form

03 Impact

Know what it costs, and the fix.

Every finding arrives with a severity, the business consequence, exact reproduction steps and a specific fix. Patch it and we retest free.

  • Severity and fix priority
  • Legal and reputational exposure
  • Closed only when the attack stops landing
A verified P0 finding in the client dashboard with its business impact, reproduction steps and recommended fix

Real screens from the platform, with demo data.

How it works

You give access.We do the rest.

An engagement typically runs one to two weeks. There is no SDK to install and nothing to integrate.

  1. 1

    Scope

    We plan the tests around your agent, not a template.

    You
    Tell us what the agent does, who it serves and what it can touch.
    We
    Agree scope, rules of engagement and a test plan.
  2. 2

    Test

    People run hundreds of scenarios against the live agent.

    You
    Give us access. Nothing to install and no SDK.
    We
    Attack it, edge-case it, and try it in the languages your users speak.
  3. 3

    Findings stream in

    Each one lands once a lead auditor has reproduced it.

    You
    Watch the dashboard as findings arrive.
    We
    Verify severity and write up impact, steps and the fix.
  4. 4

    Fix and retest, free

    A finding closes only when the attack stops landing.

    You
    Patch it and mark it fixed.
    We
    Re-run the same attack and confirm it holds.
  • Two things: can your AI agent be manipulated into doing something wrong (security), and does it make wrong decisions on its own without being attacked (judgment). Most security firms only cover the first half. We cover both, in one engagement.

Get in touch

Someone will find out what your agent does under pressure.It should be you.

We're onboarding early partners shipping AI agents in fintech, payments, and lending. Tell us what your agent does and what it can touch.

  • A person replies, usually within one working day.
  • Engagements start from $3,500, scoped first, retest included.
  • Everything you share stays under NDA and our data-handling policy.

Prefer email?info@oreset.africa

Start here

What brings you here?

Pick the closest one.

What is this about?