Attack patterns

Ten ways agents go wrong, and where an authority layer stops.

Each pattern says what happens, what Torvant does, and its limit. We state the limit because a control that overclaims is the one that fails an audit.

The patterns

What goes wrong, what Torvant does, where it stops.

They follow our own threat model and the main risks in the OWASP Top 10 for Agentic Applications. What is built and what is not is on Security and status.

01

Hijacked goal

Text in a web page, email or file redirects the agent to a goal nobody approved.

What Torvant does
Does not stop the injection. Limits what the hijacked agent can do to what its grant allows, and records each refusal.
Where it stops
Torvant does not read prompts. Harm inside the grant still happens.
02

Tool misuse

An agent uses a legitimate tool in a destructive way, such as deleting a database or sending a refund.

What Torvant does
Anything no rule allows is refused. Risky actions can be held for a person to approve.
Where it stops
Only calls that go through the checkpoint.
03

Standing credentials

A broad service account is borrowed by an agent and never handed back.

What Torvant does
The agent holds a short-lived signed grant tied to a named person's decision, not a standing credential.
Where it stops
The tool must also refuse callers that skip the checkpoint. Nothing is in production use yet.
04

Escalation through sub-agents

A helper agent ends up with more power than the agent that called it.

What Torvant does
A hand-off can only narrow. The most it can have is what its parent had, and the person at the root never changes.
Where it stops
The full re-check at every checkpoint is partly built. The others refuse hand-offs for now.
05

Runaway loops and spend

An agent repeats a paid or destructive action again and again.

What Torvant does
Limits on how many calls and how much can be spent sit inside the grant, and the grant runs out.
Where it stops
The limit has to be written into the grant.
06

Data leaving through tool output

Sensitive fields come back from a tool and flow on to the model or another system.

What Torvant does
Calls that would touch data outside the grant are refused before they run, or the sensitive fields are removed.
Where it stops
Streaming responses are not yet covered.
07

Stolen grant replay

A leaked grant is used by someone other than the agent it was made for.

What Torvant does
Grants are short. In self-hosted setups each call must also be proved by the agent's own key.
Where it stops
Pinning a grant to the platform the agent runs on is the next change.
08

Edited or deleted records

Someone rewrites what the agent did, or removes what it was refused.

What Torvant does
Each entry is linked to the one before it, back to the person's decision. A broken chain is reported loudly.
Where it stops
The cryptography has not had its external review yet. It is funded and pending.
09

Routing around the checkpoint

A misconfigured or rogue agent calls a tool directly.

What Torvant does
The tool itself can refuse callers that did not come through. In the self-hosted setup the agent's only way out is through Torvant.
Where it stops
This is the weakest point today. A tool that does not check is open to an agent that reaches it. Containment holds on Linux only.
10

A person who has left

Their agents keep acting after they are gone.

What Torvant does
Grants run out quickly and can be withdrawn, and everything handed down under a withdrawn grant stops.
Where it stops
Switching someone off in your sign-in provider does not yet withdraw their grants by itself.
Out of scope

What Torvant does not try to stop.

Tricks played on the model itself, poisoned agent memory, and bugs in the agent framework or in the tools. An authority layer limits what an agent can do. It does not make the agent right. What an authority layer is.

See it on your own agent.

A 30-minute call. We watch one of your agents, block nothing, and show you what your rules would have stopped.

Book a demo →