Ten ways agents go wrong, and where an authority layer stops.
Each pattern says what happens, what Torvant does, and its limit. We state the limit because a control that overclaims is the one that fails an audit.
What goes wrong, what Torvant does, where it stops.
They follow our own threat model and the main risks in the OWASP Top 10 for Agentic Applications. What is built and what is not is on Security and status.
Hijacked goal
Text in a web page, email or file redirects the agent to a goal nobody approved.
- What Torvant does
- Does not stop the injection. Limits what the hijacked agent can do to what its grant allows, and records each refusal.
- Where it stops
- Torvant does not read prompts. Harm inside the grant still happens.
Tool misuse
An agent uses a legitimate tool in a destructive way, such as deleting a database or sending a refund.
- What Torvant does
- Anything no rule allows is refused. Risky actions can be held for a person to approve.
- Where it stops
- Only calls that go through the checkpoint.
Standing credentials
A broad service account is borrowed by an agent and never handed back.
- What Torvant does
- The agent holds a short-lived signed grant tied to a named person's decision, not a standing credential.
- Where it stops
- The tool must also refuse callers that skip the checkpoint. Nothing is in production use yet.
Escalation through sub-agents
A helper agent ends up with more power than the agent that called it.
- What Torvant does
- A hand-off can only narrow. The most it can have is what its parent had, and the person at the root never changes.
- Where it stops
- The full re-check at every checkpoint is partly built. The others refuse hand-offs for now.
Runaway loops and spend
An agent repeats a paid or destructive action again and again.
- What Torvant does
- Limits on how many calls and how much can be spent sit inside the grant, and the grant runs out.
- Where it stops
- The limit has to be written into the grant.
Data leaving through tool output
Sensitive fields come back from a tool and flow on to the model or another system.
- What Torvant does
- Calls that would touch data outside the grant are refused before they run, or the sensitive fields are removed.
- Where it stops
- Streaming responses are not yet covered.
Stolen grant replay
A leaked grant is used by someone other than the agent it was made for.
- What Torvant does
- Grants are short. In self-hosted setups each call must also be proved by the agent's own key.
- Where it stops
- Pinning a grant to the platform the agent runs on is the next change.
Edited or deleted records
Someone rewrites what the agent did, or removes what it was refused.
- What Torvant does
- Each entry is linked to the one before it, back to the person's decision. A broken chain is reported loudly.
- Where it stops
- The cryptography has not had its external review yet. It is funded and pending.
Routing around the checkpoint
A misconfigured or rogue agent calls a tool directly.
- What Torvant does
- The tool itself can refuse callers that did not come through. In the self-hosted setup the agent's only way out is through Torvant.
- Where it stops
- This is the weakest point today. A tool that does not check is open to an agent that reaches it. Containment holds on Linux only.
A person who has left
Their agents keep acting after they are gone.
- What Torvant does
- Grants run out quickly and can be withdrawn, and everything handed down under a withdrawn grant stops.
- Where it stops
- Switching someone off in your sign-in provider does not yet withdraw their grants by itself.
What Torvant does not try to stop.
Tricks played on the model itself, poisoned agent memory, and bugs in the agent framework or in the tools. An authority layer limits what an agent can do. It does not make the agent right. What an authority layer is.
See it on your own agent.
A 30-minute call. We watch one of your agents, block nothing, and show you what your rules would have stopped.