Recover without a redeploy
Failing agents are retried or quarantined automatically, and serious incidents go to your on-call.
For teams running AI agents in production
AgentGuardian runs on your infrastructure, outside any one cloud's agent platform. It checks tool calls against your policies before they run and recovers failing agents without a redeploy.
Works with any agent that can make an HTTP call. We reply within 48 hours.
Ways to connect
Python today. Any language through the HTTP gateway.
@arep_agentPython decoratorMCPWrapperMCP tool wrapperPolicy gatewayHTTP, any languageREST APIPolicies, incidents, auditThe problem
Agents loop, time out, and drift from what they said they would do. Teams often learn about it from an alert after the damage is done, with no way to stop the next call.
Observability tells you what went wrong. AgentGuardian decides what is allowed to happen, and acts when something breaks.
Why now
Gartner published its first Market Guide for Guardian Agents in February 2026, and expects the category to reach 10 to 15 percent of the agentic AI market by 2030. AgentGuardian is a self-hosted guardian layer between your agents and the systems they touch, whichever cloud they run on.
150,000
AI agents Gartner expects per Fortune 500 enterprise by 2028.
13%
Organizations that think they have the right AI agent governance in place, per Gartner.
80%+
Unauthorized agent transactions Gartner expects through 2028 to come from internal policy violations, not malicious attacks.
40%+
Agentic AI projects Gartner expects to be canceled by 2027, citing cost, unclear value, and weak risk controls.
Sources: Gartner on agent sprawl, 2026 · Gartner Market Guide for Guardian Agents, 2026 · Gartner on guardian agents, 2025 · Gartner on project cancellations, 2025
How it works
Add one decorator to see every agent run. Put a gateway check in front of the tool calls you want to govern.
agent.py
from arep.sdk import arep_agent
@arep_agent(agent_id="finance-agent")
async def run(task):
# your agent codebefore a governed tool call
POST /gateway/invoke
{ "agent_id": "finance-agent",
"tool": "exec_shell",
"inputs": { ... } }Each governed call is allowed, denied, or held for human sign-off. Your agent gets the decision before the tool runs.
gateway decisions
read_fileALLOWwithin policyexec_shellDENYblocked by policyissue_refundAPPROVALawaiting sign-offWhen a failure is detected, an incident opens and a recovery action starts right away. Severity decides the action.
incident timeline
Platform
One layer outside your agents for policy, incidents, recovery, and audit, shared by every agent you run.
Failing agents are retried or quarantined automatically, and serious incidents go to your on-call.
Agents check each governed tool call with the policy gateway and get allow, deny, or wait for approval before it runs.
Related failures roll up into one prioritized incident, with the chain of events that caused it.
Actions that drift from what the agent said it would do are flagged as soon as they happen.
Every agent and tool on one live map. Health changes and blocked calls show up within seconds.
A tamper-evident record of every decision, verifiable on demand when compliance asks.
Enterprise governance
Write rules per agent and per tool, and check any call against them without running it. Every decision is recorded for audit.
Production guardrails
shell and system access
Data access controls
databases and PII
Deployment approvals
production releases
Payments team
team-defined rules
Security
AgentGuardian runs inside your infrastructure and keeps a verifiable record of every decision it makes.
Run it with Docker Compose on one server today. A Helm chart and a Kubernetes operator are in development. Agent data stays on your network.
Agents, operators, and internal services use separate API keys. Each key carries only the permissions its job needs.
The audit log is hash-chained, so an edit to any past entry shows up when you run an integrity check.
Ready to put a policy check in front of your agents?
FAQ
Python agents connect through the @arep_agent decorator, and MCP tools through MCPWrapper. Any language can call the HTTP policy gateway. Integrations for specific agent frameworks are on the roadmap.
If all your agents run on one platform, its built-in controls may be enough. Amazon Bedrock AgentCore, Microsoft Agent 365, and Google's Agent Gateway each tie policy to their own platform's gateway and identity. AgentGuardian runs on your infrastructure and works with any agent that can make an HTTP call, wherever it runs. It also opens incidents and recovers failing agents when a call goes wrong.
In your infrastructure, on Postgres and Redis. Today it runs on a single server with Docker Compose. A Helm chart and a Kubernetes operator that automates upgrades and scaling are in development. Agent data stays on your network.
The gateway returns a pending decision, the request lands in the approval queue, and your team is notified in Slack. Someone on your team approves or denies it, and the decision is written to the audit log.
Yes. A dry run checks a sample call against your active policies and returns the decision without executing anything. Policies are versioned, and a removed policy can be restored with one call.
When a detector fires, an incident opens with a severity from P0 to P3. The recovery engine then tells the agent to retry, quarantines the agent, or escalates to your on-call. Retry works for agents that use MCPWrapper with directive polling turned on.
Not yet. Because AgentGuardian is self-hosted, you control where agent data lives, and the hash-chained audit log gives your compliance team a record they can verify. Policy packs mapped to NIST AI RMF, the OWASP Top 10 for Agentic Applications, and the EU AI Act, with audit evidence export, are on the roadmap.
Early access
We are onboarding teams that run AI agents in production and need a policy layer they can show to compliance.
Received
You're on the list.
We'll reach out within 48 hours to set up your onboarding.
Prefer to talk first?
Book a demo