Runtime security for AI agents

Stop unsafe agent actions
before they run.

InferenceFort evaluates what AI agents are about to do—across model and tool calls—then blocks, redacts, or holds unsafe actions without rerouting your AI traffic.

Python + TypeScript Cloud + on-prem No proxy
prompt versus action protected
Prompt-only check “Review this vendor invoice and handle it.” The request looks routine. Pass

Prompts describe intent. Actions create impact.
InferenceFort evaluates the action.

Agent action Send payment to a newly introduced account The action does not match the user’s trusted context. Blocked
We judge what the agent is about to do—not just what it was told.
Measured on AgentDojo Protection that keeps agents useful.

Prompt-injection defense measured on real tool-using agent tasks.

View the benchmark methodology ↗
AgentDojo benchmark result <1% ASR without any utility loss

Measured by InferenceFort. Results vary by model and detector configuration.

Works across the agent stack you already run
LangChain LangGraph CrewAI Anthropic OpenAI Ollama vLLM Bedrock Vertex Groq Splunk Sentinel
Why runtime security

The risky part is not the prompt.
It is what the agent does next.

Agents read untrusted content, hold credentials, and act across real systems. InferenceFort evaluates that full sequence while there is still time to stop it.

01

See the complete session

Connect user intent, retrieved content, model calls, and tool actions instead of judging each string in isolation.

02

Decide before execution

Allow safe work, remove hostile instructions, require approval, or block a dangerous action before the tool runs.

03

Keep evidence your SOC can use

Record the user, agent, model, tool, destination, verdict, and reason in one trace and forward it to your SIEM.

The product

One control point for every agent action.

Discover the agents already running, enforce runtime decisions, and investigate the complete trail across users, models, tools, destinations, and costs.

app.inferencefort.com/agents Live
AI agent inventory All teams · 6 Cloud + on-prem Needs attention · 2
Set caps Manage policy Export → SIEM
$258 / $430 cap
Spend today
1,204
Calls governed today
37
Blocked / capped
7
Active agents
Agent
Owner
Model
Daily spend vs cap
Status
SB support-bot
Support
Priya Shah
Customer Success
claude-sonnet-4
$32/ $50
Active
SC sales-copilot
Sales
Sam Rivera
Revenue
gpt-4o
$47/ $50
Near cap
BA billing-agent
Finance
Alex Morgan
Finance Ops
gpt-4o
$50/ $50
Capped · blocked
CB claims-bot
Claims
Nina Kapoor
Insurance
claude-sonnet-4
$18/ $40
health data redacted
RB research-bot
R&D
Mei Lin
Research
claude-opus-4
$61/ $120
Active
ET etl-runner
Platform
Data Team
Platform Eng
llama-3-70b · on-prem
$0.00/ unmetered
Local · governed
IA intern-agent
Engineering
Jordan Park
Eng (intern)
gpt-4o
not approved
Blocked
7 agents across 6 teams · cloud + on-prem · each governed call attributed to a real human and checked against policy before it runs
DiscoverSee every agent, model, tool, and MCP server.
ProtectStop prompt injection, data leakage, and unsafe actions.
GovernEnforce access, budgets, destinations, and approvals.
InvestigateSend decision-grade traces to Splunk, Sentinel, or Elastic.
The engine

Detect what prompt filters miss.

InferenceFort evaluates intent, content, destination, and action together. That context reveals attacks that look harmless when each prompt or tool call is inspected alone.

What we watchWhat that catchesResponse
Content the agent reads Hidden instructions planted in documents, pages, tickets, and tool results. Strip
Where data is heading Private or regulated data moving to an unapproved endpoint or recipient. Block
Action vs. the request Deletes, payments, messages, or privilege changes the user did not request. Hold
Blast radius A routine operation that unexpectedly expands from one record to an entire system. Block
Who did it, and what it cost Every model and tool call attributed to a person, team, agent, and cost. Record
01

Layered detection

Combine the in-process engine with Lakera Guard, Presidio, or detectors you already license. Findings become one explainable decision, without making one service the only line of defense.

multiple detectors → one decision
02

Your model, your hardware

Use a model you choose, including one on your own infrastructure. Fully disconnected environments can run enforcement and detection without an external control-plane call.

no inspection traffic through us
03

Preserve legitimate work

Remove hostile instructions while allowing the user’s real task to continue. When a check cannot complete, the result is recorded as incomplete—never silently treated as safe.

attack removed → the task still completes
Deployment

Move from observe to enforce
with an engineer beside you.

A forward-deployed engineer maps your agents and data boundaries, tunes detection against real traffic, and enables enforcement one workflow at a time.

01

Map the real environment

Inventory agents, models, tools, owners, sensitive data paths, and approved destinations.

discover · attribute · scope
02

Prove decisions safely

Start in observe mode, review decisions against real work, and remove false positives before blocking.

watch → agree → enforce
03

Enforce and keep tuning

Turn controls on workflow by workflow, then adapt them as your agents, tools, and threats change.

enforce · measure · improve
The architecture

Security at the point of action.

InferenceFort evaluates model and tool calls inside the application runtime. That gives the decision engine identity and session context without requiring traffic rerouting, and it covers local models that never cross a network boundary.

Agent action A model or tool call fires inside your process.
Inspect context Identity, prompt, and tool + arguments.
Evaluate policy Local rules, no network round-trip.
Decision
Allow Redact Block
Log evidence Structured audit record for your SIEM.
Network gateway / proxy
Local models may bypass the gateway
Tool execution needs separate instrumentation
Adds another request hop
Requires routing or endpoint changes
Application identity context can be limited
Traffic passes through the gateway
InferenceFort
Cloud, Ollama, on-prem, and Bedrock coverage
Model and tool calls governed together
Policy runs in-process from a local cache
SDK install with no traffic rerouting
User × team × agent × cost on every record
Traffic goes direct; audit is metadata only
Also included

Runtime protection with
the controls security teams expect.

The same decision layer also enforces model access, spend limits, approved destinations, and audit requirements—so protection and governance share one policy and one evidence trail.

Model accessWhich teams and agents may use which models. Anything not approved is refused by default.
Spend capsPer-team and per-agent budgets, enforced before the call runs rather than discovered on the invoice.
Data residencyPin each provider to approved regions and destinations. The GDPR question, answered.
Audit trailEvery call, decision, and reason streamed to Splunk, Sentinel, or Elastic in your own format.
Sovereign & on-prem AI

Security that enables deployment,
instead of restricting it.

The usual answer to a regulated workload is to forbid it. Local models (Ollama, vLLM, a fine-tuned Llama on your own GPUs) make no external call, so gateways, CASBs, and DLP see nothing and the security review ends in a no. We run inside the process, so private models are protected exactly like cloud ones and the answer can be yes.

Covers local inference

A local model call never touches the wire, so perimeter tools can't inspect what never leaves. We sit at the call, not the network.

on-prem = governed

Air-gap ready. Assessor ready.

Policy evaluates in-process from a cached bundle, so enforcement works fully disconnected. Structured attribution and decision logs help produce evidence for NIST SP 800-171 and CMMC assessments.

assessment evidence · offline enforcement

One policy, cloud to on-prem

Use the same rules and audit schema for cloud models and Llama on your GPUs. Decision records support EU AI Act Art. 12 evidence needs; endpoint pinning helps enforce regional data boundaries.

EU AI Act logging · regional endpoint controls
Seamless Integration

Deploys as an SDK.
No per-call code.

Add the Python or TypeScript SDK and set one environment variable. Supported model and tool calls are governed at runtime without changing each call site.

  • No per-call code: one package, one environment variable
  • Covers cloud models and on-prem / Ollama
  • Python and TypeScript: LangChain, LangGraph, Vercel AI SDK, CrewAI, and MCP
  • Multi-tenant isolation, built for enterprise deployments
Get early access
terminal
# Python
$ pip install inferencefort

# TypeScript
$ npm install @inferencefort/ai-sdk

# Point either SDK at your control plane
$ export INFERENCEFORT_KEY=your-team-key

# Run your app: supported calls are now governed
$ python app.py

✓ governing: ALLOW priya@ · BLOCK ssn-rule · CAP $50/day
0
per-call code changes for developers
Local
policy decisions, no network round-trip
SIEM
-ready audit evidence
Py + TS
SDK coverage across agent stacks
Your team already signs in through Okta Microsoft Entra ID Connect once, and governed calls map to a real person.
Still have questions?

Questions security teams ask.

Direct answers on data handling, deployment, policy, and local-model coverage.

Does our model traffic pass through InferenceFort?
No. Model calls go directly from your application to the provider or local model; InferenceFort is not a proxy. You control audit delivery and detector configuration. Offline deployments can keep the audit trail and detection entirely inside your environment.
How do you find agents nobody told us about?
Once the SDK is installed, governance activates automatically for calls made through supported frameworks, so the moment an agent, registered or not, makes a model call, it appears in your inventory: named, attributed to an owner, with its model and spend. Shadow AI stops being invisible.
Can each team have its own models, caps, and policies?
Yes. Policies are identity-aware, so you can say "Sales can use these models up to $50/day, Finance these others" and have it enforced per user, team, and agent. Set it once in the panel; every agent picks it up automatically, with no per-team rollout.
Does this cover on-prem and local models too?
On-prem is our strongest case. Local models (Ollama, vLLM, a Llama on your own GPUs) make no external call, so network gateways and DLP are blind to them. InferenceFort runs in-process, so they're governed identically to cloud. Policy evaluates from a local cache with no round-trip, so enforcement works even fully air-gapped.
What does our team have to change to roll this out?
Install the SDK, configure identity and policy, and validate in observe mode. There is no traffic rerouting and supported integrations do not require rewriting every model or tool call.

Bring us the agent
your security team worries about.

We are onboarding a small group of design partners. We will map the workflow, run InferenceFort in observe mode, and show you exactly what it would stop before enforcement begins.