AI Agent Security

Secure what your agents do, not just what they read.

Agents read private data, take in untrusted content and call tools that change real systems, often in the same conversation. InferenceFort checks each model and tool call inside your application, with the whole session in view, and blocks harmful actions before they run.

<1%Attack success rate
93.2%Utility under attack
65%Fewer tool-output scans

AgentDojo, InjecAgent and ASB. Attack success and utility are relative to the ungoverned baseline; scan reduction is compared with screening every tool result. Methodology and limits →

The attack surface moved

A chatbot can say something wrong. An agent can do something wrong: send an email, move money, delete a record, post data to a URL. Securing an agent means deciding, for each action, whether it should happen given everything that led up to it.

Any one of these is normal. An agent holding all three is one injected instruction away from a data leak.

How InferenceFort secures agents

Checks every call before it runs

Model and tool calls are evaluated in your process before the prompt leaves or the tool executes. Policy is cached locally, so allowed calls do not wait on a network round trip.

Follows risk across the session

The lethal-trifecta guard blocks a tool that can send data out once the session has touched both private data and untrusted content.

Screens for prompt attacks

User prompts and tool results are screened with the built-in detector or your own, and findings feed the action decision.

Controls tools and MCP servers

Block tools by name pattern, require approval for sensitive ones, and allow only approved MCP servers.

Pins where data can go

Restrict each model provider to approved endpoints or regions. A call to anywhere else is blocked before it is sent.

Keeps tenants apart

Every decision carries user, agent and customer identity. Session state is kept per customer, so one tenant's activity never changes another's verdicts.

Multi-agent systems

Agents hand work to each other over queues, HTTP and MCP. InferenceFort carries trace context across those hops, including between Python and TypeScript services, so one request becomes one trace. With shared session taint enabled, risk picked up by one agent is visible to the next, so an attack split across agents is still caught.

CrewAI and LangGraph systems register their agents and wiring when they start, so you can see every agent in a system before it handles traffic. Agents that appear later are added from observed activity and marked as runtime additions.

Fits the stack you already use

LanguageGoverned surfaces
PythonLangChain and LangGraph models, LiteLLM, CrewAI tools, OpenAI and Anthropic SDKs, MCP clients and servers
TypeScript / NodeVercel AI SDK, LangChain.js, LangGraph.js, OpenAI and Anthropic SDKs, MCP clients and servers

There is no proxy to deploy and no per-call callback to add. Set identity once at your request boundary:

Python · request boundary
import inferencefort

with inferencefort.context_scope(user=user_id, customer_id=customer_id,
                                 agent_id="support-agent", thread_id=conversation_id):
    result = await agent.ainvoke(message)

Roll out without breaking production

Observe, propose, enforce
  1. Run in monitor mode and record every decision to a local audit file.
  2. Generate a starting policy from real traffic with if-policy learn. It lists the models, endpoints and MCP servers you actually use; anything else is denied.
  3. Check it with if-policy validate, review it, then switch to enforce.

Frequently asked questions

Does InferenceFort add latency to every call?

Policy decisions run in your process from a cached bundle. Network calls happen only for shared daily budgets, cross-process session taint when enabled, and any remote detectors you configure.

Do I need to change my agent code?

No per-call changes. Install the SDK, provide a policy or key, and optionally set identity at your request boundary. Node applications load a register hook or use explicit wrappers.

What happens if the InferenceFort service is unreachable?

Local enforcement continues from the cached policy. For checks that need the network, choose fail_open, fail_closed or fail_cached.

Is it only for prompt injection?

No. The same runtime enforces model access, content rules, egress, tool policy, PHI redaction, budgets and audit, attributed to the user, agent and customer.

Keep reading

Evaluate your agent workflow

Bring one agent workflow, the data it touches, and the actions you need to control. As a design partner, you shape the evaluation and review the policy decisions with our engineers.

Become a design partner →