LLM Runtime Security
Enforce policy at the moment your application calls the model.
Gateways see traffic after it leaves. Prompt filters see text without context. InferenceFort runs inside your application, where it sees the real call: which model, which endpoint, which user, which tool. It decides before the request is sent.
Why enforce in process
Nothing to route around
A gateway only governs traffic sent through it. A direct provider call or a framework's own client skips it. Interception at the framework layer governs the call wherever it goes.
Full context
In process, InferenceFort knows the user, agent, customer and conversation, and what the agent did earlier in the session. A network proxy sees a request body.
No round trip
Policy arrives as a cached bundle and is evaluated locally. Allowed calls do not wait on a network hop.
One engine, every language
Decisions are implemented once, in a dependency-free Go core loaded by both the Python and TypeScript SDKs. The same policy gives the same answer in every service.
What happens on each call
- Your app calls a model or tool through a supported framework or client. Python patches frameworks as they are imported; Node uses a register hook or explicit wrappers.
- The SDK resolves the provider, the configured destination and the caller's identity.
- The core evaluates model access, content rules, egress, detectors and tool policy from the cached bundle.
- A blocked call raises
PolicyViolation, the only exception InferenceFort raises into your code. The request never leaves. - After the call, usage and the decision are recorded asynchronously. Cost records survive outages through a bounded local write-ahead log.
Controls enforced at runtime
| Control | What it does |
|---|---|
| Model access | Allow or deny models per user, agent or team. When policy is loaded, unmatched access is denied. |
| Content rules | Substring and regex rules that block or flag, on input, output or both. |
| Egress | Pin each provider to approved endpoints or regions. Host-anchored matching; local loopback is always allowed. |
| Detectors | Built-in local prompt-attack detector, Lakera, HTTP detection APIs described in config, or your own in-process function. |
| PHI redaction | Shaped identifiers are replaced with typed placeholders such as [SSN] before the model sees the prompt. |
| Tool policy | Blocked and approval-required tools, MCP server allow-lists and session-based exfiltration checks. |
| Budgets | Daily spend caps checked against a shared ledger. The only per-call network check. |
Coverage
| Surface | Python | TypeScript / Node |
|---|---|---|
| LangChain / LangGraph | Every BaseChatModel provider, including streaming | LangChain.js and LangGraph.js, including streaming |
| Vercel AI SDK | – | doGenerate and doStream |
| LiteLLM | completion / acompletion, plus a proxy guardrail | – |
| OpenAI / Anthropic SDKs | Supported client methods | Register hook or wrapOpenAI / wrapAnthropic |
| Tools | LangChain, CrewAI, MCP | MCP |
Get started
pip install "inferencefort[langchain]"
export INFERENCEFORT_POLICY_FILE="$PWD/policy.json" # or set INFERENCEFORT_KEY
python app.py # or: if-run app.pynpm install @inferencefort/ai-sdk
export INFERENCEFORT_POLICY_FILE="$PWD/policy.json"
node --import @inferencefort/ai-sdk/register.mjs app.mjsBuilt to stay out of the way
- Inert until configured: with no key and no policy file, calls pass through untouched.
- HTTP libraries are never monkey-patched. Only framework and client interception points are wrapped.
- Transport failures never surface as exceptions.
INFERENCEFORT_FAIL_MODEdecides between fail_open, fail_closed and fail_cached. - Streaming input is checked before the provider runs. Output controls cannot recall content already delivered.
Frequently asked questions
Do I have to pass callbacks to every model?
No. Supported frameworks are intercepted at their shared entry points, so every provider built on them is governed without per-model code.
Is streaming governed?
Yes, on the supported streaming paths. Input is checked before the provider runs; output rules run as the response is produced.
What if my framework is not listed?
Many providers are reachable through a supported framework such as LangChain, LiteLLM or the Vercel AI SDK. A custom HTTP client or custom tool loop is not governed automatically; verify coverage with a deliberate blocked call.
Keep reading
Validate runtime coverage
Bring one agent workflow, the data it touches, and the actions you need to control. As a design partner, you shape the evaluation and review the policy decisions with our engineers.
Become a design partner →