Stop Prompting for Safety: Outside the Model

You cannot prompt an autonomous AI agent into being secure.
If your security model relies on instructing a Large Language Model (an AI system trained to generate text and reason through natural language) to "never share private credentials" or "refuse unauthorized refunds," you do not have an architecture. You have a wish list.
At Aniccai, I work with technical leaders every week who want to deploy agentic workflows across internal databases, ticketing systems, and customer messaging. Almost invariably, their first draft of agent safety lives right inside the system prompt. They add two paragraphs of stern rules, run twenty test queries, and convince themselves the agent is production-ready. Then an edge case hits, an adversarial customer rephrases an order, or indirect data injection poisons the context window.
Suddenly, the company ledger issues an illegal payout.
AI agents need deterministic, out-of-band authorization and strict data-flow governance built completely outside the model's reasoning loop. Treating English instructions as a security perimeter is a fatal architectural mistake.
Key Takeaways
- System prompts are probabilistic guidance, never security perimeters; models cannot reliably police themselves.
- Shared credentials give agents sweeping access without human judgment, exposing internal data stores to context poisoning.
- Deterministic policy enforcement must sit at the tool boundary, inspecting arguments before backend execution occurs.
- Multi-turn interactions create cumulative risks, requiring session-level anomaly tracking rather than isolated prompt checks.
The Fatal Flaw of In-Context Guardrails
Prompt injection is an authorization bypass. It breaks the assumed hierarchy between the application designer and the end user by conflating instructions with untrusted data.
When you place security guardrails directly into the prompt context, you force a single fallible reasoner to juggle two competing goals: fulfilling the user's intent and policing itself. No statistical model does this with perfect fidelity. If an attacker passes indirect instructions inside a Jira ticket, a scraped web page, or an email thread, the model absorbs those instructions as operational context.
In their research paper, If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary (arXiv:2608.27646), Marc Millstone and co-authors demonstrated what happens when teams abandon prompt-only defense. Against mock enterprise environments across Jira and ServiceNow, trace failures dropped from 57.6% to 0.2% when shifting from prompted checks to out-of-band enforcement. That is a massive operational delta.
No amount of prompt tweaking eliminates that 57% failure rate. Deterministic boundaries do.
Credential Inheritance: Giving Agents God-Mode by Default
Most teams build agents by generating an API token for an existing human user and feeding it to the agent's environment. This is credential inheritance, an identity anti-pattern where an automated process receives a human's complete system privileges without possessing the context or discretion of that person.
When an agent inherits an engineer's full service key, it can query every table and every ticket. If a malicious input directs the agent to scan customer records, the backend happily returns the rows because the credential itself is fully valid. The backend API sees no attack. It only sees a valid API call carrying a legitimate bearer token.
In a comprehensive review, Authorization Architectures for Tool-Using AI Agents (arXiv:2609.15906), Rakesh Kumar Surapani and colleagues mapped this exact structural breakdown. They showed that safe agent operations require a strict principal hierarchy spanning the human user, the deployer, the agent orchestrator, and the tool endpoint.
Delegating authority must involve scoping every single hop. If an agent does not explicitly require broad database read permissions to answer a support ticket, those permissions must not exist in its runtime environment.
Out-of-Band Policy Enforcement at the Tool Boundary
The correct place to govern an autonomous agent is at the Policy Enforcement Point (PEP), a deterministic gateway that intercepts proposed tool calls between the model and downstream infrastructure.
When a model decides to invoke a tool, it emits a structured payload containing the tool name and arguments. That payload must pass through an out-of-band filter before the backend runs. In practical systems, this looks like a proxy that validates arguments, checks typed policies, and evaluates current operational state.
| Governance Layer | Where It Executes | What It Governs | Failure Mode Caught |
|---|---|---|---|
| In-Context Prompts | Inside LLM Context | Probabilistic behavioral guidance | Stylistic misalignments only |
| Edge Firewalls | Ingress Gateway | Raw payload sanitization & known patterns | Obvious direct prompt injections |
| Semantic Policy Engines | Pre-Tool Execution Hook | Intent alignment & business constraints | Complex social engineering attempts |
| Out-of-Band Proxies | Network / API Layer | Typed access grants, argument ranges, data masking | Unauthorized ledger updates & credential leaks |
| Telemetry Monitors | Continuous Audit Pipeline | Multi-turn frequency & cumulative volume | Salami attacks & gradual quota drains |
Google documented this defense-in-depth model in Build zero-trust AI agents that judge intent, not just syntax (Google Developers Blog). They demonstrated how syntax-level checks fail when an attacker uses polite, perfectly typed phrasing to request refunds for non-refundable digital goods. By inserting Semantic Governance Policies right in front of tool execution, their gateway inspects the argument parameters against business rules and halts execution before sensitive signing keys are ever touched.
Hardware providers are pushing this concept even lower in the stack. In NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring (NVIDIA Technical Blog), the authors describe running agent runtimes within kernel-isolated sandboxes and using dedicated DPUs (Data Processing Units) to monitor model communications out-of-band. The model cannot bypass the watchdog because the watchdog lives on the physical network path to compute resources.
Data-Flow Contagion and Multi-Turn Exploits
Security is not a single transaction. Complex agent tasks take multiple turns, reading data from one tool and feeding it to another. This creates data-flow contagion, where untrusted content read in step one poisons the reasoning context for step six.
Consider a user who asks an agent to process an order return across eight separate messages. In each turn, the user requests a modest $20 refund on a $149 order. Every individual request passes per-turn limits, which permit small adjustments under $30 without manager sign-off. But over eight turns, the agent dispenses $160, exceeding the original order value entirely.
Single-turn guardrails cannot see this problem. They judge every request in complete isolation.
To address this, researchers introduced AgentFlow: A Flow-Centric Policy Language and Framework for Securing LLM Agent Systems (arXiv:2608.22868). AgentFlow tracks data lineage across runtime execution edges, tracking tainted data as it passes through tools, sinks, and sub-agents. Across demanding agent evaluation suites, this tracking eliminated confirmed compromise without breaking core utility.
By monitoring state across turns, your architecture stops attackers from draining resources through subtle repeated calls.
Architecting the Deterministic Floor
How do you build this without massive corporate infrastructure? You implement a deterministic floor.
First, place a hard state machine around your agent. A state machine is a mathematical model of computation where an application transitions between finite, explicit stages based on strict rules. If an agent is in the "read order details" state, the proxy literally drops any attempt to invoke the "issue refund" tool, regardless of what the LLM reasons. The agent cannot skip steps because the API route simply does not exist for that stage.
Second, narrow data responses before they return to the agent. When an agent queries a customer record, your tool proxy should strip sensitive fields like social security numbers, hashed passwords, or internal notes. Do not trust the model to ignore them. Keep them out of the context window entirely.
Third, enforce aggregate transaction thresholds outside the agent. If a support workflow touches more than three records in ten seconds, or if a financial tool commits more than a set budget in a single session, trip a circuit breaker. Halt the process and alert a human.
This approach trades a tiny sliver of edge-case autonomy for enterprise survivability. And that is the exact trade every responsible business leader should make.
Sources
- If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary (arXiv:2608.27646)
- Authorization Architectures for Tool-Using AI Agents (arXiv:2609.15906)
- AgentFlow: A Flow-Centric Policy Language and Framework for Securing LLM Agent Systems (arXiv:2608.22868)
- Build zero-trust AI agents that judge intent, not just syntax (Google Developers Blog)
- NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring (NVIDIA Technical Blog)
FAQ
Why is prompt engineering insufficient for AI agent safety?
Prompts operate within the model's probabilistic reasoning context, where instructions and untrusted data share the same channel. Because the model cannot reliably distinguish system instructions from adversarial user inputs or poisoned data, prompt-based rules can be bypassed via prompt injection, subtle social engineering, or contextual manipulation.
What is an out-of-band policy enforcement point?
An out-of-band policy enforcement point is an independent software or hardware gateway positioned between the AI model and downstream tools. It evaluates proposed actions against hard deterministic business rules, typed schemas, and session policies before any database write or external API call executes, completely outside the model's control.
How do multi-turn attacks bypass single-turn AI guardrails?
Multi-turn attacks break a prohibited action into small, individually legitimate requests spread across multiple conversation steps. Because single-turn guardrails only inspect one input or tool call in isolation, they fail to catch the cumulative impact, such as multiple tiny refunds that eventually exceed an order limit.
Can API keys alone protect backend systems from autonomous agents?
No, standard API keys only verify that the incoming request belongs to an authenticated credential. If an agent is granted an administrator's or employee's bearer token, any tool call it makes will appear authorized to the backend, even if the agent is acting on malicious instructions.
Things to Remember
- Never trust an LLM to enforce its own operational security boundaries.
- Strip sensitive internal fields at the proxy layer before payloads enter model context.
- Track cumulative session velocity and state machines deterministically to prevent multi-turn exploitation.
Take a look at your current agent architecture today: what is the single most destructive tool your agent can call, and is there anything between that model output and your database other than a system prompt?
Thinking about an agent for one of your workflows?
Most agent projects fail on scope, not on the model. A pilot picks one workflow and proves it end to end.
Related Articles
Explore all AI Agents
The 11% Club: Why Most Agentic AI Projects Die in Pilot
Discover why only 11% of agentic AI projects reach production. Learn to navigate the Gartner Paradox, avoid agent washing, and build robust AI infrastructure.

The Autonomy Trap: Why AI Agents Need Boundaries, Not Freedom
Why full AI agent autonomy fails in production. Learn how delegation envelopes, approval modes, and calibrated blast radiuses create safe, reliable ROI.

The Execution Split: Why Production Agents Require a Control Plane
Why do AI agents fail in production? Learn why decoupling reasoning from execution via a Control Plane is the key to safe, scalable enterprise automation.