Stop Prompting for Safety: Outside the Model

R
Roy Saadon
Oct 6, 2026
8 min read
Stop Prompting for Safety: Outside the Model

You cannot prompt an autonomous AI agent into being secure.

If your security model relies on instructing a Large Language Model (an AI system trained to generate text and reason through natural language) to "never share private credentials" or "refuse unauthorized refunds," you do not have an architecture. You have a wish list.

At Aniccai, I work with technical leaders every week who want to deploy agentic workflows across internal databases, ticketing systems, and customer messaging. Almost invariably, their first draft of agent safety lives right inside the system prompt. They add two paragraphs of stern rules, run twenty test queries, and convince themselves the agent is production-ready. Then an edge case hits, an adversarial customer rephrases an order, or indirect data injection poisons the context window.

Suddenly, the company ledger issues an illegal payout.

AI agents need deterministic, out-of-band authorization and strict data-flow governance built completely outside the model's reasoning loop. Treating English instructions as a security perimeter is a fatal architectural mistake.

Key Takeaways

  • System prompts are probabilistic guidance, never security perimeters; models cannot reliably police themselves.
  • Shared credentials give agents sweeping access without human judgment, exposing internal data stores to context poisoning.
  • Deterministic policy enforcement must sit at the tool boundary, inspecting arguments before backend execution occurs.
  • Multi-turn interactions create cumulative risks, requiring session-level anomaly tracking rather than isolated prompt checks.

The Fatal Flaw of In-Context Guardrails

Prompt injection is an authorization bypass. It breaks the assumed hierarchy between the application designer and the end user by conflating instructions with untrusted data.

When you place security guardrails directly into the prompt context, you force a single fallible reasoner to juggle two competing goals: fulfilling the user's intent and policing itself. No statistical model does this with perfect fidelity. If an attacker passes indirect instructions inside a Jira ticket, a scraped web page, or an email thread, the model absorbs those instructions as operational context.

In their research paper, If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary (arXiv:2608.27646), Marc Millstone and co-authors demonstrated what happens when teams abandon prompt-only defense. Against mock enterprise environments across Jira and ServiceNow, trace failures dropped from 57.6% to 0.2% when shifting from prompted checks to out-of-band enforcement. That is a massive operational delta.

No amount of prompt tweaking eliminates that 57% failure rate. Deterministic boundaries do.

Credential Inheritance: Giving Agents God-Mode by Default

Most teams build agents by generating an API token for an existing human user and feeding it to the agent's environment. This is credential inheritance, an identity anti-pattern where an automated process receives a human's complete system privileges without possessing the context or discretion of that person.

When an agent inherits an engineer's full service key, it can query every table and every ticket. If a malicious input directs the agent to scan customer records, the backend happily returns the rows because the credential itself is fully valid. The backend API sees no attack. It only sees a valid API call carrying a legitimate bearer token.

In a comprehensive review, Authorization Architectures for Tool-Using AI Agents (arXiv:2609.15906), Rakesh Kumar Surapani and colleagues mapped this exact structural breakdown. They showed that safe agent operations require a strict principal hierarchy spanning the human user, the deployer, the agent orchestrator, and the tool endpoint.

Delegating authority must involve scoping every single hop. If an agent does not explicitly require broad database read permissions to answer a support ticket, those permissions must not exist in its runtime environment.

Out-of-Band Policy Enforcement at the Tool Boundary

The correct place to govern an autonomous agent is at the Policy Enforcement Point (PEP), a deterministic gateway that intercepts proposed tool calls between the model and downstream infrastructure.

When a model decides to invoke a tool, it emits a structured payload containing the tool name and arguments. That payload must pass through an out-of-band filter before the backend runs. In practical systems, this looks like a proxy that validates arguments, checks typed policies, and evaluates current operational state.

Governance LayerWhere It ExecutesWhat It GovernsFailure Mode Caught
In-Context PromptsInside LLM ContextProbabilistic behavioral guidanceStylistic misalignments only
Edge FirewallsIngress GatewayRaw payload sanitization & known patternsObvious direct prompt injections
Semantic Policy EnginesPre-Tool Execution HookIntent alignment & business constraintsComplex social engineering attempts
Out-of-Band ProxiesNetwork / API LayerTyped access grants, argument ranges, data maskingUnauthorized ledger updates & credential leaks
Telemetry MonitorsContinuous Audit PipelineMulti-turn frequency & cumulative volumeSalami attacks & gradual quota drains

Google documented this defense-in-depth model in Build zero-trust AI agents that judge intent, not just syntax (Google Developers Blog). They demonstrated how syntax-level checks fail when an attacker uses polite, perfectly typed phrasing to request refunds for non-refundable digital goods. By inserting Semantic Governance Policies right in front of tool execution, their gateway inspects the argument parameters against business rules and halts execution before sensitive signing keys are ever touched.

Hardware providers are pushing this concept even lower in the stack. In NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring (NVIDIA Technical Blog), the authors describe running agent runtimes within kernel-isolated sandboxes and using dedicated DPUs (Data Processing Units) to monitor model communications out-of-band. The model cannot bypass the watchdog because the watchdog lives on the physical network path to compute resources.

Data-Flow Contagion and Multi-Turn Exploits

Security is not a single transaction. Complex agent tasks take multiple turns, reading data from one tool and feeding it to another. This creates data-flow contagion, where untrusted content read in step one poisons the reasoning context for step six.

Consider a user who asks an agent to process an order return across eight separate messages. In each turn, the user requests a modest $20 refund on a $149 order. Every individual request passes per-turn limits, which permit small adjustments under $30 without manager sign-off. But over eight turns, the agent dispenses $160, exceeding the original order value entirely.

Single-turn guardrails cannot see this problem. They judge every request in complete isolation.

To address this, researchers introduced AgentFlow: A Flow-Centric Policy Language and Framework for Securing LLM Agent Systems (arXiv:2608.22868). AgentFlow tracks data lineage across runtime execution edges, tracking tainted data as it passes through tools, sinks, and sub-agents. Across demanding agent evaluation suites, this tracking eliminated confirmed compromise without breaking core utility.

By monitoring state across turns, your architecture stops attackers from draining resources through subtle repeated calls.

Architecting the Deterministic Floor

How do you build this without massive corporate infrastructure? You implement a deterministic floor.

First, place a hard state machine around your agent. A state machine is a mathematical model of computation where an application transitions between finite, explicit stages based on strict rules. If an agent is in the "read order details" state, the proxy literally drops any attempt to invoke the "issue refund" tool, regardless of what the LLM reasons. The agent cannot skip steps because the API route simply does not exist for that stage.

Second, narrow data responses before they return to the agent. When an agent queries a customer record, your tool proxy should strip sensitive fields like social security numbers, hashed passwords, or internal notes. Do not trust the model to ignore them. Keep them out of the context window entirely.

Third, enforce aggregate transaction thresholds outside the agent. If a support workflow touches more than three records in ten seconds, or if a financial tool commits more than a set budget in a single session, trip a circuit breaker. Halt the process and alert a human.

This approach trades a tiny sliver of edge-case autonomy for enterprise survivability. And that is the exact trade every responsible business leader should make.

Sources

FAQ

Why is prompt engineering insufficient for AI agent safety?

Prompts operate within the model's probabilistic reasoning context, where instructions and untrusted data share the same channel. Because the model cannot reliably distinguish system instructions from adversarial user inputs or poisoned data, prompt-based rules can be bypassed via prompt injection, subtle social engineering, or contextual manipulation.

What is an out-of-band policy enforcement point?

An out-of-band policy enforcement point is an independent software or hardware gateway positioned between the AI model and downstream tools. It evaluates proposed actions against hard deterministic business rules, typed schemas, and session policies before any database write or external API call executes, completely outside the model's control.

How do multi-turn attacks bypass single-turn AI guardrails?

Multi-turn attacks break a prohibited action into small, individually legitimate requests spread across multiple conversation steps. Because single-turn guardrails only inspect one input or tool call in isolation, they fail to catch the cumulative impact, such as multiple tiny refunds that eventually exceed an order limit.

Can API keys alone protect backend systems from autonomous agents?

No, standard API keys only verify that the incoming request belongs to an authenticated credential. If an agent is granted an administrator's or employee's bearer token, any tool call it makes will appear authorized to the backend, even if the agent is acting on malicious instructions.

Things to Remember

  • Never trust an LLM to enforce its own operational security boundaries.
  • Strip sensitive internal fields at the proxy layer before payloads enter model context.
  • Track cumulative session velocity and state machines deterministically to prevent multi-turn exploitation.

Take a look at your current agent architecture today: what is the single most destructive tool your agent can call, and is there anything between that model output and your database other than a system prompt?

Thinking about an agent for one of your workflows?

Most agent projects fail on scope, not on the model. A pilot picks one workflow and proves it end to end.

Related Articles

Explore all AI Agents

We use cookies to understand how the site is used and which content helps. No advertising cookies, and we never sell or share your information for marketing. Privacy Policy