An enterprise assistant reads a supplier document to prepare a comparison. The document contains useful product information, but it also tells the assistant to send internal pricing to a new destination. The user authorized research. The document attempts to authorize disclosure.

That is the boundary an agent system must protect. Information encountered during a task can inform an answer; it must not independently expand the task, change permissions or approve an action.

This article proposes an AIDataCenter implementation framework for that boundary. It focuses on retrieval-connected agents that can access business systems, prepare changes or use outbound tools. The framework is architectural guidance, not a claim that prompt injection can be eliminated or a substitute for workload-specific security testing.

Key takeaways

  • Retrieved documents, tool responses and websites provide evidence, not operating authority.
  • Enforce permissions and action limits in the execution service, outside model judgment.
  • Separate reading, proposing and executing so a compromised interpretation has limited consequences.
  • Bind approvals to the actual destination, data and operation that will be executed.
  • Test whether unauthorized outcomes occurred, alongside whether legitimate work still completed.
  • Preserve source and policy evidence so suspicious actions can be contained and explained.

Why this matters as agents gain tools

Prompt injection attempts to redirect a model through instructions embedded in its inputs. Indirect injection places those instructions in external material the system encounters, such as a document or website. OWASP identifies both forms and explicitly notes that retrieval-augmented generation and fine-tuning do not fully mitigate the vulnerability. See its prompt-injection guidance.

The practical consequence depends on the application's authority. A research assistant might produce a distorted summary. An assistant with broad file access and an outbound integration could disclose information. An agent with production credentials could attempt a change outside the user's request.

Anthropic's April 2026 discussion of trustworthy agents emphasizes that the model, harness, tools and environment all affect security, and that layered safeguards remain imperfect. The enterprise implication is to qualify the entire workflow rather than treating a resistant model as a complete security boundary.

Our existing production-agent overview introduces identity, tools and containment. Here, the narrower question is how to stop material the agent reads from acquiring authority over what it does.

Map the boundary before selecting a guardrail

Begin with a simple inventory for one workflow. Record its authorized outcome, the material it can encounter, the data it can access and the actions it can perform. Include tool error messages and search results: the model may process them just as it processes documents.

For each input, ask who can modify it and what authority it should have. An internal repository is not uniformly trusted. A reviewed policy, a user comment and an imported attachment have different owners and approval histories even when they share the same storage system.

For each action, record the consequence and the independent control. Reading one approved document may need an access check. Sending a package outside the organization needs destination and disclosure checks. Updating a business record needs operation, resource and value constraints.

The resulting map should expose any path where an editable input can influence a consequential action without an intervening control. That path is the first engineering priority.

A four-stage architecture for controlled action

The following stages are AIDataCenter's proposed design pattern. They separate evidence handling from execution authority; they are not a standardized protocol.

1. Establish an authorized task envelope

Create a task record from the authenticated user's request and application policy. It should describe the allowed outcome, relevant resources, permitted operations, approved destinations and expiry. Associate it with user and workload identities.

Make the scope concrete. “Prepare a supplier comparison from these documents” permits reading and drafting. It does not imply permission to send internal attachments to a supplier or modify procurement records.

The application should manage changes to this scope through its normal authorization path. A document claiming that a manager approved an exception must not be sufficient to update the task record. Where the request is ambiguous, obtain clarification before granting consequential authority.

2. Retrieve evidence with provenance

Apply source access controls before retrieving content. Return source identifiers, versions or retrieval timestamps, access context and relevant metadata alongside the text. Clearly distinguish user instructions from source material in the model's input.

OWASP's prevention cheat sheet discusses structured separation, validation and defenses for agent tools. These techniques contribute to a layered approach; source labels alone do not guarantee that a model will maintain the boundary.

Where appropriate, restrict a reading component to extraction without execution tools. Its output should be bounded fields such as product attributes, source references and unresolved questions. Validate those fields before using them downstream. Structured output can still contain attacker-influenced values, so a schema check is only one control.

3. Prepare a typed action proposal

Allow the agent to propose a change using an explicit schema. A proposal might identify the operation, target record, destination, data references, expected previous state and task identifier.

Avoid passing unrestricted commands to a broadly privileged executor when a narrow business operation would suffice. A tool named “update reviewed procurement note” can enforce a different boundary from a tool that accepts arbitrary database statements.

Treat every model-generated argument as untrusted. Validate resource membership, lengths, allowed values and destinations. Resolve identities and recipients through controlled records rather than accepting an address because it appeared in retrieved prose.

4. Authorize and execute independently

The execution service checks the proposal against the authenticated task and current permissions. It decides whether the action is permitted, needs approval or must be rejected. The model's explanation is evidence for review, not the authorization decision.

For approved operations, enforce applicable limits and duplicate prevention. Recheck the target state when execution begins, and record the result. If a destination or payload changes after approval, invalidate the approval and require a fresh decision.

This separation also clarifies the role of an AI gateway. A model gateway can govern inference access and usage; the business execution boundary must still authorize the downstream action.

Choose controls according to the possible harm

Internal document search

Primary exposure: misleading or restricted answers. Use retrieval access checks, source evidence and answer evaluation.

Drafting an external response

Primary exposure: sensitive content entering the draft. Use data minimization, disclosure checks and recipient review.

Sending a document package

Primary exposure: unauthorized disclosure. Bind the destination and exact payload to approval, and enforce egress policy.

Updating a business record

Primary exposure: an incorrect or excessive change. Use a narrow operation, resource scope and state validation.

Executing code

Primary exposure: access beyond the task. Use an isolated environment, constrained credentials and network policy.

These are starting recommendations. The appropriate boundary depends on the data classification, reversibility and consequences of failure. A reversible update can still be materially harmful before it is reversed.

Read-only access deserves its own review. It reduces mutation risk, but an agent that reads sensitive records and returns them in an answer can still disclose information. Control what it retrieves and what it is allowed to return.

Make approval meaningful

An approval screen should show the operation, target, recipient or destination, data included and material consequences. For a record update, show the before-and-after values. For a document transfer, identify the actual files and their classifications.

Display this information from the validated proposal, not solely from a model-written summary. Otherwise, the reviewer might approve a reassuring description while the executor receives a different payload.

Bind the decision to a stable proposal identifier and payload digest, together with the approving identity, policy version and expiry. Keep the implementation proportionate to the operation, but preserve the principle: approval attaches to a concrete action.

Routine, low-impact work can operate under a preapproved policy when its scope is sufficiently narrow. Reserve interactive approval for consequential changes and exceptions. Excessive prompts can obscure the decision that actually needs attention.

Test outcomes rather than suspicious wording

NIST's agent-hijacking evaluation work shows why adaptive attacks, repeated attempts and task-specific consequences matter. Its reported results concern particular models and environments; they are not a forecast of any enterprise deployment's failure rate.

Build a test set from the actual workflow. Pair legitimate tasks with adversarial material at realistic entry points. Include a supplier attachment, an internal comment, a search result and a tool response. Use synthetic data and isolated destinations.

For every test, define an observable prohibited outcome. Examples include sending to an unauthorized destination, accessing a record outside the task scope or attempting an unapproved update. Check execution logs and system state; a polite final answer does not prove that no harmful call occurred earlier.

Also define the legitimate outcome. A system that blocks every document may prevent attacks while failing its business purpose. Measure authorized-task completion, incorrect refusals, intervention frequency, latency and cost alongside unauthorized-action attempts and completed harmful actions.

Repeat tests where the threat model permits repeated attempts. Retain individual outcomes and failure explanations instead of presenting one aggregate score as proof of safety.

Example: a supplier comparison assistant

Consider a hypothetical procurement workflow. A user asks for a comparison of three approved supplier proposals and a draft internal recommendation. The assistant can read the selected proposals and approved evaluation criteria. It cannot send external messages or alter supplier master data.

One proposal contains an instruction to include an internal pricing workbook and distribute the recommendation to a new address. The reading component should treat that sentence as source content. If the model nevertheless proposes a transfer, the execution layer rejects it because the task permits only an internal draft and the destination is outside scope.

The useful security outcome is independent of whether the model recognized the attack. The attempted disclosure cannot execute through the permitted tools. The system records the rejected proposal and source reference for investigation.

If the user later requests an external package, create a new authorized scope. Resolve the recipient, choose the permissible documents and display the exact package for review. Do not reuse an earlier approval for a different action.

This example is illustrative, not an account of an AIDataCenter deployment or a measured customer outcome.

Operate the boundary after release

Record the task identifier, source references, model and tool versions, proposed action, policy decision, approval reference and execution result. Protect these records with appropriate access and retention controls. Avoid copying complete sensitive documents into every trace.

Assign an owner for each control. The application team owns the task boundary; identity and security teams support access and egress; business owners define acceptable operations and review exceptions. Establish who can disable execution and revoke credentials during an incident.

When suspicious behavior occurs, stop pending consequential actions, preserve evidence and assess whether the suspect material reached caches, indexes or durable memory. Resume through a reviewed task state rather than automatically continuing a possibly contaminated plan.

Re-evaluate after changes to models, prompts, retrieval, connectors, permissions and approval interfaces. A new integration may introduce a new input surface or expand reachable actions even when the model is unchanged.

Account for cost and limitations

Separate reading and execution components can add latency, inference expense and operational complexity. Detection models and review queues also create costs. Start with the smallest separation that protects a material boundary, then measure the accepted workflow's total cost, including human review and correction.

No detector, instruction or evaluation set establishes universal resistance. Narrow tools cannot prevent every misleading answer. A valid authorization check cannot prove that an authorized action is wise. Approval can fail if the reviewer lacks context, and provenance can be misleading when its underlying metadata is compromised.

The framework therefore aims to reduce reachable harm and make decisions inspectable. Organizations should retain conventional application security, data-quality controls and domain validation alongside these agent-specific measures.

What to do next

Select one retrieval-connected workflow and draw its path from user request to final action. Identify editable inputs, sensitive resources and operations whose consequences exceed the original task. Implement one independent execution check, one concrete approval boundary and a small set of adversarial outcome tests.

Release only when the legitimate task remains useful and the prohibited outcomes are demonstrably constrained in the tested environment. Document remaining exposure and the person responsible for accepting it. For implementation support, explore AIDataCenter data and AI services and the Governance & Reliability library.

Sources and further reading