Vijay Ravi.
All insights

AI Security Notes · 01

What agentic security actually means

By Vijay Ravi · Foundations

An AI system becomes a different security problem when it can act on the world.

A chatbot that drafts an answer and an agent that changes a production system may use the same underlying model. Their risk is very different. The second system can retrieve information, select tools and execute a sequence of actions. Security therefore needs to cover the whole workflow.

Start with the action

Consider a security agent investigating a suspicious login. It can gather sign-in events, summarize evidence, and recommend a response. If it can also disable an account, revoke sessions or modify access, a mistaken decision can disrupt the business. The useful question becomes: which actions should this agent be able to take, under what conditions, and with whose authority?

The answer depends on the sensitivity of the data, the consequences of the action, and whether the result can be reversed. Retrieving approved telemetry is different from disabling a privileged account. Both need controls, but they should not require the same approval process.

Four boundaries to design

IdentityWho owns the agent and its credentials?
AuthorityWhich actions and resources are allowed?
ExecutionWhat is checked before tools run?
EvidenceCan actions be traced and stopped?

Identity. Give the agent an attributable identity and an accountable owner. Separate the identity of the initiating human from the identity executing the task. Avoid generic shared credentials that obscure responsibility.

Authority. Define the tools, resources and actions the agent can access. Use narrow, time-limited permissions where possible. A user's request should not automatically grant the agent all of that user's privileges.

Execution. Enforce policy outside the model. A prompt asking the model to behave safely is useful guidance, but it does not replace a tool gateway, scoped authorization or a required human approval. Validate the proposed action and its parameters before execution.

Evidence. Record requests, tool calls, authorization decisions, approvals and outcomes. Protect sensitive information in those records. Operators should be able to interrupt a workflow and revoke its credentials when necessary.

Treat retrieved content as untrusted

An agent may read documents, websites, emails or tool results that contain instructions. Those instructions can conflict with the user's task. A malicious document could try to redirect the agent toward exporting data or invoking a different tool.

Keep retrieved content separate from trusted instructions and constrain what downstream tools can do. Test with adversarial content across a sequence of steps, including how one tool's output influences the next action. A model refusing a malicious prompt does not prove the full system is secure.

Make autonomy proportional to consequence

A practical rollout starts with observation and recommendations. Add narrowly scoped, reversible actions once the workflow and controls are proven. Require explicit approval for higher-impact actions, supported by evidence the reviewer can understand.

For the suspicious-login example, the agent might autonomously collect approved telemetry and prepare a case. Disabling a critical service account could require an analyst's approval and a recovery plan. The specific boundary should reflect business impact, rather than an arbitrary confidence score.

Three questions for the next design review

  1. Can we enumerate the agent's permitted actions and identify the controls enforcing each one?
  2. Can we reconstruct what happened, including the authority and evidence behind an action?
  3. Can we stop the workflow and recover when it behaves unexpectedly?

Agentic security brings established security principles into a system with dynamic planning and tool use. The work is to make those principles enforceable across every action—not just describe them in a policy.

Further reading

NIST AI Risk Management Framework · OWASP GenAI Security Project