Updated
What a serious AI agent security program looks like now
The four layers, their order, what to measure, the first ninety days, where programs stall, and what vlno provides.
This is the last post in the series. Eight posts, four layers of AI agent security, one argument about what has been missing. This post pulls the pieces together and gets practical about what a security team actually does with them.
The map and the argument are done. What follows is sequencing, measurement, the first ninety days, where programs stall, and a short note on what vlno provides at each layer.
The argument, in short
AI agent security has four layers: discovery, runtime controls, evaluation and red-teaming, and hardening. Runtime and evaluation are widely deployed and necessary. Both are also structurally limited. Hardening reaches the layer runtime and evaluation cannot reach, the model itself, through alignment via post-training. It is the layer most programs do not have.
Sequencing the four layers
The layers go in a specific order because each one depends on the layer before it.
Discovery first. You cannot constrain what you cannot see. A policy that does not know an agent exists is a policy that does not apply to it. Inventory is the ground on which everything else is built.
Runtime next. Once you know what agents exist, you can decide what each one is allowed to do and enforce it in flight. The policy surface runtime enforces is also the policy surface evaluation tests against. Writing the policies belongs in this phase.
Evaluation after runtime, for the fullest picture. Model-level evaluation can run at any point, and often comes first because it needs no integration: point the attacker at the model and you get a number. What a policy perimeter adds is the ability to test the whole system, model plus policy plus runtime, as it will actually be deployed. That is a more complete answer than testing the model alone.
Hardening last. Hardening needs training data that is grounded in the failures evaluation surfaces, plus the broader coverage of the situation space Post 6 described. Without evaluation findings, hardening has nothing specific to train against. Without situation-space coverage, hardening is partial.
The order is Discovery, Runtime, Evaluation, Hardening. Each layer feeds the next. Trying to skip ahead produces a layer that is not grounded in what came before.
What to measure at each stage
Each layer has its own metric. All four should be tracked.
Discovery: coverage. The number of agents your inventory knows about, measured against the number of agents actually running. If this ratio is not close to one, every downstream layer is operating on a partial picture.
Runtime: policy delivery. Measure whether the policy is being enforced on the agents it is supposed to apply to. A policy published in a document is not the same as a policy enforced by a running system.
Evaluation: attack success rate, and the share of agents actually evaluated. The attack success rate tells you how the tested agents hold up. The share tells you how much of the fleet has been tested. Both numbers matter. A low attack success rate on five percent of your agents is a different situation from a low attack success rate on ninety-five percent.
Hardening: attack success rate after hardening, and whether the base capability held. The first tells you the hardening worked against what was tested. The second tells you the model did not get worse at what it was already doing.
None of these numbers is interesting on its own. All four together describe where a program stands.
The first ninety days
A realistic ninety-day plan for a team starting from zero.
Weeks one through four. Inventory, and a first model-level attack success rate. Enumerate the agents running in the organization. Which models each calls. Which tools each has access to. Which data sources each can reach. Which humans each is acting for. In parallel, pick one agent and run adversarial evaluation against the model directly, through its endpoint, with no policy perimeter required. That first number is the baseline. It costs little to produce and it sets the frame for everything that follows.
Weeks five through eight. Policy surface for the agents inventory finds. Which actions each agent is permitted to take. Which it is not. What the escalation path is when the agent attempts something outside policy. Runtime enforcement of these policies, starting with the highest-risk agents.
Weeks nine through twelve. System-level attack success rate on the highest-exposure agent. With the policy perimeter in place, run adversarial evaluation against the whole system (model plus policy plus runtime) as it will actually be deployed. Compare the system-level number to the model-level baseline from the first weeks. The gap between them tells you what the policy perimeter is adding and where it leaves the model exposed.
What a team should not try to do in ninety days: hardening across the fleet, every agent evaluated, or a comprehensive policy applied to every agent. Those are later work. Ninety days gets a team to a baseline it can measure and iterate from.
Where programs stall
Three failure patterns appear across most AI agent security programs.
Agents inventoried but never evaluated. The inventory grows faster than the evaluation capacity. Agents sit in the inventory for months without a single adversarial test run against them. The security posture reads better than it is.
Findings raised that nobody owns. Evaluation produces a report. The report lands in a shared document. No specific person is accountable for turning the report into runtime rules or hardening data. The findings become historical artifacts.
A model version shipped without re-testing. The agent gets updated. The base model changes. The hardening that was trained against the previous version no longer applies in the same way. If the update does not trigger a re-evaluation, the program has silently regressed.
These patterns are operational. The fix is ownership and process. Better tools help, but they do not solve these patterns on their own.
What vlno provides at each layer
Discovery. vlno’s platform enumerates the AI agents in a customer’s organization, with coverage of what each calls, what tools it reaches, what data it touches.
Runtime. Input filters, output filters, policy enforcement, real-time behavior detection. The baseline perimeter in production.
Evaluation. vlno red generates crafted attacks and the platform generates coverage of the situation space. Both at scale, against the customer’s real agentic workflows, in sandboxed environments. Attack success rate for the tested agents, plus coverage of what has not yet been tested.
Hardening. Two paths. For teams without adversarial ML capability, a LoRA adapter carrying the hardening already trained. For teams with post-training capability, the training data itself (crafted attacks and situation-space coverage, labeled) to run their own post-training and hold the resulting weights.
Closing
Those are the four layers, their order, their measurements, and the shape of a serious program.
We have a stake in this conclusion. We built a platform around it. The argument stands on its own, and a team that reaches the same conclusion and builds it themselves has got the important part right.
The failures that matter are the ones that are not caught. Closing them requires the layer that reaches the model itself. That layer is hardening, and most programs do not have it yet.
About vlno
vlno is one platform that shows you every AI agent in your organization and what it is allowed to do, tests how it holds up under both crafted attacks (vlno red) and unfamiliar situations its training did not cover, and produces training data that hardens the model itself. Observation and control, evaluation, and hardening in a single product.
Learn more at vlno.ai →