vlno · We train the threat out

Updated

Hardening is the missing piece of the AI agent security puzzle

Runtime and evaluation are necessary but structurally behind. Hardening is the layer that closes the failure family at its source, and most programs don’t have it.

October 27, 2026 · 6 min read

This series has walked through discovery, runtime controls, evaluation and red-teaming, the class of failure that has no attacker in it, the intervention that reaches into the model itself, and the training data that intervention requires. This post pulls the argument together and says plainly what the whole thing has been building toward.

The AI agent security field has a set of layers that work. Runtime constrains what a deployed agent can do. Evaluation measures how the agent holds up under pressure. Both are necessary. Both should be in every serious deployment. Both are also structurally limited in a way that leaves the failures that matter most unaddressed.

Runtime is necessary and structurally behind

We build runtime controls ourselves, which is part of why we can describe their limits precisely.

Runtime controls catch adversarial patterns. Every pattern they catch is a pattern that has been seen before, catalogued, encoded into a rule or a classifier. This is what runtime does well and where most current AI agent security investment goes.

The structural limit is that adversarial attacks are adaptive. Attackers move faster than rule writers. Every new attack that succeeds is, by definition, one that the existing rules did not anticipate. The distance between what runtime already catches and what attackers are already trying has never been zero. Against a determined adversary it will not be zero.

Runtime is not going to close that gap by getting better at recognizing patterns. Pattern recognition is downstream of the attacker’s move, and the attacker sets the pace.

Evaluation is a floor, not a fix

Evaluation generates adversarial pressure at scale and scores how the model holds up. The output is a measurement, usually an attack success rate. The number is real, it is valuable, and any serious program has it.

The structural limit is that a measurement describes the model as it is. It does not change the model. The day after the evaluation is finished, the same weaknesses that produced the number are still there. The attacks that succeeded still work. The report tells you where you stand. It does not move you.

For a measurement to move you, it has to be fed into an intervention that closes what the measurement found. Runtime patches close one attack at a time. Hardening closes the family. Without an intervention on the other end, evaluation is a report about a problem that has not been solved.

Hardening closes the failure family at its source

Hardening reaches the layer runtime and evaluation cannot reach. The model itself. The mechanism is alignment via post-training. It teaches the model to recognize when it is outside its training coverage and to respond correctly. Sometimes that means resisting a crafted attack. Sometimes it means handling an unfamiliar situation the way the operator would want. Both cases are corrected at the same layer.

Runtime blocks a specific known attack. Evaluation reports on a specific known failure. Hardening closes the whole family of failure the training data represents. When a new attack from the family arrives, hardening holds. When an unfamiliar situation from the family arrives, hardening holds. The model’s disposition has changed at the source.

This is a different kind of closing than runtime provides. It does not require the attack to have been seen and catalogued. It requires the family the attack belongs to to have been covered in the training data. Once the family is covered, the specific instances follow.

Hardening is the missing piece

Serious AI agent security programs already have discovery. They have runtime. They increasingly have evaluation. Very few of them have hardening.

The reason is not that hardening does not work. It works. Adversarial ML research has demonstrated for years that post-training can shift a model’s disposition on targeted subspaces without touching the rest of the model. The reason hardening is not deployed at scale is that most programs do not have the training data or the post-training skill it requires. The data and the post-training skill have rarely been available together, which is a market condition more than a product problem.

This is what has kept AI agent security programs stuck at runtime and evaluation, layers that are necessary but that do not close the failure family.

What a serious program looks like now

A serious AI agent security program has all four layers. Discovery, so it knows what agents exist and what they are allowed to do. Runtime, so it constrains what the deployed agent can do. Evaluation, so it knows where the model stands and where it fails. Hardening, so the failure family is closed at its source.

The layers are complementary. Discovery makes the problem visible. Runtime blocks what can be blocked. Evaluation measures what runtime cannot catch. Hardening addresses what evaluation surfaces at the layer where the failure originates.

A program with only the first three layers is doing serious work. It is also perpetually behind, because runtime is downstream of the attacker’s move and evaluation is downstream of the last round of testing. A program with all four is not perpetually behind, because hardening addresses the failure at the source, and the source does not move at the pace of the next attack.

Where the argument leaves off

Hardening is the missing piece of the AI agent security puzzle.

The final post in this series is about what a security team that decides to close the gap does now. What to deploy, what to measure, and what a serious AI agent security program looks like in practice.

About vlno

vlno is one platform that shows you every AI agent in your organization and what it is allowed to do, tests how it holds up under both crafted attacks (vlno red) and unfamiliar situations its training did not cover, and produces training data that hardens the model itself. Observation and control, evaluation, and hardening in a single product.

Learn more at vlno.ai →
Hardening is the missing piece of the AI agent security puzzle – Chapter 7 · vlno