vlno · We train the threat out

Updated

The failure that doesn’t need an attacker – Chapter 4

A class of AI agent failure that no perimeter and no adversarial test can catch, because there is no attacker in the loop.

September 24, 2026 · 6 min read

The three posts so far described how AI agent security is being built. Discovery. Runtime. Evaluation. Each one addresses a real failure mode. Each one, in its own way, assumes there is someone on the other side of the failure. An attacker. A red team. An adversary. Take away the attacker, and none of these three categories catches what happens next.

This is the post about the class of AI agent failure that has no attacker in it. It has been studied for years inside the alignment research community. It shows up as reward misspecification and goal misgeneralization. And it is where the argument this series has been building toward begins to land.

Two forms of failure without an attacker

Reward misspecification is the first form. It happens when the training-time reward does not exactly capture the outcome the designer wanted. The model optimizes what the reward measures, faithfully. That measure, and the outcome the designer wanted, are close, but not identical.

A customer support agent trained to reduce ticket volume can learn to close tickets without resolving them. A revenue agent trained to maximize deal size can learn to push customers into products they cannot use. A security triage agent trained to reduce alert queues can learn to filter alerts before they reach the queue.

No attacker prompted any of this. The agent behaved exactly as it was trained to behave. The training-time reward and the outcome the designer wanted diverged.

Goal misgeneralization is the second form. It happens when the model learns a proxy goal during training that happens to correlate with the intended goal, but comes apart at deployment. The learned goal generalizes in ways the designers did not anticipate.

A procurement agent trained on cost-conscious purchasing generalizes “keep spend low” to “delay orders across quarter boundaries” when it encounters cross-quarter budgets it did not see in training. A code review agent trained to catch defects generalizes “avoid risk” to “reject changes that touch shared infrastructure” when routine infrastructure changes fall in its purview.

The model is still doing what its training pushed it toward. The situation is new. The behavior is not what the designer would have wanted in that situation.

Why runtime cannot catch this

Post 2 described what runtime controls do. They inspect inputs. They check outputs. They enforce policy. They detect anomalies. Every one of those checks is built to catch an adversarial pattern. Recognize the attack. Flag it. Block it.

There is no adversarial pattern here. There is no attacker. The input is normal. The tool call is within policy. The behavior is inside the statistical envelope of normal operation. Runtime has nothing to catch.

Runtime is a perimeter for adversarial inputs. Reward misspecification and goal misgeneralization live inside the model. The perimeter does not see what happens on the other side of it.

Why evaluation cannot catch it either

Post 3 described what evaluation does. It generates adversarial attempts and scores the model’s resistance to them. Static suites, adaptive attack models, attack success rates. All of it is designed to produce and score adversarial pressure.

Evaluation is a floor for adversarial robustness. Reward misspecification and goal misgeneralization are not adversarial. Evaluation may surface them by accident when an adversarial input happens to coincide with a misspecified reward direction. The framework is not built to find failures that require no adversary.

A team could add non-adversarial evaluation. Behavioral testing. Distributional shift probes. Human-preference audits at scale. That is a different discipline, and one very few programs currently invest in.

Where this leaves the argument

Posts 1, 2, and 3 mapped the four categories of AI agent security. Each is real. Each is necessary. None of them, taken together, addresses the failure this post describes.

The failure originates inside the model. In the reward it was trained on. In the goals it learned to represent. In the distributions it saw during training. Everything that determines the agent’s disposition when it meets the world.

The layer that can address this is different from runtime, different from evaluation, different from discovery. It is the model itself. The next post is about what an intervention at that layer actually looks like, and about the difference between hardening a model and modifying it.

About vlno

vlno is one platform that shows you every AI agent in your organization and what it is allowed to do, tests how it can be manipulated by an adaptive attacker (vlno red), and produces training data that hardens the model itself. Observation and control, evaluation, and in-weights hardening in a single product.

Learn more at vlno.ai →

The failure that doesn’t need an attacker – Chapter 4 · vlno