Updated
What model-level hardening actually looks like
The intervention that reaches inside the model, what it changes, and what it deliberately leaves alone.
The previous post described a class of AI agent failure that no perimeter and no adversarial test can catch. Reward misspecification and goal misgeneralization. Failures that originate inside the model, not in the traffic around it. If the failure lives inside the model, the intervention has to reach inside the model too.
That intervention exists. It is called alignment, done through post-training. This post is about what alignment reaches inside the model, the two capabilities it takes to do it well, and how a security team gets access to it without becoming an adversarial ML lab.
A targeted intervention, not a general modification
The word “training” makes product owners nervous. They hear “you are going to train our model” and they picture something expensive, slow, and risky. They picture the model getting worse at what it was already doing. They picture the whole capability profile of the system shifting. Then they say no.
That reaction is reasonable. It is also not what alignment post-training does.
Alignment reaches into a specific part of the model’s behavior. The part that governs what the model does when it meets a situation its training did not cover. Sometimes that situation is crafted by an attacker. Sometimes it is simply unfamiliar. Both produce the same underlying failure, the model pursuing what it learned in a way nobody intended, and both are corrected in the same place. Alignment updates that part. It leaves the rest alone.
The base capabilities of the model are preserved. What the agent can help with, what it can reason about, what tools it can use, what it knows about the domain, how it handles traffic its training already covered. All of that remains. What changes is how the agent responds when its training did not cover the situation in front of it.
This is what hardening the model looks like. The change reaches only the part that governs behavior outside training coverage. The rest of the model is left untouched.
Modifying a model changes its capability profile. Retraining. Fine-tuning for a new task. Shifting what the model can do. Hardening leaves the capability profile alone. It changes only how the model behaves when it meets something its training did not cover.
Take the procurement agent from the previous post. Trained on cost-conscious purchasing. At deployment it meets cross-quarter budgets its training never covered and starts delaying orders across quarter boundaries. No attacker in the loop. Alignment reaches that behavior. The training data has to include the shape of the situation (an unfamiliar budget structure) and the response the operator considers correct. The trained model then handles that situation the way the operator intended, without any change to how it handles the purchasing situations it was already good at.
At the technical layer, alignment post-training is training data plus a training procedure. The training data consists of examples spanning both crafted attacks and unfamiliar situations, each paired with the correct response. The training procedure updates the model on that data with parameter changes small enough to keep the update targeted, and with an evaluation loop that checks the base capabilities have not shifted.
The result: the model learns to recognize when it is outside its training coverage and to respond in the way alignment trained it to. Sometimes that is a crafted attack. Sometimes it is an unfamiliar situation with no attacker in it. When the trained model meets traffic that its training already covered, it responds exactly as it always did.
The two capabilities it takes to do this well
Alignment post-training done well requires two capabilities. If either is missing, the alignment does not stick.
The first is high-quality training data that covers both kinds of situations the model will fail in. The crafted case is adversarial data: attacks generated at scale, each paired with the correct response. The uncrafted case is coverage of the situation space: unfamiliar but benign situations where the model’s learned goals come apart with no adversary present. Both cases have to be labeled with what the correct response is.
The second case is the harder one. Generating attacks is a discipline the field has developed over the last several years. Adversarial ML teams know what they are doing. Generating the situations where a model’s learned goals break down, with no malicious content at all, is a different discipline. It requires coverage of the situation space rather than adversarial ingenuity. Very few programs do it, and it is where a lot of the failures that matter live.
The second capability is post-training skill. Knowing which parameters to update. How much learning to apply. How to preserve base capabilities while shifting behavior in a targeted subspace. How to evaluate whether the alignment held under both crafted attacks and unfamiliar situations, and whether the base capabilities held under normal load. This is a specialized craft.
Either capability missing, and the alignment fails in a specific way. Without the data, a good training procedure has nothing worth training on. Without the skill, the data is applied in ways that damage the base model.
Most security teams have neither
Most security teams do not have this capability in-house. Their expertise is security, policy, and the products they protect. Building a training corpus that covers both crafted attacks and unfamiliar situations, and running a post-training procedure to embed alignment into a production model, is a capability most organizations have never had reason to develop.
Some teams have part of the capability. A red team that can generate adversarial examples but no post-training skill to use them, and no discipline for generating non-adversarial coverage of the situation space. A machine learning team that can post-train but no training data to feed it. A security architect who knows what the intervention should do but no path from concept to hardened model.
This is where the whole approach has been stuck for most of the field. The failure mode is understood. The intervention is understood. The gap between the intervention and any specific security team’s ability to execute it is what has kept the intervention from being deployed at scale.
Two paths from vlno
For teams without the data and the skill, vlno provides a LoRA adapter. A LoRA adapter is a small overlay on the model’s parameters. It sits alongside the base model without changing the base model’s weights. It carries the specific behavior updates that alignment training produced. When the model runs with the adapter loaded, it responds to situations outside its training the way alignment trained it to. When the adapter is not loaded, the base model behaves exactly as it did before. The security team installs the adapter alongside the base model in the inference stack. The alignment is now in production.
For teams that do have post-training capability, vlno provides the training data itself. Crafted attacks generated at scale by vlno red. Coverage of unfamiliar situations where a model’s learned goals come apart without an adversary. Both cases labeled and structured for post-training. The team runs their own post-training procedure against the base model, applies the alignment they consider appropriate, and holds the resulting weights themselves.
Both paths deliver the same underlying alignment. The choice depends on whether the team has the post-training skill to run it themselves.
Where this leaves the argument
The intervention exists. Alignment via post-training. It reaches the layer runtime and evaluation cannot reach. It preserves the base model. It closes families of failure at their source.
The remaining question is where the training data behind the alignment comes from. The alignment is only as strong as the training data it was built from. If the data is thin, the alignment holds against the tested cases and fails against everything else. If the data covers both crafted attacks and the space of unfamiliar situations the model will meet in production, the alignment closes whole families of failure at once.
That is the subject of the next post.
About vlno
vlno is one platform that shows you every AI agent in your organization and what it is allowed to do, tests how it holds up under both crafted attacks (vlno red) and unfamiliar situations its training did not cover, and produces training data that hardens the model itself. Observation and control, evaluation, and in-weights hardening in a single product.
Learn more at vlno.ai →