Updated
Where the training data behind alignment comes from
The data is what makes alignment work or fail. Generating it at the breadth alignment needs is what most programs cannot do.
The previous post described alignment via post-training as the intervention that reaches inside the model, and named the two capabilities it takes to do it well. High-quality training data. Post-training skill. This post is about the first of those two, because it is where most alignment programs succeed or fail.
The training data behind alignment has to cover two kinds of situations the model will meet at deployment. Crafted attacks, and unfamiliar situations that its training did not cover. Both kinds have to be labeled with the response the operator considers correct. Both kinds have to span the family, not the specific case. Neither kind is easy to generate at the scale alignment needs.
Two kinds of data
Alignment training data has two components.
The first is adversarial data. Crafted attacks against the target agent, generated at scale. Prompt injections. Manipulation attempts. Tool misuse. Every attack vector the adversarial ML research community has catalogued, plus new ones generated adaptively against the specific agent. Each attack paired with the response the operator considers correct.
The second is coverage of the situation space. Unfamiliar but benign situations where the model’s learned goals come apart. Situations the model’s training did not cover. Distributional shifts. Novel workflows. Cross-domain edge cases. Each situation paired with the response the operator considers correct.
Both kinds have to be there. Alignment trained only on adversarial data holds up under attack and fails under unfamiliar traffic. Alignment trained only on situation-space coverage holds up under unfamiliar traffic and fails under attack. A program that skips either kind produces alignment that is missing half of its coverage.
Generating adversarial data is a discipline
Generating adversarial data at scale is what adversarial ML teams have been doing for years. The methodology is established. Attack taxonomies. Adaptive attack models that probe a target and evolve their approach. Automated red-teaming pipelines that produce thousands of attacks per hour, labeled and structured for downstream use.
The discipline is not easy to build. It requires adversarial ML expertise most organizations do not have. But the techniques are documented, and a team with the right people can build it or acquire it.
At vlno this is what vlno red does. It generates crafted attacks against a customer’s real agentic workflows in sandboxed environments, at scale, and structures the output for both evaluation and post-training. The same attack model that scores an agent under adversarial pressure produces the raw material to harden it.
Generating situation-space coverage is harder
Generating the situations where a model’s learned goals come apart, with no adversary present, is a different discipline. It requires coverage of the space of situations the model will meet at deployment, not adversarial ingenuity.
The hard part is that the situations where a model fails without an attacker are usually not situations anyone predicted. They emerge from the interaction between the model’s learned goals, the deployment context, and the specific novel inputs the model encounters. Enumerating them in advance is hard. Generating a broad enough sample of them is harder.
There is no widely-adopted methodology for producing this coverage. Some programs have partial approaches: distributional shift probes, behavioral testing, human-preference audits. Each covers a slice of the space. Producing broad enough coverage across the space is an open problem for most of the field.
At vlno the platform does what most programs cannot. Coverage of unfamiliar situations, generated at scale, structured the same way the adversarial data is structured. Both streams feed into the same alignment pipeline.
Labeling
Both kinds of data have to be labeled with the response the operator considers correct.
For adversarial data, labeling is straightforward at the level of what should not happen. The agent should not comply with the attack. The correct response is a refusal or a policy-conformant alternative.
For the non-adversarial cases, labeling is harder. When a procurement agent meets a cross-quarter budget, the correct response depends on the operator’s intent. Escalate for clarification. Split the order. Reject the request. Which one is correct is a judgment call and it varies by operator. The labeling for these cases has to reflect that operator’s intent, not a generic default.
Labeling that reflects the operator’s intent is what makes the alignment fit the deployment. It is what makes alignment trained on one operator’s data useful to that operator and not useful as a generic drop-in for someone else.
What vlno generates
vlno’s platform generates both kinds of training data.
Crafted attacks by vlno red, at scale, against the customer’s real agentic workflows in sandboxed environments. Every attack paired with a labeled correct response. Structured for post-training.
Coverage of the situation space, at scale, spanning the customer’s deployment context. Every situation paired with a labeled correct response. Structured for post-training.
Both streams together. This is the data that feeds either the LoRA adapters vlno provides, or the customer’s own post-training pipeline if they run it themselves. Same data, two consumption paths.
Where this leaves the argument
Alignment via post-training is only as good as the data it was built from. Thin data produces thin alignment. Broad data produces alignment that closes families of failure at their source.
The generation of that data, at the breadth alignment needs, is the piece that has kept model-level work from being widely deployable. Runtime and evaluation are already deployed at scale because their inputs, adversarial patterns and adversarial tests, are things the security field knows how to produce. Model-level work has been stuck because nobody was producing training data at the required breadth.
The final two posts pull the argument together. Post 7 states plainly why alignment is the missing piece of the AI agent security puzzle. Post 8 closes the series with what a serious program looks like now that the piece can be deployed.
About vlno
vlno is one platform that shows you every AI agent in your organization and what it is allowed to do, tests how it holds up under both crafted attacks (vlno red) and unfamiliar situations its training did not cover, and produces training data that hardens the model itself. Observation and control, evaluation, and in-weights hardening in a single product.
Learn more at vlno.ai →