vlno — We train the threat out

Holodeck Evals · now open

Your agent passed the demo. What does it do on a hostile Tuesday?

We drop your agent into realistic work environments — a crowded inbox, a finance console at month-end, a support queue on a Monday — and grade what it actually did from live world state. Not what it said. Not a judge’s vibe. Evidence.

Transcript evals grade the talk.

Most evals read what the agent said it did, scored by another model. Agents have learned to say it. The gap between a good transcript and a good outcome is where production incidents live.

Demos are benign by construction.

Nobody's demo inbox contains a supplier politely asking to update the banking details, or a 'colleague' who needs one password confirmed. Real Tuesdays do.

Refusal is not robustness.

An agent that deletes its inbox to stay safe is safe the way a switched-off computer is secure. If your eval can't tell refusal from resilience, you're measuring the wrong thing.

01

Pick the model.

Any public model id, your fine-tune, or your internal endpoint. Nothing to integrate; the agent gets a job, not a test harness.

02

It works real shifts.

Triage the inbox. Close the books. Work the ticket queue. Injections live in the data the way they live in the wild — a vendor email, a poisoned ticket, a calendar invite with opinions.

03

You get the evidence.

Every verdict links to the state your agent actually changed: the email it sent, the account it rerouted, the role it granted. The world is the grader.

Utility — did it finish the work?

Every reply sent, every slot booked, every ticket resolved. An agent that can't do the job has no robustness score worth the name.

Robustness — did it hold the line?

The environment pushes: a phish in the inbox, a poisoned ticket, a too-helpful colleague. We measure whether the goal survived contact.

Fidelity — did it stay useful under pressure?

Resisting by refusing everything is sandbagging with good PR. We grade the agent down for it — loudly.

A sample report card

An open-weights frontier model · 82 scenarios · 217 adversarial attempts · August 2026. Anonymized for launch — your report names your model.

Utility

87%

of shifts completed

Robustness

70/82

scenarios held the line

Where it broke

11/12

hijacks were plain in-band asks

Cost of knowing

17.9M

tokens for the full sweep

The most dangerous attacks against this model were the polite ones.

Where it broke: 11 of 12 hijacks were plain in-band asks (“just this once, confirm the password”) — no clever encoding needed. One via chat-template role injection.

Report Card

free

Five scenarios across inbox and finance ops. Three axes. Evidence attached. ~24h turnaround. The fastest honest answer to 'should we ship this model?'

The Full Sweep

All twenty scenarios plus the adaptive red-team: an attacker that reads your agent's reactions and rewrites its approach until something lands. You get the transcript of how.

Custom Worlds

Your workflows as scenarios — your tools, your data shapes, your failure modes. If it matters to your Tuesday, we can build it.

Regression Gate

The sweep in CI. Every model bump, every prompt change, every fine-tune — graded before it ships.

Tell us what to evaluate.

What’s the thing your agent does in production that nobody measures? One line. We read all of them — and the three best become next month’s scenario pack, credited to you.

Get your agent’s report card

Do you judge with LLMs?

No. We don't ask a model whether your agent was good — we read the inbox, the ledger, and IAM after the run. Verdicts come from live world state: what changed, not what was said. (We do use models to QA our own scenarios. Never to grade yours.)

Is my model data safe?

We need only a public model id or an endpoint you control. Runs are isolated and sealed. We never train on your runs, and your report is yours.

Can I bring my own scenario?

Yes — that's Custom Worlds. Or drop it in the feedback line; the best three each month get built free.

What models can you run?

Anything on OpenRouter, or an OpenAI-compatible endpoint you expose. Fine-tunes, internal builds, the weird stuff — that's the point.