Notes 5 min read

Agree the standard before you build

The deployments that last agreed quality, cost, time and human effort on real examples before anything was built. The pilots that die never agreed what good meant.

Laurens Nys Founder, Ortelian

View Markdown

Most agent pilots I see start with a demo. Someone shows the agent doing something impressive, the room nods, and a build starts. Three months later the pilot is quietly shut down, and nobody can say whether the agent was doing the job, because nobody ever wrote down what the job was. That is the usual ending, not the unlucky one: MIT’s 2025 survey found 95% of organisations see no measurable P&L impact from their generative AI pilots. The deployments that last are the ones where quality, cost, time and human effort were agreed on real examples before anything was built. That is the only way anyone can later say whether the agent is doing the work. The pilots that die never agreed what “good” meant.

The standard is four questions, asked of real work

We ask four things of every piece of work an agent will take over. Quality: does the work meet the standard on real examples? Cost: what does each completed piece of work cost? Time: how long before the result is ready? Human effort: what review, correction or intervention is still needed?

GENERATIVE AI PILOTS, 2025 95% of organisations see no measurable P&L impact from their generative AI pilots 1 in 20 pilots extracts value
Generative AI pilots, 2025
95%
of organisations see no measurable P&L impact from their generative AI pilots
1 in 20 pilots extracts value
Most pilots never reach a measurable result. Source: MIT NANDA, "The GenAI Divide: State of AI in Business 2025", July 2025.

Each one has to be asked of a specific piece of work, not of the agent in general. “Is the meeting prep good?” cannot be answered. “Is this brief, for this meeting, what the rep needed on screen before the call?” can. The same goes for cost. A pilot that costs nothing to run because it does nothing useful is not cheap. What matters is the cost of one completed brief, one ranked list, one enriched contact. At a deep-tech company, the one place the platform came up short was phone numbers. It didn’t pull them natively, and every phone endpoint we could add was expensive. That was a cost question about one piece of work, and we could only answer it because we knew what one completed contact was supposed to contain.

Time and human effort are the two that get skipped. Time means when the rep can use the result, not how fast the model answers. The top-50 list at that company refreshes at 7am so the rep has it by 9am. If it landed at 11am it would fail the standard even if every account on it were right. Human effort is what is left for a person to do after the agent has done its part. If a rep has to check every account against the CRM before calling, the agent has moved the work, not done it.

Real examples become the evaluations

The standard lives in real cases with agreed answers, not in a document. For the meeting-prep agent, the standard is what a good rep would want on screen before a call: previous calls, emails, deal history, similar companies. The examples are real past meetings. We take meetings that already happened, run the agent against them and ask whether a good rep would have wanted that brief. Once we have a set of those, every change to the agent is checked against them.

For the ranked account list, the standard was simpler and harder. The rep had a spreadsheet of about 2,700 partner companies. What he wanted was a short list he trusted. The standard was him saying “this is exactly what I wanted.” We got there at 64 ranked accounts, 48 of them with contacts already pulled. That sentence from him is the evaluation. Every later version of the list gets measured against the one he said it about.

Some cases go into the set because they need a human. A deal at a late stage, a contact who has just changed company, an account that sits in two territories. For those the expected outcome is “stop and ask”, not “act”. An agent that acts confidently on one of them is failing, and the evaluation has to say so.

The standard has to exist before the build, because it decides what the build is

Agreeing the standard first looks like process, but it is what sets the scope. What the agent may access, what it may do and which decisions need approval all follow from what “good” means for the work. If the standard for the account list includes writing contacts to the CRM, the agent needs write access and a rule about when it may write. If the standard is a list in Slack for a rep to act on, it doesn’t. You cannot set those limits after the build without rebuilding.

BEFORE THE BUILD BUILD AND OPERATE Real examples Agreed standard Build Evaluate Deploy Past work with an agreed answer Quality Cost Time Human effort Scope follows the standard Against the same examples In use by the team Falls short investigate change knowledge or workflow rerun evaluations
Before the build
Real examples
Past work with an agreed answer
Agreed standard
Quality · Cost · Time · Human effort
Build and operate
Build
Scope follows the standard
Evaluate
Against the same examples
Deploy
In use by the team
Falls short → investigate → change knowledge or workflow → return to Evaluate.
The standard is set on real examples before anything is built. Every later fix is checked against the same examples, so a change is known to be better, not just different.

The second reason is what happens after launch. The work falls short. It always does, somewhere. A signal gets weighted wrong, a rule in the ontology turns out to be too strict, a rep gives feedback in the channel that changes what the brief should contain. When that happens we investigate, change the knowledge or the workflow and rerun the evaluations. That last step is what makes the fix safe. Without a baseline, every fix is a guess about whether things got better or just different. The usage review that follows (one week at one company) only works because there is a number from day one to compare against.

The audit is where the standard gets agreed

We agree this on site, in the two-week audit. Week one is observe and map: trace the inputs, decisions, handoffs and exceptions in the work, and inspect the systems and information people actually rely on. That is where the real examples come from. You cannot pick good past meetings from a distance. You have to sit with the rep and ask which briefs he would have wanted.

Week two is redesign and rank: decide what belongs with agents, software or people, rank the opportunities and check what the first build requires. By the end, the customer leaves with a map of the work, ranked opportunities, the first scope and a baseline. The baseline is the standard, written against real cases, with the current numbers for quality, cost, time and human effort filled in. The build starts from there, not from a demo.

”We’ll know it when we see it” is how pilots die

The obvious objection is that the standard will emerge. Build something, put it in front of people, and they will tell you if it is good. In my experience they won’t, and neither will the agent. People are good at saying a specific output is wrong. They are bad at saying what the output should have been, and worse at saying it consistently across fifty outputs. Without agreed examples, the feedback is a stream of individual reactions, and the agent gets tuned to the loudest one. The rep at the deep-tech company found the meeting-prep channel on his own and now checks it before every external meeting. That happened because the first briefs were already close to what he wanted, and they were close because we had agreed with him, on real past meetings, what close meant.

Agree what good looks like on real work before you build, or you will never be able to say whether the agent did the job.

Follow the work.

Roughly monthly. Notes only. Unsubscribe anytime.