Two Clocks, One Library: Evaluating Agents Before and After You Publish
A launch scorecard with no successor is a photograph of a moving object. Here is the release clock and the surveillance clock, joined by one versioned case library.
Pre-publish evals tell you a change is safe to ship. Post-publish evals tell you the thing you shipped still works. Here is why you need both, on one shared case library.
September 4, 202616 min read
Agent EvaluationLLM-as-JudgeMLOpsEU AI ActDrift Detection
Pre-publish evals tell you whether a change is safe to ship. Post-publish evals tell you whether the thing you shipped is still working. Doing one does not give you the other.
Agent Evaluation · MLOps · Continuous Monitoring
Key Facts
Public agent benchmarks have shifted from scoring a single answer to scoring a trajectory — SWE-bench, WebArena, GAIA, and tau-bench all measure whether an agent took sensible steps, used the right tools, and reached a correct end state.
tau-bench introduced the metric that matters most in production: reliability across repeated attempts (pass-at-k), not best-of-one. The gap between the two is where most enterprise disappointment lives.
Zheng et al.'s work on MT-Bench and Chatbot Arena showed strong judge models agree with human preference at rates comparable to human-to-human agreement — and documented the failure modes that matter: position bias, verbosity bias, self-preference, and drift when the judge model updates.
The EU AI Act does not stop at pre-market conformity. Article 15 requires accuracy, robustness and cybersecurity across the lifecycle; Article 72 requires providers of high-risk systems to run a post-market monitoring system; Article 73 adds serious-incident reporting duties.
1 to 5% of live traffic, judged continuously and sampled by risk rather than uniformly, is enough to see drift before customers do.
TL;DR
Ask a team when their agent was last evaluated — if the answer is launch week, they have a photograph of a moving object, not an evaluation programme.
Two clocks are needed: a release clock (event-driven, blocking) and a surveillance clock (continuous, sampled). Most teams run only the first.
The asset joining them is one versioned case library. A confirmed production failure is promoted into a permanent case, so the release clock can never let the same defect ship twice.
A defect caught at the pre-publish gate costs roughly 1x to fix. The same defect found by a customer or regulator costs 30 to 100x.
Continuous evaluation has moved from good practice to regulatory obligation — Articles 15 and 72 of the EU AI Act require exactly the surveillance clock described here.
Keep reading → the seven signals, the eight-week build sequence, and where AgentTrust OS fits.
Ask a team when their agent was last evaluated. If the answer is the week it launched, they do not have an evaluation programme, they have a photograph of a moving object. Agents rarely fail on release day. They fail three months later, quietly, when a model was updated or an index was refreshed and nobody was measuring.
A team builds an agent. Before launch they do the right thing — assemble a few hundred test cases, run them, tune until the numbers look good, present a scorecard to a governance forum. The agent goes live. The scorecard is filed. Then the evaluation stops, not by decision, by omission. There is no owner, no schedule, no trigger.
Definition — The Evaluation Cliff
The moment an agent leaves the only environment where anyone was measuring whether it was any good. Three months later, a model point update, a re-indexed knowledge base, a downstream API field rename, or a new class of question breaks quality — and none of it raises an alert. Latency is fine, error rates are fine, the dashboard is green. The only instrument pointed at it is customer complaints, which arrive weeks late, heavily filtered, and attached to a reputational cost.
Two clocks, one library
Pre-publish and post-publish evaluation are not two projects and not two toolchains. They are two consumers of a single versioned case library, running on different clocks with different tolerances.
The release clock
Event-driven. Fires on every change to prompt, model, retrieval config, tool schema, or policy. Blocking, exhaustive against the library, deliberately biased toward false positives — it should occasionally stop a safe release rather than ever pass an unsafe one.
The surveillance clock
Continuous. Samples live traffic, scores it, watches distributions. Biased toward cheapness, because it runs forever. Not blocking — it triggers investigation, rollback, or a new release-clock run.
The line I keep coming back to
An eval suite that does not grow from production incidents is a fossil. The best predictor of whether an organization's agents will still be trustworthy in a year is not the size of their launch eval set — it is whether last month's production failures are in this month's regression suite.
Figure 1: Two clocks, joined by one artefact. Every production failure the surveillance clock confirms is promoted into the case library, so the release clock physically cannot let that defect ship again.
Building the case library, from four sources
Everything else depends on this asset, so it gets built first and owned explicitly. A production-grade library draws from four sources — teams that use only the first are the ones who hit the cliff.
01
Golden cases
100 to 300, written by domain experts.
02
Sampled production traffic
Replayed, stratified by intent.
03
Adversarial cases
Mapped to OWASP LLM and MITRE ATLAS.
04
Incident-derived cases
The source most teams lack — and the one that compounds.
Version it like source code. A case whose expected output is stale because policy changed is worse than no case at all.
What a blocking release gate actually contains
The gate runs on every change to any component that can alter behavior — system prompt, tool definitions, retrieval configuration and index contents, model identifier or version, temperature and sampling parameters, guardrail thresholds, and any policy the agent must follow.
Dimension
What it measures
Typical gate condition
Task success
Did the agent reach the end state, scored programmatically where possible
No regression vs. baseline beyond the noise band
Reliability
Same task, repeated k times, all must pass
Report pass-at-k, gate on the strictest tier
Safety
Adversarial suite, refusal correctness, data leakage
Zero critical findings, hard block
Groundedness
Claims supported by retrieved sources, citations resolvable
Hard floor for any customer-facing agent
Cost & latency
Tokens and wall time per successful task
Budget ceiling — a quality win that triples cost is a decision, not a default
Two rules make the gate trustworthy. Always run a baseline in the same job — the current production configuration, same cases, same conditions — because absolute scores drift with judge versions and infrastructure; the delta is what you can defend. And establish a noise band before setting thresholds — run the identical configuration five times and measure the spread. A gate tighter than your own measurement noise will block real releases at random and get switched off within a month.
Shadow and canary, the stage that catches what offline evals cannot
Shadow mode runs the new configuration against real production traffic without exposing results to users — the only cheap way to discover the input distribution has moved away from your case library. Canary exposes a small slice of real users with automatic rollback wired to guardrail triggers, override rate, abstention rate, and cost per task.
The surveillance clock: four instruments, running forever
Post-publish evaluation is a sampling problem — you cannot judge every run affordably, and you do not need to.
Programmatic checks on every run. Schema validity, citation resolvability, forbidden phrases, tool-call legality, numeric range checks — nearly free, and they catch hard failures immediately.
Judge scoring on a stratified sample. One to five percent, weighted toward high-risk intents and low-confidence runs. Track the distribution, not the mean — the mean hides a growing tail.
Behavioral telemetry as a proxy. Abstention rate, escalation rate, override rate, retry rate, conversation length — cheap, covers all traffic, and moves before quality scores do.
Drift detection on inputs. A new intent cluster appearing is a signal to extend the library before quality drops — the only genuinely proactive instrument in the set.
Governing the judge
If a model judge gates your releases, that judge is production infrastructure. Pin its version, version the rubric alongside the code, keep a human-labelled calibration set of 50-100 cases, and re-run calibration whenever the judge model or rubric changes. Use pairwise comparison against the baseline wherever you can — judges are markedly more reliable at "which of these two is better" than at absolute scoring.
Evaluating any agent type: change the unit, keep the method
"Our agent is different, it cannot be evaluated that way" is the most common objection. In practice the method stays constant and only the unit of measurement changes.
Agent type
Dominant failure
Score this unit
Hard gate
Conversational
Confident wrong answer
Final turn & policy adherence
Zero unsupported claims
Retrieval & knowledge
Plausible answer, wrong source
Groundedness & citation resolution
Every claim cited
Workflow & tool-using
Right answer, wrong side effect
Trajectory & end state in the system
Zero illegal tool calls
Code-writing
Passes tests, breaks intent
Test pass, diff review, build
No reduction in coverage
Multi-agent
Error amplified down the chain
Per-hop scoring, not just end hop
Level attribution exists
Long-horizon autonomous
Silent divergence from goal
Checkpoint states, cost per outcome
Budget & time kill-switch
Standing it up in eight weeks
Week 1
Define good. A written definition of success per agent, agreed with the business owner, risk tier, and failure severities.
How you know it worked: the business owner signs the definition.
Weeks 1–3
Case library v1. 100-300 golden cases, 50 adversarial, 100 replayed production runs, versioned in the repo.
How you know it worked: any engineer can run the full set locally.
Weeks 3–4
Scorers. Programmatic checks, rubric judge with a calibration set, trajectory scoring where relevant.
How you know it worked: judge-to-human agreement above 80%.
Week 4
Baseline and noise band. Five identical runs to measure spread, thresholds set outside the noise band.
How you know it worked: repeat runs do not flip the gate.
Week 5
Release gate in CI. Blocking job on every change to prompt, model, index, tools, or policy.
How you know it worked: a deliberately regressed prompt is blocked.
Week 6
Shadow and canary. Shadow runner against live traffic, canary with automatic rollback on guardrail and cost triggers.
How you know it worked: rollback fires in a drill.
Week 7
Surveillance. Programmatic checks on all runs, sampled judge, override and drift telemetry, alert routing.
How you know it worked: an injected quality drop raises an alert.
Week 8
The feedback edge. Incident-to-case pipeline with an owner and a service level — the step that makes the system compound.
How you know it worked: last month's incident is in this month's suite.
What it is worth
The whole economic argument rests on one measurable asymmetry: where a defect is caught determines what it costs.
1×
cost to resolve a defect caught at the pre-publish eval gate — engineering hours, before anyone sees it
10–20×
cost of the same defect found by production surveillance — incident handling, some customer impact
30–100×
cost if a customer or regulator finds it first — remediation, disclosure, review of every affected case
AgentTrust OS: Running Both Clocks on One Runtime
The release clock and the surveillance clock in Figure 1 are not two toolchains to build and maintain separately. AgentTrust OS runs both against the same runtime, so the case library, the judge governance, and the incident feedback loop are one system rather than three.
Trust Certify
Is the release clock. The Reliability Engine runs the case library thousands of times to produce a consistency score, so the blocking release gate becomes a certification threshold rather than a manual sign-off someone has to remember to enforce.
Trust Runtime
Keeps the surveillance clock honest between releases. Every production decision gets a confidence score at the moment it is made, so drift shows up as a change in the score distribution rather than a customer complaint three weeks later.
Trust Audit
Is the case library itself, and the record of what the judge decided and why. Both clocks write to the same store, so a governed judge and a defensible history are properties of the runtime, not a separate pipeline someone maintains by hand.
Frequently Asked Questions
It answers whether a change is safe to ship, which is necessary but not sufficient. It says nothing about whether the thing you shipped is still working three months later, after a silent model point update, a re-indexed knowledge base, or a shift in what users are asking. Those failures do not show up in error rates or latency dashboards — the agent keeps answering, just worse. You need a second, continuous instrument pointed at production, sampling live traffic rather than a static test set.
Less than the alternative. Sampling one to five percent of traffic, weighted toward high-risk intents and low-confidence runs rather than sampled uniformly, is a small and bounded cost. Compare that to the 30-to-100x cost multiple of a defect a customer or regulator finds instead of your surveillance clock. The rough sizing rule is 15-20% of build cost for the eval system in year one, and 5-8% of run cost after — below that range, teams build a launch scorecard and stop; above it, the programme starts optimising the eval rather than the agent.
Yes — the method stays constant and only the unit of measurement changes. A workflow agent's dominant failure is the right answer with the wrong side effect, so you score trajectory and end state in the system of record rather than final-turn quality, and the hard gate is zero illegal tool calls rather than zero unsupported claims. The versioned case library, the baseline run in the same job, pass-at-k rather than pass-at-1, and the incident-to-case feedback loop are identical across agent types.
That is exactly the failure this framework is built to catch, because the judge is production infrastructure and needs its own evaluation. Pin the judge's version, version the rubric alongside the code, keep a human-labelled calibration set of 50-100 cases, and re-run calibration whenever the judge model or rubric changes. Report judge-to-human agreement as a governance metric — when it falls below roughly 80%, the rubric is ambiguous, not the model weak.
No, and it would be dishonest to claim otherwise. Trust Certify's pre-production gate and Trust Runtime's continuous scoring both need a case library that reflects your actual failure modes, your policy, and your production incidents — nobody outside your organisation can write those cases for you. What the platform removes is the plumbing: the release gate wiring, the judge governance, the drift telemetry, and the incident-to-case pipeline that most teams spend a quarter building before they can focus on the cases themselves.