Image

Your Agent Observability Stack Should End in a Regression Test

October 7, 2026
By
Joshua Goldfein

A trace tells you what the agent did. Reliability depends on what happens next, and the object that step has to produce is a regression case that remembers the failure after the engineer who fixed it has moved on.

In short

The trace is open on the screen: forty spans, the retrieval step, the tool call with the wrong argument, the confident final answer. The engineer finds the problem in twenty minutes, patches the tool description, reruns two examples, and closes the ticket. Three weeks later the model gets upgraded and the same failure comes back with a different span ID.

Observability did its job in that story. The run was no longer a black box, and nobody argued from screenshots. What it did not do is turn the evidence into a change that was approved, tested against the original failure, and remembered. That gap, between agent observability and agent operations, has a one-line diagnostic: when an agent fails, what object gets created? If the answer is a ticket, the lesson leaves with the engineer.

Evidence is the first step, and only the first

Teams instrument agents for the right reason. Without traces, incident review becomes folklore. Nobody can separate a retrieval failure from a reasoning failure, a tool-contract problem from a stale source, an instruction-hierarchy problem from a missing escalation rule. A good trace changes the review surface: which context the agent saw, which tool it called, which constraint it missed, where the run diverged.

Then four things still have to happen outside the dashboard. Someone classifies the failure. Someone decides whether the fix belongs in a prompt, a tool schema, a retrieval policy, an eval, a source hierarchy, or an escalation rule. Someone tests the change against the original failing input and against the cases that already worked. Someone decides whether it is safe to promote. When those steps live in people’s heads and chat threads, the organization gets visibility without institutional memory, and the same pattern returns under a new name after the next model, tool, or prompt change.

The repair loop, in four steps

TRACE TO REGRESSION CASE
01 Observe
02 Diagnose
03 Rerun
04 Lock

Observe is the trace itself: inputs, sources, model calls, tool calls, outputs, cost, and timing, captured under the same data-handling rules as the production run.

Diagnose turns the trace into a claim with evidence attached: the failure category, the causal chain through the spans, and a proposed change. Automated diagnosis earns its place here, since a model reads forty spans faster than a person does, and what it produces stays a proposal. A diagnosis that cannot point at the span it is explaining is a story, and stories do not get promoted.

Rerun is where authority enters. An accountable owner approves the proposed change for a sandbox replay against the original failing input and against the cases that already passed, with the sandbox conditions recorded. Old trace and new trace sit side by side. Nothing has touched production yet, and nothing should until the comparison shows the failure gone and nothing else broken.

Lock is the step most teams skip, and the one this post is named for. The failing input, its trace, the failure category, the approved fix, the expected behavior, and an owner become a regression case. The next model upgrade, prompt revision, retrieval change, or tool update passes through that case before anyone accepts the new behavior. A failure that produces only a fixed prompt is a temporary improvement. A failure that produces a regression case changes what every future change is inspected against.

“Self-correcting harness” needs care as a phrase, because it suggests a system that patches itself end to end. The version that holds up in production is more autonomous at evidence gathering and controlled experiment, and governed at promotion, the boundary where a change reaches customers, operations, or the compliance file.

Provenance keeps the library honest

A regression library stays useful only while each case keeps its relationship to the failure that created it: the input, the trace, the category, the approved fix, the expected behavior, the owner. Strip that away and the suite becomes a pile of assertions whose importance nobody remembers, which is the moment someone deletes the one that mattered to make the build go green.

Early cases are informal: representative inputs, expected outcomes, forbidden behaviors, source-grounding requirements, escalation triggers. Later they formalize into trace-based suites, judge checks with calibrated human review, policy assertions, and adversarial edge cases. Test maintenance becomes a real operational cost at that point, which is why the owner field exists from the first case.

Where does your loop stop?

WHERE THE INCIDENT STOPS TRAVELLING
Stage What the team can do Where it stops
Trace-only observability Inspect a run after it fails Diagnosis and repair stay manual, and nothing is remembered
Eval-aware operations Compare behavior against known quality checks Production failures never become new checks
Sandbox reruns Replay failing inputs safely before deployment Rerun conditions vary, so comparisons are hard to trust
Approval-gated fixes Route changes through accountable owners Reviews happen without the trace and the eval in front of the approver
Regression-compounding harness Convert every failure into a durable, owned case Test maintenance becomes a standing responsibility

The jump between stages is procedural before it is technical. A team at the second stage usually already owns everything the fifth needs: traces, evals, a staging environment, pull requests, an incident review. What is missing is the loop that connects them, and a name next to each step.

When an agent fails, what object gets created?

The design test

Where Mercury fits

Mercury reviews the repair loop on one production agent workflow. We take one real failure of yours and walk it through the four steps, and you keep what that produces: the diagnosis packet with its spans cited, the approval record, the sandbox comparison, the regression case with its provenance, and a list of the steps in your loop that nobody currently owns. If your agent’s last incident produced a ticket and nothing else, review the repair loop on that workflow with us.