Custom LLM Evaluation Harness
A financial-technology company engaged Mercury to build an evaluation and monitoring layer for a custom LLM support system that had passed its initial validation.
The work established overnight evidence collection, regression detection, and review surfaces so the team could know each morning whether the system still behaved as it had when it was validated.
The Objective
The company came out of the build phase of a custom LLM support system with something most teams would call a success. The system answered well in validation, escalated when it was supposed to, and cleared its launch checks. What the company asked Mercury for next was harder to point at. They wanted to be able to say, on any given morning, that the system still behaved the way it had when it was validated, and to say it with evidence rather than confidence.
Success had a concrete shape. Overnight, checks would run against a known set of questions. By morning there would be a record of what changed, which categories held steady, where behavior regressed, and which items needed a human decision. Nobody would have to remember to run anything, and nobody would be asked to vouch for the system on goodwill.
Challenge
Custom LLM systems fail without errors. Behavior changes whenever anything underneath the answers changes: the model provider ships an update, a source document is revised, retrieval re-ranks as content grows, a prompt gets tuned, or an escalation policy is edited. The system keeps answering fluently, and fluency survives degradation longest. A team watching only for outages sees nothing.
A single aggregate quality score can hide exactly the failures that cost the most. If most answers improve slightly while a handful of safety cases begin to fail, the average moves up while the risk moves up with it. This company's system carried escalation duties, and those are precisely the categories that fail quietly inside a healthy-looking total.
A scheduled check is only useful if its failures go somewhere. Without persistent records and a review path, overnight testing becomes ritual instead of instrument. And an evaluation can slowly stop testing the real system as the prompts it exercises diverge from the ones in production and the questions in the test set diverge from the ones users actually ask.
Solution
Mercury started with measurement. Instead of relying on a single quality score, answers would be graded by a scoring council across five dimensions: accuracy, completeness, safety, groundedness, and tone. Scores would also be tracked by category, so a safety failure could not disappear inside a good average. The council preserves its grading rationale, which means a reviewer can later see why a score moved and disagree with it.
Mercury built a versioned corpus of 143 items: 70 golden cases pinning known-correct answers, 58 safety cases, and 15 conversational flow cases. The safety set was shaped around escalation: sensitive topics, categories the system must never handle, and both false positives and false negatives. Golden cases catch drift in what the system knows. Safety cases catch drift in what it is allowed to do.
The corpus needed a machine to run it. Mercury built a six-phase evaluation harness: retrieval checks, safety evaluation, council scoring of final answers, flow smoke tests, regression detection, and report generation. Each phase had independent error handling, so a fault in one phase could not silently erase the rest of the evidence.
Each run leaves a trail. The harness writes markdown reports and persists run and item records. Baselines are maintained as exponential moving averages with per-category metrics. When a category crosses its regression threshold, the configured design raises an alert and routes the run into review. The intended morning sequence is short: an alert names the category, the run detail shows which items moved, the item records show the answers side by side, and a human decides what happens next.
Mercury added two parity gates, one for prompt parity and one for frontend-flow parity, to watch for divergence between what the evaluation exercises and what the live answer path actually does. That kind of mismatch can make an otherwise green run misleading.
Mercury built an administrative dashboard exposing runs and run detail, a candidate queue, a review queue, launch controls, model comparison, grading rationale, and a cost-and-scope preflight, behind audit and role-based access boundaries. The dashboard exists because the loop ends at human judgment.
The last decision governed how the corpus itself would grow. Candidate questions pass through a human review and export lifecycle before they can enter the corpus. The runtime cannot write to the authoritative corpus at all. A second corpus version followed through that path: 166 items, with the safety set extended to 63 and 18 long-tail cases added. Nightly coverage was made deterministic and route-based, with an explicit weekly deep-run contract.
Results
The engagement produced a built and QA-closed evaluation and monitoring system: the six-phase harness, both corpus versions, council scoring with preserved rationale, EMA baselines, category thresholds, configured alerting, regression detection, the parity gates, the admin dashboard, and the human review lifecycle for corpus change.
The schedules are configured: nightly runs, weekly deep runs, and candidate-mining jobs. Configuration is confirmed by the engineering record. A claim that those runs have executed in production over time is a different claim, requires separate approval, and is deliberately absent here. No production accuracy, safety, cost, or reliability figure appears in this case.
What changed for the company is the shape of the question they can ask. Whether the system still behaves as validated now has a designed answer path with evidence at every step.
Summary
The engagement turned a question that never closes into an operation designed to answer it on a schedule. The pattern is a fixed evaluation contract, a versioned corpus, scheduled runs, regression thresholds, persistent evidence, a review queue, and governed change. Nothing in the loop remediates anything on its own. Its product is evidence, arranged so that human judgment lands in the right place with the right records in hand.
For teams operating a custom LLM past its initial validation, Mercury offers an evaluation and monitoring review: a structured look at corpus, schedules, regression policy, and review surface.




