Governed improvement loop for a custom AI deployment
A financial-technology company engaged Mercury to build a governed improvement path for a custom AI support system after its initial build. The work turned post-build feedback into classified, testable evidence and moved it through knowledge, answer-policy, and escalation lanes with review at each step.
The Objective
Mercury had completed the first build of a custom AI support system for the company: a knowledge base built from the company's own support material, an assistant that answered from that knowledge, and an evaluation corpus that tested expected behavior. The next question was practical: once real use exposed gaps, how would the team improve the system without losing control of it?
Success meant more than making a weak answer stronger. An observed failure had to become inspectable evidence pointing to the right kind of change—source material, retrieval behavior, answer policy, or escalation—then be reviewed, verified, approved, and monitored so the people responsible could explain what changed and why.
Mercury's objective was a post-build improvement loop that proved feedback could move from observation to controlled change without being hidden inside ad hoc prompt edits or one-off fixes.
Challenge
The hardest feedback did not arrive in clean categories. It arrived as support conversations that drifted, questions with no source behind them, citations that pointed near the right answer, policy edge cases no one had written down, and human handoffs that happened awkwardly or too late. Each item was useful, but each looked small enough to patch in isolation. That is where drift begins.
The clearest example was a short question. Real users often ask in fragments. The golden corpus had tested well-formed questions well but had not modeled short customer-style questions with the same depth, so the system could pass expected tests while struggling with input support teams see constantly.
Those short questions exposed several failure shapes. In some cases the system abstained even though a usable answer existed. In others, retrieval landed on an adjacent source instead of the governing one. In one class, the answer could reach beyond the context it cited, creating a faithfulness gap. A single motion to fix the answer would have blurred those differences.
Support leads and operators answer for system behavior after the build team has moved on. They cannot govern a change they cannot see, and they cannot defend an answer policy that lives only in someone's memory of a prompt edit.
Solution
Mercury began with the operating habit already present in the first build: let real demand shape the system. The knowledge base had been developed in waves from support-ticket and chat-history analysis. The post-build work extended that habit from content coverage into system behavior.
The first step was to make the short-question problem durable. Mercury built a real-query test set that captured short customer-style questions and recorded how the system handled them. That record included abstentions, adjacent-source retrieval, and the faithfulness gap where an answer could exceed the cited context. Once written down, the short question became a repeatable case the team could retest after a change.
That evidence made classification possible. Mercury traced each documented failure to the kind of change it required. Missing or weak source material belonged in the knowledge base. A source-backed answer that came out in the wrong shape belonged in answer policy or scripted flow. A conversation that should have left the assistant belonged in escalation and handoff design. Each lane carried a different risk and needed a different review path.
The knowledge lane addressed gaps in real conversations. Transcript-audit findings became validated knowledge-base additions; when a short question exposed a missing source, the team could add supportable knowledge or preserve the need to abstain rather than stretching answers beyond their sources.
The answer-behavior lane handled a different problem. Some audit findings showed that the source existed, but the assistant still needed clearer rules for how to respond. Chat-audit findings became governed answer-policy and scripted-flow changes, each with a review lifecycle. That moved behavior out of hidden prompt adjustment into a form an operator could inspect and approve.
The escalation lane handled the point where the system should stop. Mercury delivered human off-ramps, identity capture, escalation rules, and a handoff journey validated through user-acceptance testing. A person can tolerate a system that admits its limit but is far less forgiving of a handoff that disappears.
Verification tied the lanes together. The evaluation system supports candidate questions from production-shaped use and QA flags, review and export, long-tail suite routing, and persistent review queues. A short question that exposed a gap could become a candidate test, but it did not silently rewrite the authoritative corpus. Promotion remained governed.
The loop closed when later retrieval work improved short-query behavior and the remaining residuals stayed visible—a truthful record rather than a clean story. This behavior improved, these cases remain open, and the next cycle begins from evidence still on the books.
Results
The engagement produced a documented improvement path. The real-query test record preserved the short-question class that the original corpus had under-modeled. Transcript-audit findings became validated knowledge-base additions. Chat-audit findings became reviewed answer-policy and scripted-flow changes. Human off-ramps, identity capture, escalation rules, and the handoff journey were delivered and validated through user-acceptance testing.
Feedback no longer has to remain an anecdote or become a quiet prompt patch. It can be captured, classified, routed to the right lane, verified against evaluation coverage, approved through a lifecycle, and monitored in the next cycle. Later retrieval work improved short-query behavior, and the residuals that survived were preserved rather than hidden.
This case demonstrates post-build and production-like testing of an improvement loop. Customer-facing rollout status requires separate validation and is not claimed here. No deflection rate, satisfaction score, response-time improvement, accuracy figure, staffing effect, or revenue result is used.
Summary
The short question is useful because it looked too small to build a program around. It could have become a quick edit, a note in a backlog, or a recurring complaint nobody owned. Instead it became a testable evidence class, then a classification, then knowledge work, answer-policy work, escalation work, evaluation coverage, verification, approval, and residual monitoring. That is the difference between fixing an answer and operating a support system.




