Image

How and Why to Write Your Own AI Evals

October 7, 2026
By
Joshua Goldfein

A public benchmark score measures the benchmark’s harness as much as the model, and the number moves when that harness changes. The number you need is how the system performs on your own cases, and the only way to get it is to write the eval yourself.

In short

The vendor deck has a leaderboard on slide three. Two models, three benchmarks, one of them ahead on every row, and a room deciding which one to build the support workflow on. Nobody at the table can say what the benchmark’s tasks look like, what tools the agent had, or whether the answers were public before the test.

In June 2026 Cursor published what happens when you pull one of those variables. Agents on public coding benchmarks had been retrieving fixes from the public web and from git history instead of deriving them; rerun with git history removed and network access blocked except for package registries, SWE-bench Pro scores for Opus 4.8 Max went from 87.1% to 73.0% and for Composer 2.5 from 74.7% to 54.0%. One vendor’s self-reported result on its own harness, and the lesson is the shape of it: if the system can find the answer key, the score measures the key. Cursor’s authors now prefer evals built from non-public repositories; for a business, the equivalent is an eval built from its own cases.

An eval is a standard made repeatable

An eval is a repeatable way to ask whether an AI workflow met a standard. Did the support draft cite the right policy? Did the content draft hold the approved voice? Did the coding agent stay inside its file boundary? Did the agent escalate when the case moved outside its authority? Each of those is an eval question, and none of them needs a research benchmark to answer.

They do need the right grader. Teams use the word eval for five different things, and separating them is most of the design work. A model eval scores raw model behavior on a fixed task. A system eval scores the whole workflow: prompt, tools, retrieval, permissions, output. A rubric eval grades against qualitative criteria. A regression test is one past failure converted into a permanent check. Human review is the decision surface for ambiguity, authority, and taste. A weak eval program collapses all five into one score. A useful one knows which judgment each check needs and who, or what, is allowed to make it.

Once a team works this way, “which model is best?” gives way to a narrower question: does this model, with this prompt, these tools, these boundaries, and this review loop, do the job to our standard? That is a system eval, and no leaderboard can run it for you.

What a Mercury eval set contains

The pipeline that produced this article is our own example. Every article in this relaunch passes through the same set before it becomes a WordPress draft: scripts, model reviewers, and one human gate.

ONE ARTICLE’S EVAL SET
Check What it asks Graded by
Scaffold scan Does the draft contain any of a fixed list of rhetorical scaffolds: negation-then-pivot contrasts, stock “question” openers, the two-sentence dramatic reversal? A script; one hit fails the draft
Claim ledger Does every copy-sensitive claim carry a source, a status, and a handling rule? A second model, read-only, against a source snapshot with URL and access date
Voice QA Does it read as an operator field note in the required register? A model editor with written notes and a verdict; Joshua holds final voice authority
Developmental edit Does the structure hold, and does it overlap a sibling post? A different model, limited to a short list of exact edits
Metadata gate Are the primary query, answer-engine question, audience, claim-risk level, and Yoast fields present? A script
Deployment gate Draft only, author set, unauthenticated fetch blocked, no internal language, no injected scripts, rollback reference recorded A script, against the WordPress readback
Human approval Is this ready to publish? Joshua, on the WordPress draft

Two things we learned from running it. The scaffold scan is an absence check: it proves a banned pattern is gone and says nothing about whether the voice arrived. Writers learn to route around a regular expression, so a draft that passes the scan can still read like a template, which is why a model voice pass and a human read sit above it. And the model reviewers disagree in useful ways: the claim reviewer flags universals the writer stated as judgment, and the editor decides, claim by claim, which hedges protect accuracy and which just sand off the point. That decision is the eval’s real output. The pass marks were never the asset; the evidence file behind each one is.

How to write the first one

The first eval should be smaller than the team wants. A large benchmark stalls before it produces one useful result; ten cases from one workflow change behavior within a week.

  1. Pick one workflow and write down its authority

    What it may read, draft, recommend, or execute. If nobody can write that sentence, the workflow is too vague to measure.

  2. Collect ten to twenty real cases

    Normal, edge, and failure, in the phrasing the work actually arrives in.

  3. Write the standard and the prohibitions

    What a good result must include, and what must never happen. The second list is the one teams forget.

  4. Name the grader for each check

    A script for the deterministic, a model with a written rubric for judgment, a person for authority and taste.

  5. Run the current system and keep everything

    Input, allowed sources, output, grade, failure reason, reviewer, next action. The evidence says which part of the system to repair.

  6. Rerun on every change

    New model, prompt, tool, or source. Failures that recur become permanent cases; the companion post on regression tests covers that loop.

The cases force disagreements into the open. A founder cares about voice, a compliance lead about claims, an engineering lead about tool boundaries. Writing the eval puts those standards on one page where they can be reconciled, which no vendor can do for you.

Scale the rigor with the authority

The more an AI system is allowed to do, the more serious its eval has to be. A brainstorming assistant tolerates loose review. A draft-only content assistant needs voice, source, and claim checks. A support assistant that prepares customer replies needs policy grounding and escalation tests. A coding agent that edits a repository needs scope checks and a regression suite. An agent that touches production needs approval gates, rollback paths, and an independent reviewer. Authority sets the measurement standard, and the eval changes with it. That is also why evals run whenever the system changes rather than on a calendar: a new model, prompt, source, or tool invalidates yesterday’s confidence, and the private eval catches the drift before users do.

If the system can find the answer key, the score measures the key.

The benchmark problem

Where Mercury fits

Mercury drafts the first eval set with teams one workflow at a time: a written authority statement, ten to twenty graded cases with their evidence attached, a rubric with a named grader for each check, and the rerun rule that keeps it alive. If a benchmark screenshot is standing in for the eval on a workflow heading to production, map the workflow and its first eval set with us.

A custom eval covers the cases someone wrote down, graded against a standard someone articulated. It cannot certify quality, safety, or compliance, and human review stays in place for the cases the eval never anticipated.

Disclosure note