Failure-mode map
Which errors are in scope, and who owns them
Solutions
An evaluation strategy is the written system that turns opaque GenAI behaviour into a decision: ship, hold, investigate, or roll back. It starts from failure modes and the decision you must own — not from the metrics a tool happens to expose.
For teams with at least one LLM, RAG, or agent system in or near production, in settings where a fluent error is expensive — financial services, insurance, healthcare, legal, or any organisation whose risk committee will ask for evidence.
Public leaderboards do not predict your distribution. A model that looks calm on a general benchmark can still invent a citation, skip a policy, or take the wrong tool on the one case that matters. Strategy work begins with the failures that would actually hurt: product, engineering, risk, and domain SMEs in the same room, naming what must never ship.
From that list we write a task taxonomy and slices — happy path, edge, adversarial, demographic — so the later dataset and harness have somewhere to live. A score without a slice is a number you cannot act on.
Ownership is named in the same pass. If no one owns the rubric when the model card changes, the strategy was a workshop, not a system. The memo says who may move a threshold, and who must be in the room when they do.
The rubric is the judgment, written down. Pointwise or pairwise, pass/fail or graded, single-turn or trajectory: the form follows the failure. Each dimension has a definition, a method, and a threshold. Faithfulness is not groundedness. Calibration is not confidence. Operational risk is not a style score.
The release policy is the line. What is good enough to ship, what is a hold, what is a rollback. Decision thresholds — the name the site used first — live here: a go-live memo the organisation can keep after we leave.
A threshold workshop is this work in two days: the line written, the dimensions agreed, the first cases named. A full strategy engagement is the same line, then the taxonomy and the mapping that let a harness enforce it.
Where the system is high-stakes, metrics map to internal policy, model-risk language, and — where it applies — obligations such as the EU AI Act. The mapping is explicit. A dashboard that cannot be read by legal is not a governance artifact.
We do not issue certificates. We produce the evaluation design and the first evidence pack so your risk function has something other than a vendor narrative. What they then file is theirs.
Typical artifacts. Scope follows the use case, not a template dump.
Which errors are in scope, and who owns them
What is tested, including the rare path
Dimensions, methods, and agreement rules
Ship, hold, investigate, rollback
How a score becomes a committee sentence
Typical shape: Discovery / Audit Sprint, or Framework Design + Pilot, typically 2–4 weeks for a sprint; 6–10 weeks for a first framework. Duration is typical, not contractual. Formats and deliverables are written on the evaluation ledger.
Get Started
Write with a system in mind. We will tell you whether it can be evaluated, and what the first ledger would contain.
Request an evaluation