Provenance
Source system, time, sampling rule
Solutions
A golden dataset is not a handful of favourite prompts. It is a versioned, provenance-tracked sample of the distribution you will actually serve — including the cases you would rather not look at.
For teams whose current “eval set” is a spreadsheet of happy paths, or whose production logs cannot yet be used as evidence. Required before a harness or a release gate can mean anything.
We sample from logs or historical cases under the data rules you already have. Client data stays in the agreed environment. Anonymisation is part of the artifact, not an afterthought. Every item carries provenance: where it came from, who labeled it, which version of the rubric applied.
The set is sliced. A single average hides the slice that fails. Challenge cases — rare, high-risk, adversarial — are constructed when production does not yet contain them, then validated by a human. Synthetic is a supplement, never the whole ledger.
A pipeline audit without a set is a reading of whatever happened to be in the room. The dataset is how that reading becomes repeatable when the model, the prompt, or the index moves.
Ground truth is a claim. We treat it as one: guidelines, more than one annotator where the stake requires it, agreement measured, disagreements adjudicated. The dataset card records that process so a later auditor is not asked to take the score on faith.
Where the label is preference rather than truth — two fluent answers, one better — the set still carries the guideline version and the pair. Rankings without a protocol are taste.
When the product, the policy, or the model changes, the set changes with it — or it is frozen and named. Access control belongs with the rest of your evaluation assets. A golden set that cannot be reproduced is a demo.
Retention follows your existing records policy. In regulated settings the evaluation set is itself an artifact that may need to be produced later. We write that down at the start, not after a request from audit.
Minimum fields. Further columns follow the rubric.
Source system, time, sampling rule
Happy path, edge, adversarial, demographic
What “correct” meant on that date
Where humans disagreed, and how it was closed
Who may see it, and when it is destroyed
Typical shape: Framework Design + Pilot, or Full Build, typically inside a 6–10 week pilot; longer when several use cases share a library. Duration is typical, not contractual. Formats and deliverables are written on the evaluation ledger.
Get Started
Write with a system in mind. We will tell you whether it can be evaluated, and what the first ledger would contain.
Request an evaluation