Skip to content

Solutions

Risk, safety & compliance evaluation

Some failures are not quality in the ordinary sense. They are policy, leakage, unfairness, or a sentence you cannot defend to a regulator. Risk evaluation is a scored red team and a file the committee can keep.

For high-stakes and regulated deployments — and for any team whose board has started asking whether the system can be shown, not only demoed.

01

Red-teaming with a catalog

Adversarial tests are designed from the failure-mode map: jailbreaks, leakage, disallowed advice, prompt injection, retrieval of the wrong document. Findings are a catalog with patches, not a slide of scary examples. The two-week red team on the ledger is this work in concentrated form.

We re-test after the patch. A red team that does not return is a performance. The catalog is versioned with the system it was run against.

02

Policy, fairness, safety

Instruction adherence is scored against system, policy, and format constraints. Bias and demographic performance are sliced, not averaged. Content safety is a dimension with a threshold, including “zero critical” where that is the only acceptable line.

Operational risk on the ledger — policy, leakage, escalation — is this family of failures. A 94% faithfulness score does not excuse a single critical leak.

03

Evidence that can leave the room

Audit-ready packages — eval factsheets or the format your model-risk function already uses — record who authored the rubric, what data was used, how agreement was measured, and what the release threshold was. If it cannot be shown, it is not evidence.

We write in the register the committee already reads. We do not invent a parallel vocabulary that only the vendor understands.

04

What the file contains

For internal risk, legal, and — when required — external review.

Failure catalog

Engineering and security

Policy adherence scores

Product and compliance

Fairness / slice report

Risk and the business owner

Eval factsheet

Model risk, audit, regulator

05

Engagement

Typical shape: Red Team (concentrated) or a Risk & Compliance workstream on a sprint / retainer, typically 2 weeks for a red team; longer when documentation must meet a named control framework. Duration is typical, not contractual. Formats and deliverables are written on the evaluation ledger.

06

Questions

Is this a legal opinion?
No. We produce evaluation evidence. Counsel and your risk function decide what it means under the rules that bind you. The point is that they have something other than a demo.
Do you certify against the EU AI Act?
We do not issue certifications. We map evaluation design and results to the questions those obligations raise, in language a compliance reader can use.
Will findings be published?
No. Client work stays under the SOW and data rules. Public writing uses anonymised patterns, never a named incident, unless you ask otherwise.

Get Started

When a model is ready is not a feeling.

Write with a system in mind. We will tell you whether it can be evaluated, and what the first ledger would contain.

Request an evaluation