Skip to content

Platform

Evaluation ledger

Representative figures for how Seuil presents evidence. Not a live product, and not a vendor ranking — a ledger of measure.

01

Model comparison

Faithfulness, hallucination, and citation as percentages. Cost is an index against Frontier A.

Frontier A

Faithfulness
94.2
Hallucination
1.8
Citation
91.0
Latency
2.4s
Cost
1.00

Frontier B

Faithfulness
92.8
Hallucination
2.3
Citation
88.4
Latency
1.9s
Cost
0.86

Frontier C

Faithfulness
90.1
Hallucination
3.1
Citation
86.2
Latency
2.1s
Cost
0.72

Open-weight

Faithfulness
86.4
Hallucination
4.6
Citation
79.8
Latency
1.1s
Cost
0.18

Fine-tune, anon.

Faithfulness
88.9
Hallucination
3.4
Citation
84.1
Latency
0.8s
Cost
0.31
02

Judgment rubric

The same dimensions, applied to a pipeline rather than a model card.

Faithfulness

≥ 0.92

Claims supported by the source context actually provided

Groundedness

≥ 0.90

No unsourced assertion in a high-stakes answer

Instruction adherence

≥ 0.95

Compliance with system, policy, and format constraints

Calibration

ECE ≤ 0.06

Agreement between stated confidence and observed accuracy

Operational risk

0 critical

Policy, leakage, and escalation failures on critical cases

03

Engagements

How the work is framed. Duration is typical, not contractual.

Pipeline Audit

3–4 weeks

Scored ledger and threshold memo

Continuous Evaluation

Quarterly

Living scorecard and drift notes

Red Team

2 weeks

Failure catalog and patches

Threshold Workshop

2 days

Written go-live decision policy

Get Started

When a model is ready is not a feeling.

Write with a system in mind. We will tell you whether it can be evaluated, and what the first ledger would contain.

Request an evaluation