← All projects

Framework · GenAI trust

Faithful by Design

AI makes mistakes — and it's very good at hiding them. A framework for making them visible, measurable and rare.

FaithfulnessEvaluationGuardrailsFramework

The problem

Language models are fluent by design. A wrong answer reads exactly like a right one — same confidence, same polish, sometimes with a citation that doesn't say what it claims. The risk isn't that AI makes mistakes; people do too. It's that AI's mistakes are hard to see.

Faithfulness, not just accuracy

Faithfulness asks a narrower, checkable question: is every claim in this answer supported by the sources the system was given? It can be measured claim by claim — which makes it controllable, and it's the property that lets people trust an answer without redoing the work.

In my work

The same ideas run through what I build: grounding answers in a governed knowledge graph with agreed definitions, and designing tools like Match Studio to show the underlying matched work so users can check the result for themselves.

A confident answer, checked claim by claim.

Why did card spend fall in Q2?
AI assistant

Card spend fell 4.2% quarter on quarter, driven mainly by lower travel spend, which dropped 18% after airlines raised fares. Spend is expected to rebound in Q3.

Looks confident. Reads well. Is it true?
Sources the AI was given
S1 Q2 spend report
Total card spend: −4.2% QoQ
Active cards: +0.6% QoQ
S2 Category breakdown
Travel: −11.0% QoQ
Dining: −1.3% QoQ
Groceries: +0.8% QoQ

Illustrative example — not real data.

Six ways AI gets it wrong — fluently.

Fabrication

Invents facts, figures — even citations.

Unfaithful summary

The source says one thing; the answer says another.

Wrong definition

Uses “active customer” differently from the business.

Silent omission

Drops the caveat that changes the conclusion.

Plausible wrong SQL

Runs, returns numbers — joined the wrong tables.

False confidence

Same fluent tone whether right or guessing.

Five gates, one feedback loop.

Qquestiontrusted answer1 GROUNDgoverned sources2 CONSTRAINanswer from context3 CITEclaim → source4 VERIFYcheck each claim5 DECIDEabstain or escalateAbstain or routeto a humanMONITORGolden test sets · Faithfulness score over time · User flags · Regression tests on every model change
1 · Ground
  • Retrieve only from governed, entitled sources
  • Shared definitions from the semantic layer
2 · Constrain
  • Answer only from what was retrieved
  • “I don't know” is a valid answer
  • Tools — not the model — do maths and SQL
3 · Cite
  • Every claim links to a source span
  • No citation, no claim
4 · Verify
  • Split the answer into claims; check each against its source
  • Validate SQL results and numbers independently
5 · Decide
  • Below the confidence threshold → abstain
  • High-stakes answers → human sign-off

Not every answer needs every guardrail.

ExploreDrafts, brainstormingInformInternal analytics, searchDecideFinancial, regulatory, client-facing
Ground in governed sources
Cite every claim
Deterministic tools for numbers & SQL
Claim-level verification
Abstain below confidence threshold
Full audit trail
Human sign-off
Required Recommended Not needed

What gets measured gets trusted.

≥ 95%

Faithfulness

Supported claims ÷ all claims

≥ 98%

Citation precision

Citations that truly support their claim

≤ 1%

Hallucination rate

Unsupported claims on a golden test set

≥ 90%

Right to abstain

Says “I don't know” when sources don't cover it

Example targets — set them per use case and risk tier.

The point

The goal isn't an AI that never makes mistakes. It's one that can't hide them.

Next projectMatch Studio
→