Framework · GenAI trust
Faithful by Design
AI makes mistakes — and it's very good at hiding them. A framework for making them visible, measurable and rare.
The problem
Language models are fluent by design. A wrong answer reads exactly like a right one — same confidence, same polish, sometimes with a citation that doesn't say what it claims. The risk isn't that AI makes mistakes; people do too. It's that AI's mistakes are hard to see.
Faithfulness, not just accuracy
Faithfulness asks a narrower, checkable question: is every claim in this answer supported by the sources the system was given? It can be measured claim by claim — which makes it controllable, and it's the property that lets people trust an answer without redoing the work.
In my work
The same ideas run through what I build: grounding answers in a governed knowledge graph with agreed definitions, and designing tools like Match Studio to show the underlying matched work so users can check the result for themselves.
See it happen
A confident answer, checked claim by claim.
Card spend fell 4.2% quarter on quarter, driven mainly by lower travel spend, which dropped 18% after airlines raised fares. Spend is expected to rebound in Q3.
Looks confident. Reads well. Is it true?Illustrative example — not real data.
Where mistakes hide
Six ways AI gets it wrong — fluently.
Fabrication
Invents facts, figures — even citations.
Unfaithful summary
The source says one thing; the answer says another.
Wrong definition
Uses “active customer” differently from the business.
Silent omission
Drops the caveat that changes the conclusion.
Plausible wrong SQL
Runs, returns numbers — joined the wrong tables.
False confidence
Same fluent tone whether right or guessing.
How to control it
Five gates, one feedback loop.
- Retrieve only from governed, entitled sources
- Shared definitions from the semantic layer
- Answer only from what was retrieved
- “I don't know” is a valid answer
- Tools — not the model — do maths and SQL
- Every claim links to a source span
- No citation, no claim
- Split the answer into claims; check each against its source
- Validate SQL results and numbers independently
- Below the confidence threshold → abstain
- High-stakes answers → human sign-off
Match controls to stakes
Not every answer needs every guardrail.
| ExploreDrafts, brainstorming | InformInternal analytics, search | DecideFinancial, regulatory, client-facing | |
|---|---|---|---|
| Ground in governed sources | |||
| Cite every claim | |||
| Deterministic tools for numbers & SQL | |||
| Claim-level verification | |||
| Abstain below confidence threshold | |||
| Full audit trail | |||
| Human sign-off |
Measure it
What gets measured gets trusted.
Faithfulness
Supported claims ÷ all claims
Citation precision
Citations that truly support their claim
Hallucination rate
Unsupported claims on a golden test set
Right to abstain
Says “I don't know” when sources don't cover it
Example targets — set them per use case and risk tier.
The goal isn't an AI that never makes mistakes. It's one that can't hide them.