TECHNOLOGY
Disagreement is the measurement.
Most AI tooling treats model disagreement as an error to suppress. We treat it as the only cheap signal available about whether an answer can be trusted.
A single model is fluent when it is right and fluent when it is wrong, so its confidence tells you very little. Asking that same model to check itself returns the same blind spots with more reassurance attached. Two models from rival vendors do not share a training pipeline, and where they diverge is where you should look.
A ruling, not an average
Averaging two answers hides the disagreement, which is the part worth having. A judge produces a ruling and the dissent stays attached to it.
A confidence score that means something
The score reflects how much survived contest, not how fluent the output sounded. It is a probability raiser with receipts, not an oracle, and we say so on the product site too.
A disagreement map you can hand to someone
The output is designed to be forwarded. Where the models split, what each claimed, and how it was resolved, in a form a colleague or a regulator can read without taking your word for it.
What we do not publish.