Compare
Original answers stay visible
Read each model response side by side before the consensus layer summarizes it.
No account needed. The demo replays a real run: six independent answers, then one consensus with the contradictions marked inside it.
Live model pulseSee which model answers most often earn the judge's pick in real runs. View leaderboard ›01 · Ask
One field, one run switch, one send. The app opens exactly here, with no modes to learn before the first question.
What should the models cross-check?
Please avoid personal or sensitive data. AI models can make mistakes.
02 · Run
Six models answer in parallel, then the consensus is written and an uninvolved model checks it for contradictions. Scroll to move the run.
The Nile is the conventional longest river, but the Amazon has a credible claim.
03 · Decide
One answer, read top to bottom — with the disagreement written into it rather than hidden behind it. Four things carry that, and you can try each one here.
01
Every run opens with one number: how much the six answers actually agreed. An uninvolved model works it out, so the engine that wrote the answer is not also grading it.
02
The small number behind a sentence is how many models supported it. Rest on it, or tap it, and it names them: who agreed, who deviated, and in what words.
Try it on the 3/6.
By convention and in most references the Nile leads at about 6,650 km
The disputed 2007 expedition puts the Amazon at 6,992 km
03
Agreement gets no decoration at all — only the count. A difference in detail gets a fine rule under the sentence, a real contradiction a heavier amber one. The mark sits on the sentence it belongs to, so you never have to hold two documents open at once.
04
Click a mark and the disagreement opens: both camps, named, with the sentence each side actually wrote and a note on what is worth checking yourself.
Try the marked sentence.
Whether the Amazon overtakes the Nile depends on where its mouth is drawn.
the Nile is traditionally listed as the longest river at about 6,650 km
if the Pará estuary is counted, the Amazon may exceed 7,000 km
The same lens on other questions
04 · Monitor
Turn any consensus into a Watch. consens.io reruns the same question on your schedule and alerts you only when the result materially moves.
The latest run now includes new transparency requirements. The core conclusion moved materially.
Create a WatchProof, not promises
On 314 closed-book MMLU-Pro questions, the consensus ranks first overall, and pulls clear exactly where the models disagree.
When the models disagree
Accuracy on the questions where they split · MMLU-Pro
Why consens.io
Compare
Read each model response side by side before the consensus layer summarizes it.
Verify
Where the models pull apart, the sentence itself carries the mark, instead of the doubt being smoothed into one confident answer.
Control
Select providers, switch reasoning modes, and decide which outputs count toward consensus.
Model coverage
Direct comparison across major model families, so every answer can be checked against independent alternatives.
What it costs
consens.io is a one-person project and still very much a test. Every question you send calls six model providers at once, and that bill lands on me. That's the only reason there are limits at all.
Today
No checkout, no price, no waiting list. Where something is switched off, it's one of the expensive runs, and those I enable per account rather than for everyone, because they cost a multiple of a normal one.
The limits
Your remaining runs are shown in the app, and that number moves with what the month allows. You can also put your own provider API keys into the settings; then my limits stop applying and the costs are yours.
Later
There's no membership today and I'm not collecting sign-ups for one. But if people keep telling me they want the expensive runs regularly, that's how I'd make it possible. Write to me either way at contact@consens.io, and have fun with it :)
Get the product view