Six models answer. See where they disagree.

GPT-6 Luna+6 Try the demo
ReasoningAuto Attach ChatGPT Mistral Claude Gemini DeepSeek Grok

01 · Ask

Start from one prompt.

Agent is where every question starts: it researches, asks six models and writes one answer. Compare and Consensus wait in the (+) menu, next to files and reasoning.

GPT-6 Luna+6
ReasoningAuto Attach ChatGPT Mistral Claude Gemini DeepSeek Grok

Please avoid personal or sensitive data. AI models can make mistakes.

  • The field. One question, nothing to configure first.
  • The models. The agent that writes, and the six it checks with.
  • Send. The only filled control on the page.

02 · Run

The agent shows its work.

It asks six models in parallel, writes one answer from what they said, and has that answer checked against all six. Where they contradict each other on a fact, a separate judge reads the sources they cited. Scroll to move the run.

Can a heat pump heat our 1978 house with the original radiators?

Working for 0s

Probably yes, with a few radiators changed, not all of them.

  • Get a room-by-room heat-loss calculation before anyone quotes a size.
  • Rooms that stay warm at 50 °C keep their radiators.
  • Run at 55 °C, or replace more radiators for 45 °C: three models each way.

A single answer would have handed you all three as settled.

43/100 agreement Contradictions 2 Answers 6

03 · Decide

The doubt is marked in the answer.

One answer, read top to bottom, with the disagreement written into it rather than hidden behind it. Four things carry that, and you can try each one here.

  1. 01

    How far apart they were

    Every checked answer carries one number: how much the six answers actually agreed. An uninvolved model works it out, so the model that wrote the answer is not also grading it.

    43 Partial agreement 1 critical · 1 minor detail · disputed: 55 °C with most radiators, or more radiators for 45 °C, 6 models compared Analysis by Gemini independent of the consensus engine
  2. 02

    Who backed each claim

    The small number behind a sentence is how many models supported it. Rest on it, or tap it, and it names them: who agreed, who deviated, and in what words.

    Try it on the 4/6.

    Get a room-by-room heat-loss calculation before anyone quotes a heat pump size

    Rooms that stay warm keep their radiators; only the rooms that fall short need larger or fan-assisted ones

  3. 03

    Where they pull apart

    Every checked sentence is highlighted in the colour of its result: green where every model agreed, amber where they split, red where they contradict each other. The mark sits on the sentence it belongs to, so you never have to hold two documents open at once.

    • Size the heat pump from the calculation green · all of them agreed
    • Rooms that stay warm keep their radiators amber · support is split
    • Keep the gas boiler as a backup, or not grey · a difference in detail
    • 55 °C, or more radiators for 45 °C red · they contradict each other
  4. 04

    What they actually said

    Click a mark and the disagreement opens: both camps, named, with the sentence each side actually wrote and a note on what is worth checking yourself.

    Try the marked sentence.

    Three models would keep most radiators and run at 55 °C, three would replace more of them to run at 45 °C or below.

    Contradiction · critical Run at 55 °C with most radiators, or replace more of them for 45 °C?
    ChatGPT, Gemini, Grok
    55 °C. Keep every radiator that passes the test; the last ten degrees save little.
    replacing every radiator to get there costs more than that difference will save in a decade
    Claude, Mistral, DeepSeek
    45 °C. Replace more radiators now; the efficiency is paid every winter.
    The efficiency difference is around a quarter, and you pay it every winter for twenty years
    Worth verifying: ask for the heat pump's rated efficiency at 35 °C and at 55 °C from its datasheet, and get a price for the extra radiators. That turns the disagreement into arithmetic.

The same lens on other questions

  • 94 “Is nuclear power a low-carbon energy source?” High agreement · nothing marked in the answer
  • 84 “Which programming language should I learn first?” Strong agreement · one difference in emphasis, no contradiction
  • 41 “Can I take ibuprofen while on blood thinners?” Partial agreement · one critical contradiction, marked in the answer

Proof, not promises

We benchmarked the consensus.

On 314 closed-book MMLU-Pro questions, the consensus ranks first overall, and pulls clear exactly where the models disagree.

When the models disagree

Accuracy on the questions where they split · MMLU-Pro

72.1%
70.6%
69.1%
64.7%
58.8%
44.1%
39.7%
consens.io Claude GPT-5.5 Gemini Grok DeepSeek Mistral
#1Highest accuracy overall and on the hard questions. The top of the field is a statistical tie.
+32ptsLead over the weakest model on questions where the answers split.
See the full benchmark

Live from real runs

No model wins every time.

Even ChatGPT, the model with the highest best-answer rate, gives the judge's pick in 51% of the runs it is in. In the other 49%, the best answer came from another model.

Open the model pulse
  1. ChatGPT 51%
  2. Gemini 17%
  3. Muse 15%
  4. DeepSeek 14%
  5. Kimi 14%

Share of its runs in which a model's answer was the judge's pick · 385 judged runs · the tick marks a fair share by chance.

04 · Monitor

Know when the answer changes.

Turn any question into a Watch. consens.io reruns it on your schedule and alerts you only when the result materially moves.

Consensus Watch Active

Has the EU guidance for general-purpose AI models changed?

WeeklyMaterial changes onlyTelegram
Jul 14
Baseline72/100 model agreement
Jul 21
ChangedNew transparency requirements are now included.
34/100 movement
Consensus changedTelegram · now

The latest run now includes new transparency requirements. The core conclusion moved materially.

Create a Watch

Why consens.io

Built for decisions that need a second view.

Compare

Original answers stay visible

Open Model answers to read each original response or compare two answers side by side.

Verify

Disagreement is marked in place

Where the models pull apart, the sentence itself carries the mark, instead of the doubt being smoothed into one confident answer.

Control

You choose the checks

Pick the agent's model, the six it checks with, and whether it reasons first. Where the models contradict each other on a fact, a separate judge reads the sources they cited, and its verdict sits with the contradiction under the answer. It does not fact-check the whole answer.

Model coverage

One workspace for leading AI providers.

Direct comparison across major model families, so every answer can be checked against independent alternatives.

Zero data retention on every model call. Requests are routed only to provider endpoints that don't keep your prompts or answers.How it works ›