Colay / Guides

When AI debate adds value to an answer

The case for Consensus starts with a testable question: does discussion catch an error that the first answer missed? Several models may contribute useful observations, but participant count alone proves little. This article examines a research result and turns it into a practical evaluation method, while keeping published experiments separate from claims about a particular product.

External research results and illustrative calculations are not measurements of Colay performance.

A small experiment worth reading carefully

In Du et al.’s 2023 preprint, preceding their ICML 2024 publication, Bard solved 11 of 20 GSM8K problems, ChatGPT solved 14, and joint debate solved 17: 55%, 70% and 85%. These are results for specific models on a small sample, not measurements of Colay. Du et al., 2023/ICML 2024

Three additional correct answers in a 20-question sample are a signal, not a universal improvement rate. Testing another language, model or document changes the conditions.

Identify the improvement you actually need

Consider a hypothetical launch decision with three options. One answer recommends the fastest route. A second spots a supplier dependency. A third notices that the estimated timeline covers development but excludes approval. A useful synthesis corrects the scope of the estimate and explains whether that changes the recommendation. Merely listing three perspectives does not complete that work.

Our suggested criterion is a change in the decision, not an impressive discussion. Was a missing constraint recovered? Was an incorrect unit fixed? Did the result label an unknown that the first answer treated as a fact? If the only change is a longer response, the benefit remains unproven. This criterion works across products and model combinations.

Agent count and model diversity are separate variables

ReConcile, published at ACL 2024, investigates a distinct protocol: different models discuss answers and use confidence-weighted voting. It evaluates that arrangement, rather than validating every service that coordinates several agents. Chen, Saha and Bansal, ACL 2024

In your own review, distinguish the number of candidate answers, the differences in their starting approaches, and the rule used to form a final answer. Three roles assigned to one model do not become three independent knowledge sources. Different model names do not establish independent errors either: every participant can inherit the same false premise from the request.

Count regressions alongside corrections

Here is a hypothetical example, not a research finding. Across 100 preselected tasks, the initial answer is correct on 60. Discussion repairs 14 wrong answers but changes six correct answers into wrong ones. The result is 68 correct answers: a net gain of eight percentage points. Reporting only the 14 repairs would conceal a material part of the outcome.

Also examine the consequences. Fixing a typo and losing a budget constraint have different costs. Keep the initial answer beside the final one and record why each material claim changed. This can show where a workflow helps even when its overall score barely moves. For consequential decisions, a reviewer who understands the subject must determine whether those changes are justified.

Hypothetical evaluation across 100 tasks
TransitionCountInterpretation
Wrong → correct14Corrections
Correct → wrong6Regressions
Net change+868 correct instead of 60

Run a small but fair comparison

Choose recent real tasks with an established acceptable outcome. Write down the grading criteria, spending limit and maximum acceptable delay before generating responses. Give both workflows the same evidence. Do not select only questions the first model has already failed: that would tilt the comparison toward any second attempt, regardless of whether discussion contributed anything.

Review outputs without exposing the workflow name. Score correctness, coverage of material constraints and the amount of human revision separately. Preserve unsuccessful examples alongside successes. If colleagues help evaluate the results, let them score independently before discussing them. Otherwise, agreement in the review can hide genuine differences in how people understand the task.

Apply the question to Colay Consensus

In Consensus, participating agents discuss a task and a coordinator produces a synthesis. A practical request asks them to compare alternatives against explicit criteria, identify disputed assumptions and avoid converting missing data into confidence. Supply the factual basis and ask for important claims to be connected to particular passages in that material. Review those connections yourself.

This article does not report a measured Colay accuracy improvement. The research supports investigating discussion when competing interpretations matter; the value of a particular run depends on its actual output. Preserve unresolved disagreements in the final document when evidence is insufficient. A useful result may end with a request for more information instead of unanimous approval.

Sources and methodology

  1. Du et al., 2023/ICML 2024

    Section 3: 20 GSM8K problems.

  2. Chen, Saha and Bansal, ACL 2024

    ReConcile: a distinct protocol using diverse models, discussion and confidence-weighted voting.

Bring your next question to Colay

Choose a model, use Auto, or bring several perspectives together with Consensus.

Open Colay