Colay / Guides
When AI debate adds value to an answer
The case for Consensus starts with a testable question: does discussion catch an error that the first answer missed? Several models may contribute useful observations, but participant count alone proves little. This article examines a research result and turns it into a practical evaluation method, while keeping published experiments separate from claims about a particular product.
External research results and illustrative calculations are not measurements of Colay performance.
A small experiment worth reading carefully
In Du et al.’s 2023 preprint, preceding their ICML 2024 publication, Bard solved 11 of 20 GSM8K problems, ChatGPT solved 14, and joint debate solved 17: 55%, 70% and 85%. These are results for specific models on a small sample, not measurements of Colay. Du et al., 2023/ICML 2024
Three additional correct answers in a 20-question sample are a signal, not a universal improvement rate. Testing another language, model or document changes the conditions.
Identify the improvement you actually need
Consider a hypothetical launch decision with three options. One answer recommends the fastest route. A second spots a supplier dependency. A third notices that the estimated timeline covers development but excludes approval. A useful synthesis corrects the scope of the estimate and explains whether that changes the recommendation. Merely listing three perspectives does not complete that work.
Our suggested criterion is a change in the decision, not an impressive discussion. Was a missing constraint recovered? Was an incorrect unit fixed? Did the result label an unknown that the first answer treated as a fact? If the only change is a longer response, the benefit remains unproven. This criterion works across products and model combinations.
Agent count and model diversity are separate variables
ReConcile, published at ACL 2024, investigates a distinct protocol: different models discuss answers and use confidence-weighted voting. It evaluates that arrangement, rather than validating every service that coordinates several agents. Chen, Saha and Bansal, ACL 2024
In your own review, distinguish the number of candidate answers, the differences in their starting approaches, and the rule used to form a final answer. Three roles assigned to one model do not become three independent knowledge sources. Different model names do not establish independent errors either: every participant can inherit the same false premise from the request.
Count regressions alongside corrections
Here is a hypothetical example, not a research finding. Across 100 preselected tasks, the initial answer is correct on 60. Discussion repairs 14 wrong answers but changes six correct answers into wrong ones. The result is 68 correct answers: a net gain of eight percentage points. Reporting only the 14 repairs would conceal a material part of the outcome.
Also examine the consequences. Fixing a typo and losing a budget constraint have different costs. Keep the initial answer beside the final one and record why each material claim changed. This can show where a workflow helps even when its overall score barely moves. For consequential decisions, a reviewer who understands the subject must determine whether those changes are justified.
| Transition | Count | Interpretation |
|---|---|---|
| Wrong → correct | 14 | Corrections |
| Correct → wrong | 6 | Regressions |
| Net change | +8 | 68 correct instead of 60 |
Run a small but fair comparison
Choose recent real tasks with an established acceptable outcome. Write down the grading criteria, spending limit and maximum acceptable delay before generating responses. Give both workflows the same evidence. Do not select only questions the first model has already failed: that would tilt the comparison toward any second attempt, regardless of whether discussion contributed anything.
Review outputs without exposing the workflow name. Score correctness, coverage of material constraints and the amount of human revision separately. Preserve unsuccessful examples alongside successes. If colleagues help evaluate the results, let them score independently before discussing them. Otherwise, agreement in the review can hide genuine differences in how people understand the task.
Apply the question to Colay Consensus
In Consensus, participating agents discuss a task and a coordinator produces a synthesis. A practical request asks them to compare alternatives against explicit criteria, identify disputed assumptions and avoid converting missing data into confidence. Supply the factual basis and ask for important claims to be connected to particular passages in that material. Review those connections yourself.
This article does not report a measured Colay accuracy improvement. The research supports investigating discussion when competing interpretations matter; the value of a particular run depends on its actual output. Preserve unresolved disagreements in the final document when evidence is insufficient. A useful result may end with a request for more information instead of unanimous approval.
Sources and methodology
- Du et al., 2023/ICML 2024
Section 3: 20 GSM8K problems.
- Chen, Saha and Bansal, ACL 2024
ReConcile: a distinct protocol using diverse models, discussion and confidence-weighted voting.
Bring your next question to Colay
Choose a model, use Auto, or bring several perspectives together with Consensus.