Colay / Guides
Many answers from one model, or discussion between models?
You can solve a question several times and choose the most frequent answer. You can also ask participants to discuss competing solutions and produce a synthesis. Both approaches use more than one attempt, but they test different ideas. Understanding that distinction makes it easier to decide which part of Consensus is useful for a particular job.
External research results and illustrative calculations are not measurements of Colay performance.
What the self-consistency paper actually measured
Wang et al., ICLR 2023, report PaLM-540B accuracy on GSM8K rising from 56.5% with greedy chain-of-thought to 74.4% with self-consistency: 17.9 percentage points. Results average 10 runs with 40 sampled outputs per run. These are research results, not Colay measurements. Wang et al., ICLR 2023
GSM8K’s test split contains 1,319 elementary mathematics word problems. It provides checkable final answers, rather than representing every research, writing or decision-making task. GSM8K dataset, OpenAI
A practical interpretation is that the first response need not be the best answer available from a model. Repetition also consumes resources. A comparison between forty attempts and one attempt, without accounting for cost, does not settle which workflow is preferable under your spending limit or acceptable waiting time.
Separate three operations in the workflow
For a practical evaluation, distinguish candidate generation, candidate assessment and final composition. Generation determines which possible solutions become available. Assessment decides which meet the task requirements. Composition produces a usable answer for the reader. Each operation can be improved separately, and success at one stage does not automatically repair a failure at another.
A correct solution might already be present but lose a vote. Alternatively, the right candidate may be selected while the final prose drops an essential condition. Preserve intermediate answers when evaluating a workflow. Without them, it is difficult to tell whether you need more candidates, a better selection rule, or a clearer instruction to the coordinator.
When answer frequency is informative
Consider a hypothetical arithmetic question producing five answers: 42, 42, 42, 24 and 24. Majority voting selects 42. That selection is reproducible, but correctness still depends on the problem and the calculation. If the three matching responses all confuse minutes with hours, the vote preserves the same error. Vote share is not automatically the probability that an answer is true.
Open questions introduce an earlier difficulty: deciding what counts as agreement. Two responses might recommend the same product for incompatible reasons. Two different recommendations might both work under different constraints. Before counting matches, define the output you need: a decision, its applicability conditions, required facts and any information that remains missing.
Why discussion might be useful for your task
For a short question with one numerical answer, checking the calculation may be enough. A software architecture decision benefits from comparing assumptions about scale, maintenance and unknown constraints. In that setting, discussion can be evaluated for its ability to reveal incompatible premises, rather than simply producing the most frequently repeated recommendation.
Suppose two participants recommend different designs because they assume different data volumes. A useful synthesis identifies the point at which the preference changes. Ask for a compact account of what is known, what is assumed and what observation would settle the disagreement. This is our proposed task-design method, not a guarantee that a model will discover every dependency.
Compare workflows without changing the question
Set a spending ceiling and a shared scoring rubric in advance. Compare an ordinary answer, separate attempts with an explicit selection rule, and discussion. Equal call counts do not establish equal costs: context length, output size and model choice all matter. Record actual spending and delay beside quality, instead of treating additional work as a free improvement.
A hypothetical calculation makes the trade-off visible. Workflow A produces 18 acceptable results from 20 tasks for 40 accounting units. Workflow B produces 19 for 80 units. The additional acceptable result costs 40 units. That does not mean A is preferable: the extra successful task may be especially valuable. It makes the cost of improvement explicit enough to discuss.
What this means for Consensus in Colay
Colay Consensus uses discussion among participants and a coordinator’s synthesis. Self-consistency numbers do not describe that mode: a different arrangement needs its own evaluation. To investigate the benefit, choose a recurring task, supply the same factual material and keep an ordinary-mode answer beside the Consensus result. Evaluate both against the criteria established beforehand.
Ask for the basis of the recommendation, its strongest material objection and the conditions that would change it. Check sources and calculations separately. Where a formula or an executable test can establish correctness, use that check instead of the number of agreeing participants. Multiple attempts expand the candidate pool; the workflow still needs a defensible rule for accepting an answer.
Sources and methodology
- Wang et al., ICLR 2023
Table 2; section 3.2.
- GSM8K dataset, OpenAI
The main test split contains 1,319 elementary mathematics word problems.
Bring your next question to Colay
Choose a model, use Auto, or bring several perspectives together with Consensus.