Colay / Guides

When several AI models share the same blind spot

Three matching answers can feel like three checks. That interpretation depends on how the answers were produced. If every participant relies on the same mistaken premise, agreement repeats the premise instead of testing it. To judge an AI consensus, examine the routes to the answer as well as the answer itself.

External research results and illustrative calculations are not measurements of Colay performance.

Similarity and correctness are different questions

Artificial Hivemind, a NeurIPS 2025 study, found within-model repetition and cross-model similarity on open-ended prompts. Those prompts admit multiple plausible answers. This is evidence about response diversity, not a measured probability that models share factual errors. Artificial Hivemind, NeurIPS 2025

Consider a product team asking several agents why customers leave. Similar explanations might be correct, fashionable, or borrowed from an incomplete brief. The useful question is what observation could disprove each explanation. Without that test, different wording may hide a single underlying assumption, while identical wording might simply reflect an obvious fact.

Dietterich’s ensemble analysis explains why voting can benefit from accurate classifiers with differing errors. Its independent-error calculation is a conditional mathematical model, not a property automatically supplied by using different model names. Dietterich: Ensemble Methods in Machine Learning

A calculation with deliberately simplified assumptions

The following is our hypothetical binary classification example, not a Colay experiment. Assume three voters, a single correct label, an error probability of 20% for each voter, and majority voting without discussion. With independent errors, a wrong majority needs exactly two errors or three errors. Its probability is 3 × 0.2² × 0.8 + 0.2³ = 0.104, or 10.4%.

Now change only the dependence structure. If all three always succeed or fail together, the majority is wrong 20% of the time. As a third constructed case, let half the tasks use that shared-error mechanism and half use independent errors. Each voter still has a 20% error rate, but majority error becomes 0.5 × 20% + 0.5 × 10.4% = 15.2%.

Constructed example: the same individual error rate, different joint outcomes
Assumed relationshipIndividual errorMajority error
Independent errors20%10.4%
Half shared, half independent20%15.2%
Completely shared errors20%20%

Why a discussion is not a majority-vote experiment

A conversational coordinator can accept a minority objection, combine two partial solutions, or introduce a fresh mistake. Participants can also change their answers after reading one another. The calculation above therefore cannot predict the accuracy of a discussion. It isolates one issue: counting outputs does not tell you how much additional evidence they contain.

In Colay Consensus, several agents discuss a request and a coordinating agent synthesizes the conclusion. That workflow does not establish statistical independence between participants. Separate initial answers can make differences easier to inspect, but even separately generated answers can share training influences, supplied documents, omissions, or interpretations of the task.

Inspect dependencies before adding participants

For a concrete review, create a small evidence ledger. Record the disputed claim, the document supporting it, the assumption connecting document to conclusion, and a possible counterexample. Ask each reviewer to fill missing fields. Three references to the same press release remain one origin of evidence, even if three websites reproduce it.

Assign checks with different observable outputs. One reviewer can test arithmetic, another can trace quotations to source passages, and another can search the brief for excluded cases. These assignments are proposed working methods, not guarantees of independent errors. Their value is that you can see whether a participant contributed a new check rather than another endorsement.

  • Keep an objection until its supporting evidence has been addressed.
  • Record which facts came from the prompt and which were independently checked.
  • Prefer a demonstrated counterexample to a confident count of agreeing agents.

Measure useful disagreement

On tasks with known answers, mark where both agents fail, where only one fails, and where both succeed. Investigate shared failures first: they reveal gaps that another similar participant may not repair. For creative tasks, replace a single correctness label with explicit requirements, such as distinct concepts, audience fit, or coverage of competing constraints.

Stop adding participants when new responses no longer change the evidence, reveal a constraint, or improve a tested outcome. Colay subscriptions have credits and limits, so repeated agreement also has a usage cost. A compact discussion that preserves one decisive objection can be more useful than a long discussion ending in unanimous but unsupported prose.

Sources and methodology

  1. Artificial Hivemind, NeurIPS 2025

    Open-ended generation study.

  2. Dietterich: Ensemble Methods in Machine Learning

    Classical ensemble assumptions.

Bring your next question to Colay

Choose a model, use Auto, or bring several perspectives together with Consensus.

Open Colay