Colay / Guides
The risks of one AI model: checking a confident answer
A single AI model can produce a persuasive memo while leaving its most important assumption untested. The useful question is which claims could change your decision. This article separates those claims from the prose and explains how a second perspective can lead to an actual check.
External research results and illustrative calculations are not measurements of Colay performance.
What the research measures
Farquhar et al., published in Nature on 19 June 2024, reported a mean AUROC of 0.790 across 30 task–model combinations for semantic entropy. The method detects unstable confabulations. AUROC measures discrimination between correct and incorrect answers; it does not mean that 79% of answers are correct. This was not a Colay evaluation. Farquhar et al., Nature, 2024
NIST AI 600-1 identifies confidently presented false information as a distinct generative AI risk. It is a risk-management framework, not a model leaderboard or a product error-rate study. NIST AI 600-1, 2024
Our practical recommendation is to assess evidence separately from presentation. An elegant table can contain an unsupported input. A long explanation can omit the exception that matters to your situation. Before requesting more analysis, mark the few statements that would change what you do. Those statements deserve explicit checks even when the rest of the answer looks convincing.
wrong and arbitrary
Farquhar et al., Nature, 2024
Nature uses this phrase for confabulations that are sensitive to incidental generation details: a specific class of error.
Four ways a plausible answer can fail
Start with the factual layer: does the cited document exist, and is its date correct? Next, check applicability: does the statement cover your location, product and time period? Then inspect arithmetic, especially units and denominators. Finally, ask about completeness: which realistic alternative has been left out? A response can pass three of these checks and still fail the decision.
With one answer, those questions can collapse into a single impression that everything looks reasonable. Another model is useful when it separates them and identifies a specific weakness. Rephrasing the same conclusion adds little evidence. The relevant unit of progress is a tested assumption, rather than another paragraph or another vote for a preferred option.
Worked example: correct arithmetic, incorrect scope
Consider a hypothetical support automation proposal. Its draft assumes 1,200 tickets per month, four minutes saved per ticket and a staff-hour value of €20. Multiplication produces 80 hours, equivalent to €1,600. These invented figures illustrate the workflow; they are not a customer case study or observed Colay results.
The first answer recommends proceeding. A reviewer points out that only 40% of tickets may be eligible for automation. Holding the other assumptions constant changes the estimate to 32 hours, or €640. The original model could have calculated perfectly while supporting a poor decision because the coverage assumption was missing.
The next step is to establish that 40% figure using actual tickets. If it is also invented, the reviewer has merely substituted one unsupported premise for another. A useful synthesis specifies the required sample, alternative scenarios and the threshold for proceeding. Agreement does not resolve the missing observation; measurement does.
| Assumption | Hours per month | Equivalent at €20/hour |
|---|---|---|
| All 1,200 tickets | 80 | €1,600 |
| 40% of tickets | 32 | €640 |
Using Consensus to make assumptions visible
In Colay Consensus, participants discuss a task and a coordinator synthesizes the outcome. For this example, the user can explicitly request assumption checks, unit checks, alternative explanations and unresolved questions. These are suggested instructions, not a claim that the product automatically provides independent experts, verified sources or a completed audit.
Provide the same input table and distinguish measured values from estimates. Ask each objection to identify the row it depends on. Request a final answer that separates established facts, conditional conclusions and unknowns. Then open the underlying documents and reproduce the decisive calculation. Several models accepting an estimate does not turn it into a measured fact.
When an extra perspective adds little
If every participant receives an incomplete document, all may overlook the same exception. A question that embeds your preferred conclusion can also frame the entire discussion around defending it. A coordinator that smooths away objections may remove its most valuable output. Additional credits are best spent where another participant contributes a new, testable challenge.
A short translation of your own text or a draft invitation may only need ordinary proofreading. A recommendation that commits spending, changes a customer promise or depends on external facts deserves more scrutiny. Match the checking process to the consequences of an error and the ease of reversing it. More participants are a means of organizing scrutiny, not a reason to skip it.
Keep a decision record
Save the original question, decisive claim, supporting document or calculation, unresolved limitation and decision owner. A week later, this compact record will explain why the team proceeded and what new evidence should trigger a review. It also gives you something concrete to compare when testing a single model against Consensus on your own work.
Start with one memo whose weakest assumption could affect a real decision. Ask Consensus to identify that assumption, then verify it outside the discussion. Success may be a revised recommendation, a confirmed calculation or a decision to wait for one missing measurement. All three are more useful than confidence unsupported by evidence.
Sources and methodology
- Farquhar et al., Nature, 2024
Semantic entropy; published 19 June 2024.
- NIST AI 600-1, 2024
Generative AI risk profile, section 2.2.
Bring your next question to Colay
Choose a model, use Auto, or bring several perspectives together with Consensus.