Colay / Guides

AI confidence and human judgment: keeping ownership of a decision

A useful AI result should help a person understand a decision, not merely accept a finished document faster. The challenge is to preserve your own criteria, make objections actionable and establish whether a multi-perspective discussion actually improves the work you do.

External research results and illustrative calculations are not measurements of Colay performance.

What confidence research can tell us

A CHI 2025 study by Microsoft Research and Carnegie Mellon surveyed 319 workers describing 936 uses of AI. Higher confidence in AI was associated with less self-reported critical thinking. These are survey associations, not proof that AI causes a decline in intelligence. Lee et al., CHI 2025

A 2024 Nature Human Behaviour meta-analysis covered 106 experiments and 370 effects from 2020–2023 studies. Human–AI combinations performed worse, on average, than the better of either alone: Hedges’ g = −0.23, 95% CI −0.39 to −0.07. This standardized effect does not mean a 23% loss of accuracy. Vaccaro et al., Nature Human Behaviour, 2024

Our recommendation is to give human oversight a specific job. A person should define the goal, choose the criteria, check decisive evidence and explain why the conclusion is accepted. Simply rereading a polished answer and approving it is too vague a procedure to evaluate. The presence of a person is less informative than the checks that person actually performs.

Write your criteria before reading the answer

Before requesting analysis, record what you are choosing, which constraints are mandatory and what would change your view. For a software choice, that might mean a required integration, an acceptable learning period and the ability to export your data. This brief note does not need to contain the right answer. Its purpose is to preserve your criteria before an attractive external framing arrives.

After the discussion, compare its conclusion with the note. If a new criterion appears, identify where it came from: an actual requirement, a newly discovered limitation or an appealing description? Revising criteria is reasonable when you can explain the reason. The record helps distinguish learning from silently adopting the most recent confident opinion.

Worked example: a table that looks more complete than it is

Imagine a fictional team selecting software for weekly reports. It compares three options against five criteria. One AI answer assigns scores and announces a winner. Yet two criteria, the required export format and a specific integration, have not been verified. The table looks finished even though decisive input values remain unknown.

Mark those cells as unknown instead of assigning a neutral score. Ask the discussion to test the decision under alternative values: what if export is unavailable? Would the winner change if setup takes a week? Those questions produce a short product-verification plan. The three options and five criteria are part of this teaching example, not findings about Colay customers.

A person then tests the export with a sample file and the integration with a real workflow before updating the table. If the choice changes, the reason matters: new observations arrived. If it remains the same, the check still creates a defensible explanation for colleagues. The aim is to improve the basis of the choice, not to maximize the number of times humans override AI.

A Consensus request that leaves judgment visible

Colay Consensus has participants discuss a request and a coordinator synthesize the result. Try this instruction: “Check the criteria first. Identify unknown inputs. Show which unknown could change the choice. Propose the smallest useful verification.” You can also ask the final answer to preserve an alternative that remains reasonable under different assumptions. This is a suggested practice, not a measured product guarantee.

Do not ask the coordinator to close the question simply because most responses agree. Request the evidence-to-conclusion connection and keep material uncertainty visible. Then explain the decision in your own words without copying the summary. If that explanation depends entirely on saying that the models agreed, return to the inputs that should justify it.

Measure usefulness on your own work

Choose several recurring tasks and define an error before comparing methods. Give a single-model run and a discussion the same assignment and supporting material. Inspect concrete defects such as an unsupported feature claim, a wrong calculation or a missed mandatory constraint. Record review time and credit use too. Keep these observations separate from an impression that the answer sounds deeper.

Where practical, ask a colleague to assess anonymized outputs without knowing which mode produced each one. A small internal sample cannot establish universal product superiority, but it can reveal where discussion earns its additional time in your workflow. Pay particular attention to unanimous errors. They expose limitations that a successful demonstration might leave invisible.

Preserve the option to stop without a winner

Allow the conclusion that evidence is insufficient. If a format demands a winner under every condition, it encourages filling gaps. For a small reversible task, you might accept an assumption and check later. For a consequential choice, record the observation required before proceeding. The task owner should set that threshold rather than infer it from the answer’s tone.

For your next work decision, write criteria, request discussion, verify one decisive fact and explain the outcome yourself. This makes the human role observable. Consensus can provide an organized place for different arguments, while the person making the decision retains responsibility for the criteria, factual checks and eventual action.

Sources and methodology

  1. Lee et al., CHI 2025

    Microsoft Research and Carnegie Mellon; survey, not a causal experiment.

  2. Vaccaro et al., Nature Human Behaviour, 2024

    Published 28 October 2024; preregistered meta-analysis.

Bring your next question to Colay

Choose a model, use Auto, or bring several perspectives together with Consensus.

Open Colay