Colay / Guides

Test Consensus against the work you actually need done

A useful evaluation asks whether a workflow improves your deliverables under constraints you can afford. It does not start by collecting impressive examples. Compare the same tasks, decide how success will be scored before seeing the answers, and preserve failures. That produces a decision you can revisit instead of a number that only looks precise.

External research results and illustrative calculations are not measurements of Colay performance.

Choose the decision before choosing the score

NIST AI RMF’s Measure function calls for documented test sets, metrics, uncertainty, and evaluation in conditions resembling deployment. It supplies a measurement framework, not a prescribed winner or a certification of an AI product. NIST AI RMF: Measure

Write down your adoption rule. For example: use the reviewed workflow for requirements analysis if it finds more material omissions within the same spending allowance and does not exceed an agreed waiting time. This is a proposed rule for your experiment. Decide what counts as material by listing concrete examples before any model produces an answer.

Separate factual correctness from completeness, usefulness, and style. A polished response that invents a constraint should fail a factual check even if reviewers prefer its tone. For open-ended tasks, use a short rubric with anchored examples and record disagreements between human reviewers. Agreement between judges is itself something to inspect, not assume.

Pair the tasks and control the budget

Build a sample from actual work: ordinary tasks, difficult cases, ambiguous requests, and examples with missing information. Keep a separate set for tuning prompts. Run each evaluation task through both workflows with the same source material. Conceal workflow names from scorers and randomize answer order so position and branding do not decide the comparison.

Use the same total spending envelope and the same completion deadline for each workflow across the task set. Include every generation, coordination step, retry, and tool call in the accounting. A single-agent baseline should have a reasonable prompt and access to the same relevant information. Giving one side extra attempts while counting only its final answer defeats the comparison.

Equal envelopes do not require wasting unused budget. Record actual spending, and predefine what happens when a workflow reaches its cap: an unfinished task should remain visible. You can also run a separate equal-quality comparison to estimate cost, but do not mix its result with the equal-budget experiment.

A small sample needs an uncertainty interval

NIST describes the Wilson interval for a binomial proportion. Our calculation for 95 successes in 100 independent, representative trials gives a 95% interval of 88.82%–97.85%, using z = 1.9599639845. These are illustrative observations, not Colay measurements. NIST: Confidence intervals for proportions

The interval concerns the success rate of the defined task population under the sampling assumptions. It is not a probability that a particular answer is correct. It also cannot repair a biased sample. Repeating nearly identical requests from one document may create dependence, and testing only easy drafts says little about complex analysis.

Do not announce that a 95% observed score proves a 95% reliable product. Record the date, model versions, prompts, task selection, exclusions, scoring rules, and budget. A changed model or a new task mix can make the old estimate a poor guide to the next month.

Read the paired outcomes, not just two percentages

Consider another explicitly invented result on the same 100 tasks: both workflows pass 89, only Consensus passes 6, only the single-agent workflow passes 1, and both fail 4. Totals are 95 and 90 successes. The paired difference is five tasks, but the most informative observations are the seven cases where the workflows disagree.

For this constructed table, a two-sided exact McNemar test conditions on those seven discordant pairs. Under equal win probability, its p-value is 2 × (1 + 7) ÷ 128 = 0.125. This would not meet a preselected 0.05 threshold. It does not prove equivalence; it shows why the apparent lead needs more evidence. Overlap of separate confidence intervals is not a replacement for a paired analysis.

Hypothetical paired results on 100 tasks
OutcomeTasks
Both pass89
Only Consensus passes6
Only single-agent passes1
Both fail4

Turn findings into a limited operating rule

Inspect every regression. A workflow that improves many low-impact drafts but introduces one serious factual failure may not meet your adoption rule. Expand promising categories with fresh tasks, and keep a holdout set that has not shaped prompts. Repeatedly inspecting results and stopping at the first attractive score can exaggerate apparent gains.

In Colay Consensus, agents discuss the request and a coordinating agent synthesizes a conclusion. Evaluate that final deliverable and the effort required to verify it. Record credits because Colay subscriptions have credits and limits. Publish only what your test supports: the sampled tasks, tested configuration, observed trade-offs, and uncertainty. None of this example establishes a measured Colay benchmark.

Sources and methodology

  1. NIST AI RMF: Measure

    Evaluation framework.

  2. NIST: Confidence intervals for proportions

    Wilson interval method.

Bring your next question to Colay

Choose a model, use Auto, or bring several perspectives together with Consensus.

Open Colay