Colay / Guides
Mixture-of-Agents: a synthesis must earn its value
Collecting several AI responses is straightforward; turning them into a coherent, checkable deliverable is harder. Mixture-of-Agents research provides a reason to take synthesis seriously, but its headline percentages are easy to misread. Here we examine the measurement, distinguish preference from factual accuracy and propose a practical way to assess the final output.
External research results and illustrative calculations are not measurements of Colay performance.
The result, the metric and the reference answer
Wang et al.’s 2024 MoA paper reports a 65.1% length-controlled win rate on AlpacaEval 2.0’s 805 instructions, versus 57.5% for GPT-4o (05/13). Both are compared with gpt-4-1106-preview. This measures AI-judge preference, not the share of correct facts or Colay performance. Wang et al., 2024
The difference is 7.6 percentage points on that metric. It does not establish workplace accuracy or the proportion of customers who will choose a product.
Why adjusting for length matters
Dubois et al. introduced length-controlled AlpacaEval to adjust automatic-judge preferences for differences in answer length. That adjustment does not turn the score into an audit of every factual statement. Dubois et al., Length-Controlled AlpacaEval, 2024
In your own comparison, a long synthesis may feel complete even when the additional detail does not improve the decision. Give both workflows the same space limit and required content. Then score usefulness, factual correctness and readability separately. A single combined score can hide an outcome in which presentation improves while the basis of the recommendation becomes weaker.
Preserve the conditions behind each recommendation
Consider a hypothetical example. One response proposes a two-week implementation if the data is ready; another estimates six weeks if cleaning is necessary. A synthesis saying implementation takes four weeks sounds balanced, but represents neither scenario. A better answer preserves the branch and asks about data readiness. Averaging the statements destroys useful information.
For this reason, we suggest treating a claim and its applicability conditions as one unit during synthesis. Each recommendation needs a basis, a boundary and identifiable input evidence. You can ask a coordinator for a conflict table, then check that compression has not removed an inconvenient condition. Good editing makes differences easier to understand instead of smoothing them away.
Audit what disappeared between drafts and final text
Suppose a supplier brief contains 12 mandatory conditions. In a hypothetical review, the candidate answers collectively cover all 12, but the final synthesis preserves only nine. Required-condition coverage is 75%, even if the final document reads better than any draft. This is a separate measure and has no connection to the AlpacaEval percentage.
List the mandatory conditions from the brief, locate each in the final text and record omissions. Next, inspect claims introduced only during synthesis. A new number or promise needs supporting evidence. If no support exists, label it as an assumption or remove it. This tests whether useful content survived the workflow, rather than measuring how polished the result feels.
| Check | Example outcome | Action |
|---|---|---|
| Required conditions | 9 of 12 retained | Restore three omissions |
| New numerical claims | 2 unsupported | Verify or remove |
| Disputed assumptions | 1 concealed | Show the decision branch |
Use a strong single-model comparison
To test the value of multiple models in your work, give the single-model workflow the same prepared brief, sources and grading criteria. Comparing an organized process with a careless one-line request confounds task preparation with architecture. Record the effort required to prepare inputs and review outputs, as well as model usage and waiting time.
Try both a short factual question and a longer decision memo with alternatives. They need different success criteria. The former requires support for a specific claim; the latter also needs coverage of material conditions and a clear explanation of the choice. If improvement appears only in decision memos, keep the practical recommendation limited to that type of work.
Apply the idea to Consensus without borrowing its score
Colay Consensus also ends with a coordinator’s synthesis, but that does not make it identical to the researched MoA configuration. This article contains no published Colay benchmark. Treat the external result as a reason to test the workflow on your material, not as a ready-made quality score for every Consensus run.
Request a conclusion with four parts: recommendation, evidence, unresolved disagreement and the next checkable step. Provide a space limit and a list of mandatory conditions. Review the result against the brief and its sources. The workflow earns a useful role when you can identify what it preserved or added and how much human work still remains.
Sources and methodology
- Wang et al., 2024
Table 2a; section 3.1.
- Dubois et al., Length-Controlled AlpacaEval, 2024
The metric adjusts model-judge preferences for response length; it is not a factual-accuracy score.
Bring your next question to Colay
Choose a model, use Auto, or bring several perspectives together with Consensus.