Colay / Guides

Compare AI models on the prompts your product actually needs

Before building an evaluation pipeline, you can run a small, structured comparison in Colay. Give available agents the same permitted inputs, inspect their separate answers and record which requirements they meet. This is a practical first screening without writing code or supplying provider API keys; it is not an automated benchmark runner or proof that the model will behave identically in your product.

Turn a feature into testable cases

Start with the output your user needs, not a model leaderboard. If the feature extracts action items, success means preserving the right owner, task and deadline while leaving unknown fields unknown. An elegant summary that invents a deadline should fail that case.

As an illustrative starting set, prepare 12 permitted examples: 4 ordinary inputs, 4 ambiguous inputs and 4 edge cases. These numbers are a manageable exercise, not a statistically validated sample size. Remove personal and confidential information you are not authorized to share. Keep an expected result or explicit acceptance rule beside each case.

Run a comparison you can interpret

Anthropic's evaluation guide distinguishes a task from repeated trials of that task because outputs can vary. This article applies that basic distinction to a manual screening exercise. A few clean answers cannot establish a production failure rate.

Avoid giving one candidate a corrected prompt after another has already failed. If the prompt needs improvement, create a new version and rerun every candidate on it. Otherwise you are comparing instructions as well as models.

  1. Choose the candidate agents explicitly from the current catalog. Do not use Auto when you need to know which candidate answered.
  2. Keep the brief, source material and requested format identical. Use Ask separately for the first answers and retain their labels.
  3. Record the date, displayed agent name, prompt version and any visible settings. Note missing responses separately from incorrect responses.
  4. Review each result against the rules before comparing prose quality. Repeat important or ambiguous cases and retain every attempt, including failures.

Example: action items from a meeting note

Fictional input: “Mira will send the draft on Tuesday. The team still needs an owner for the pricing review.” The desired result contains one assigned task and one unassigned task. It does not assign both to Mira or invent a date for the pricing review.

The useful artifact is a case sheet, not a screenshot of the nicest answer. Attach the exact returned text to each row so a teammate can challenge your score. Keep formatting failure separate from factual failure: a fixable table layout and an invented commitment have different consequences.

Manual scoring sheet for this fictional case
CheckPass conditionFailure to record
OwnerMira owns only the draftPricing review assigned without evidence
DeadlineTuesday belongs to the draftInvented pricing deadline
Unknown fieldPricing owner remains unknownMissing information silently filled
UsabilityBoth tasks appear in the requested structureTask omitted or format unusable

A prompt for the screening run

Keep evaluation instructions in a separate checklist that you apply to the answer. You can ask another agent to identify possible errors, but its verdict is another output to inspect. Consensus can help summarize the observed tradeoffs after the first comparison; explicitly supply the labelled results and your scores.

Extract action items from the supplied meeting note: [permitted note]. Return a table with task, owner, deadline and supporting excerpt. Use only the supplied text. Write unknown for an unspecified owner or deadline. Do not turn a suggestion into an agreed commitment. Preserve unresolved questions separately. Input ID: [case ID].

Use the shortlist for the next test

Summarize the cases each candidate handles, its critical failures and the questions still unanswered. Retain a candidate only if its behavior fits the feature's minimum requirements. If none qualifies, narrow the feature or improve the supplied context before selecting a winner.

An integration trial must separately check the actual model endpoint, tool access, system instructions, latency, usage cost and failure handling. Colay's credits are not a quote for that API workload. Start in Colay with one real, de-identified case you can grade yourself, then expand only when the comparison changes your shortlist.

Questions, answered

Can I compare models without API keys?

Yes, for the end-user comparison in Colay. Building those models into your own product is a separate integration with its own access, terms and costs.

Is this a statistically reliable benchmark?

No. A small manual exercise reveals concrete behaviors and failures. Broader claims require representative cases, repeated trials and an evaluation design appropriate to the decision.

Should I use Consensus to choose the winner?

First score separate outputs against your acceptance rules. Use synthesis to organize the evidence, keeping your original results and human decision visible.

Sources and methodology

  1. Anthropic — Demystifying evals for AI agents

    Primary engineering guidance for the distinction between tasks and repeated trials. It does not evaluate Colay or validate the example sample size.

Bring your next question to Colay

Choose a model, use Auto, or bring several perspectives together with Consensus.

Compare my prompts