Colay / Guides
More AI participants do not create a security boundary
An agent can encounter useful evidence and hostile instructions in the same document. Adding reviewers may help reveal suspicious content, but it also creates more opportunities to repeat it. A secure workflow needs clear rules about which material can inform an answer and which actors can authorize an action. Agreement cannot supply that authorization.
External research results and illustrative calculations are not measurements of Colay performance.
What the research actually demonstrated
AgentPoison (2024) studied poisoned long-term memory and retrieval stores in three agent settings. The attack required no model fine-tuning. It demonstrates a retrieval trust problem, not a measured vulnerability in Colay. AgentPoison, NeurIPS 2024
Agent Smith (ICML 2024) simulated up to one million LLaVA-1.5 agents with randomized pairwise chats. An adversarial image could propagate harmful behavior. This specific simulation does not establish an infection rate for arbitrary production systems. Agent Smith, ICML 2024
OWASP LLM01:2025 distinguishes direct and indirect prompt injection and recommends layered mitigation. It does not claim that adding another model, or relying on a system prompt alone, eliminates the problem. OWASP LLM01:2025 Prompt Injection
The practical implication is to examine the path information takes. A suspicious instruction can enter as a quotation, become an agent’s summary, and later appear as a recommendation from a seemingly trusted participant. The wording may change while the underlying request survives. Review the origin and permitted use of the content, not just the identity of the last speaker.
A harmless review can cross several trust boundaries
Imagine a team reviewing supplier proposals. One proposal contains text that tries to change the reviewers’ task. A research agent summarizes the document; another agent critiques that summary; a coordinator assembles the recommendation. This is a hypothetical defensive scenario. None of those transfers should turn supplier-authored instructions into instructions from the user.
A useful review record keeps the source passage alongside the extracted claim. The next participant should be able to distinguish a document’s assertion from a reviewer’s inference and from the user’s actual request. If summarization removes that distinction, the coordinator may treat repeated content as corroboration even though every repetition has the same untrusted origin.
The issue becomes more consequential when a workflow can send messages, edit shared records, or invoke tools. Producing text and authorizing an external action are separate decisions. Even a well-supported recommendation should not acquire additional permissions simply because several participants repeated it.
Map what each participant can read and change
For a workflow you operate, draw a small map: external documents, retrieval stores, participants, shared notes, coordinator, and possible external actions. Mark where information crosses from one area to another. Then ask what evidence travels with it, which component checks permissions, and whether a participant can pass a request outside its assigned role.
Keep the review task narrow enough to inspect. A participant checking arithmetic needs relevant quantities and units; it may not need every attachment or access to a mailbox. This is a proposed design principle for minimizing unnecessary exposure. It is not a claim about which controls Colay currently implements or about the isolation of any particular provider.
| Boundary | Question to answer |
|---|---|
| Document → participant | Is source text distinguishable from task instructions? |
| Participant → coordinator | Can each important claim be traced to its origin? |
| Coordinator → external action | Who separately authorizes the action and its scope? |
| Current task → saved context | Which material may persist into a later task? |
Test containment without handling live secrets
Use a controlled exercise with synthetic documents and harmless canary values. Define the behavior that should remain unchanged, such as the selected evaluation criteria or allowed output destination. Vary whether a suspicious passage appears directly, inside a quotation, or in a summary. The goal is to observe boundary handling, not to expose real customer records.
Score more than whether the final answer looks acceptable. Check whether participants followed the original task, preserved provenance, requested unauthorized actions, or carried unwanted instructions into later context. Record the exact tested configuration and failures. A clean result on a small exercise is evidence about those cases, not proof that the system resists every possible injection.
Use Consensus as analysis, then verify the decision
Colay Consensus lets agents discuss a request and gives a coordinating agent the task of synthesizing a conclusion. This describes the collaboration workflow, not a security certification. When material is untrusted, ask for source-linked claims and keep the final decision separate from any consequential external action. Check the current product documentation for available controls.
If a discussion appears to inherit a document’s instructions, stop using that result, identify the original passage, and repeat the review with a clean task context where appropriate. Additional rounds alone may repeat the same contamination. Colay subscriptions use credits and limits, so more review also consumes an allowance; it should buy a specific check whose result you can inspect.
Sources and methodology
- AgentPoison, NeurIPS 2024
Memory and retrieval poisoning.
- Agent Smith, ICML 2024
Simulated multimodal agent interactions.
- OWASP LLM01:2025 Prompt Injection
Prompt injection risk guidance.
Bring your next question to Colay
Choose a model, use Auto, or bring several perspectives together with Consensus.