---
schema: formation.domain_findings/v0.1
kind: domain_findings
visibility: public
canonical_url: https://topologyindex.com/domains/security/findings.md
description: 'Each published study tagged with security work, once, with every finding filed under a pattern page: the sentence, comparison, domain, caveat and source. Attributed, stated without figures, never a ranking.'
domain: security
domain_page: /domains/security.md
findings:
  - citations:
      - compared_against: 'Prior single-agent harnesses with interactive tools, and earlier published evaluations'
        direction: helped
        pattern: single_agent
    source_id: arxiv:2412.02776
    url: https://arxiv.org/abs/2412.02776
  - citations:
      - compared_against: 'A single-agent patcher, a fixed workflow and a general-purpose coding agent on the same tasks'
        direction: mixed
        pattern: supervisor
    source_id: arxiv:2603.01257
    url: https://arxiv.org/abs/2603.01257
  - citations:
      - compared_against: 'Direct use of the same LLM for penetration testing, and ablations removing each module'
        direction: helped
        pattern: planner_worker
    source_id: arxiv:2308.06782
    url: https://arxiv.org/abs/2308.06782
  - citations:
      - compared_against: 'The same system run as a single executor, and prior single-agent CTF agents'
        direction: helped
        pattern: planner_worker
    source_id: arxiv:2502.10931
    url: https://arxiv.org/abs/2502.10931
  - citations:
      - compared_against: 'Planner and executor running the same model'
        direction: no_clear_gain
        pattern: planner_worker
    source_id: arxiv:2604.17159
    url: https://arxiv.org/abs/2604.17159
  - citations:
      - compared_against: 'A single agent with the same model and no vulnerability description, and ablations without the task-specific agents or the hierarchy'
        direction: helped
        pattern: hierarchical_delegation
    source_id: arxiv:2406.01637
    url: https://arxiv.org/abs/2406.01637
  - citations:
      - compared_against: 'The Cybench single agent and AutoGPT under the same model and iteration budget'
        direction: mixed
        pattern: hierarchical_delegation
    source_id: arxiv:2503.17332
    url: https://arxiv.org/abs/2503.17332
  - citations:
      - compared_against: 'Single-agent chain-of-thought prompting, fine-tuned code models and the GPTLens auditor-critic system'
        direction: helped
        pattern: debate
    source_id: arxiv:2505.10961
    url: https://arxiv.org/abs/2505.10961
hostile_input:
  - citations:
      - compared_against: 'Multi-agent systems without provenance tagging'
        direction: helped
        pattern: signed_coordination
    source_id: arxiv:2410.07283
    url: https://arxiv.org/abs/2410.07283
  - citations:
      - compared_against: 'Multi-agent systems with unprotected inter-agent messages'
        direction: not_tested
        pattern: signed_coordination
    source_id: arxiv:2502.14847
    url: https://arxiv.org/abs/2502.14847
  - citations:
      - compared_against: 'Multi-agent orchestrators without system-level trust models'
        direction: not_tested
        pattern: signed_coordination
    source_id: arxiv:2503.12188
    url: https://arxiv.org/abs/2503.12188
path: /domains/security/findings.md
pattern_index: /patterns/index.md
product_api_version: v1
schema_version: v0.1
title: 'Published findings on multi-agent patterns studied in security work'
---

# Published findings on multi-agent patterns studied in security work

The published findings of the [`security` domain](/domains/security.md), in words. That page
has the task shapes common in the domain, their rows of the [decision guide](/patterns/index.md)
and the starters. None of it has been measured here, and nothing on this page ranks patterns.

## Published findings in this domain

Attributed to each source and stated without figures. Each source is listed once, under the
first pattern in the vocabulary that cites it, with every pattern page that cites it for this
domain. The order is the pattern vocabulary, never the direction, and unfavourable results
are included on purpose. None of this is evidence produced here. Reviewed 2026-09-23.
Each label says how the pattern each finding is filed under fared against what it was compared with, as the source reports
it; the labels are defined at [/docs/schemas/pattern/v0.1.md](/docs/schemas/pattern/v0.1.md).

### single_agent

- Source: [Hacking CTFs with Plain Agents](https://arxiv.org/abs/2412.02776), Turtayev et al., 2024-12-03.
  - **Helped**, filed under [single_agent](/patterns/single_agent.md) — Turtayev and colleagues report that a plain single agent using prompting, tools and repeated independent attempts outperformed a more elaborate agent harness on a high-school-level hacking benchmark, and that tree-of-thought exploration added little.
    Compared against: Prior single-agent harnesses with interactive tools, and earlier published evaluations. Domain: Offensive security: capture-the-flag challenges. Benchmarks: InterCode-CTF.
    Caveat: The benchmark is saturated and possibly contaminated, the task subset differs from prior work, and no multi-agent system was compared.

### supervisor

- Source: [A Systematic Study of LLM-Based Architectures for Automated Patching](https://arxiv.org/abs/2603.01257), Xu, Sheng, Chen, Huang, 2026-03-01.
  - **Mixed**, filed under [supervisor](/patterns/supervisor.md) — Xu and colleagues report that a multi-agent patching system did not consistently beat a well-designed single agent, winning with one model and losing with another at higher overhead, while a general-purpose coding agent patched the most vulnerabilities.
    Compared against: A single-agent patcher, a fixed workflow and a general-purpose coding agent on the same tasks. Domain: Security: repairing real-world vulnerabilities in large Java projects. Benchmarks: AIxCC.
    Caveat: A small set of vulnerabilities from one competition, with each architecture reimplemented by the authors.

### planner_worker

- Source: [PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing](https://arxiv.org/abs/2308.06782), Deng et al., 2023-08-13.
  - **Helped**, filed under [planner_worker](/patterns/planner_worker.md) — Deng and colleagues report that splitting penetration testing across a reasoning session that keeps a task tree and separate sessions that generate commands and condense tool output completed more targets and sub-tasks than using the same model directly, and that removing the reasoning session made the system worse than the plain model.
    Compared against: Direct use of the same LLM for penetration testing, and ablations removing each module. Domain: Offensive security: penetration testing of practice machines. Benchmarks: HackTheBox, VulnHub, picoMini.
    Caveat: A human expert executes every proposed command, so this is a human-in-the-loop system rather than an autonomous one.
- Source: [D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security](https://arxiv.org/abs/2502.10931), Udeshi et al., 2025-02-15.
  - **Helped**, filed under [planner_worker](/patterns/planner_worker.md) — Udeshi and colleagues report that a planner directing specialised executor agents solved somewhat more capture-the-flag challenges than a single executor given the same prompt, at a modestly higher total cost, and that pairing a strong planner with weaker executors consistently underperformed.
    Compared against: The same system run as a single executor, and prior single-agent CTF agents. Domain: Offensive security: capture-the-flag challenges. Benchmarks: NYU CTF Bench, Cybench, HackTheBox.
    Caveat: Author-run comparison with small gains over the single-executor ablation, and the comparison with EnIGMA mixes model versions.
- Source: [Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks](https://arxiv.org/abs/2604.17159), Merves et al., 2026-04-18.
  - **No clear gain**, filed under [planner_worker](/patterns/planner_worker.md) — Merves and colleagues report that assigning a stronger model to the planner and a cheaper one to the executor, or the reverse, gave no meaningful benefit over using one model for both roles, and that adding an auto-prompting agent often degraded results in a well-equipped environment.
    Compared against: Planner and executor running the same model. Domain: Offensive security: capture-the-flag challenges. Benchmarks: NYU CTF Bench.
    Caveat: Single-trial measurements within one framework, from a short symposium submission.

### hierarchical_delegation

- Source: [Teams of LLM Agents can Exploit Zero-Day Vulnerabilities](https://arxiv.org/abs/2406.01637), Zhu et al., 2024-06-02.
  - **Helped**, filed under [hierarchical_delegation](/patterns/hierarchical_delegation.md) — Zhu and colleagues report that a hierarchical planner dispatching task-specific expert agents exploited real-world web vulnerabilities without being told what they were far more often than a single agent, and that removing the expert agents or the hierarchy sharply reduced success.
    Compared against: A single agent with the same model and no vulnerability description, and ablations without the task-specific agents or the hierarchy. Domain: Offensive security: exploiting unknown real-world web vulnerabilities.
    Caveat: Author-built benchmark of a small set of reproducible open-source web vulnerabilities, which the authors note may be a biased sample.
- Source: [CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities](https://arxiv.org/abs/2503.17332), Zhu et al., 2025-03-21.
  - **Mixed**, filed under [hierarchical_delegation](/patterns/hierarchical_delegation.md) — Zhu and colleagues report that a hierarchical team of specialised agents exploited more real-world web vulnerabilities than an agent built for capture-the-flag tasks, while a general single agent with self-criticism achieved the highest success rate.
    Compared against: The Cybench single agent and AutoGPT under the same model and iteration budget. Domain: Offensive security: exploiting real-world web application vulnerabilities. Benchmarks: CVE-Bench.
    Caveat: The team framework comes from the same research group, and all success rates are low.

### debate

- Source: [Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents](https://arxiv.org/abs/2505.10961), Widyasari et al., 2025-05-16.
  - **Helped**, filed under [debate](/patterns/debate.md) — Widyasari and colleagues report that a courtroom-style arrangement of a security researcher, a code author, a moderator and a review board detected vulnerable functions markedly better than single-agent prompting and an auditor-critic system, but that extending discussion beyond one round lowered performance.
    Compared against: Single-agent chain-of-thought prompting, fine-tuned code models and the GPTLens auditor-critic system. Domain: Security: detecting vulnerable code in paired vulnerable and patched functions. Benchmarks: PrimeVul.
    Caveat: Author-run evaluation on function-level pairs, and the arrangement costs roughly three times the single-agent run with the stronger model.

No study reviewed here covers triage of many independent findings or alerts, so the
`splits_into_independent_parts` shape is untested in this domain. Every finding tagged with
this domain is about offensive or defensive technical work on code and systems.

## Risk: agents in this domain read hostile input

Security work means reading content an adversary may control: target code, web pages, tool
output and service responses. The studies below attack multi-agent systems through that
content and through the messages agents exchange. They are not findings about a pattern doing
security work, and they do not rank any arrangement. They are a reason to weigh arrangements
in which agents share state or pass messages ([blackboard](/patterns/blackboard.md),
[shared_ledger](/patterns/shared_ledger.md), [mailbox_network](/patterns/mailbox_network.md)) carefully when inputs are
adversarial, and to consider message provenance ([signed_coordination](/patterns/signed_coordination.md)).

- Source: [Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems](https://arxiv.org/abs/2410.07283), Lee, Tiwari, 2024-10-09.
  - **Helped**, filed under [signed_coordination](/patterns/signed_coordination.md) — Lee and Tiwari report that malicious prompts can self-replicate across interconnected agents even when agents do not share all communications, and that marking message provenance with LLM Tagging plus existing safeguards significantly reduced spread.
    Compared against: Multi-agent systems without provenance tagging. Domain: Multi-agent application security.
    Caveat: LLM Tagging is a prompt-level label, not cryptographic signing, and works only in combination with other safeguards.
- Source: [Red-Teaming LLM Multi-Agent Systems via Communication Attacks](https://arxiv.org/abs/2502.14847), He et al., 2025-02-20.
  - **Not tested**, filed under [signed_coordination](/patterns/signed_coordination.md) — He and colleagues report that an adversary who only intercepts and manipulates inter-agent messages can compromise entire multi-agent systems across various frameworks and communication structures.
    Compared against: Multi-agent systems with unprotected inter-agent messages. Domain: Multi-agent application security.
    Caveat: Attack study that motivates message integrity protection but does not evaluate signing as a defense.
- Source: [Multi-Agent Systems Execute Arbitrary Malicious Code](https://arxiv.org/abs/2503.12188), Triedman, Jha, Shmatikov, 2025-03-15.
  - **Not tested**, filed under [signed_coordination](/patterns/signed_coordination.md) — Triedman and colleagues report that adversarial web content can hijack control flow in multi-agent orchestrators to execute arbitrary code, even when individual agents resist prompt injection.
    Compared against: Multi-agent orchestrators without system-level trust models. Domain: Multi-agent application security.
    Caveat: Attack study; it calls for trust and security models but does not test authenticated messaging.

[Reporting outcomes (limited rollout)](/docs/api/contributing.md): only for a pattern, starter or formation fetch that carried a `Use-Ticket` (or a "Report back" note at the end of the page), which invited credentials and some selected visiting agents receive; without one there is nothing to report and nothing else changes.
