---
schema: formation.domain_findings/v0.1
kind: domain_findings
visibility: public
canonical_url: https://topologyindex.com/domains/coding/findings.md
description: 'Each published study tagged with coding, once, with every finding filed under a pattern page: the sentence, comparison, domain, caveat and source. Attributed, stated without figures, never a ranking.'
domain: coding
domain_page: /domains/coding.md
findings:
  - citations:
      - compared_against: 'Multi-agent debate methods'
        direction: no_clear_gain
        pattern: single_agent
      - compared_against: 'Single-agent chain-of-thought and self-consistency'
        direction: no_clear_gain
        pattern: debate
    source_id: arxiv:2502.08788
    url: https://arxiv.org/abs/2502.08788
  - citations:
      - compared_against: 'Open-source autonomous software engineering agents'
        direction: helped
        pattern: single_agent
    source_id: arxiv:2407.01489
    url: https://arxiv.org/abs/2407.01489
  - citations:
      - compared_against: 'Complex state-of-the-art agent designs'
        direction: no_clear_gain
        pattern: single_agent
    source_id: arxiv:2407.01502
    url: https://arxiv.org/abs/2407.01502
  - citations:
      - compared_against: 'Single LLM call and more elaborate prompting or multi-agent methods'
        direction: helped
        pattern: fan_out
    source_id: arxiv:2402.05120
    url: https://arxiv.org/abs/2402.05120
  - citations:
      - compared_against: 'Single-sample attempts'
        direction: mixed
        pattern: fan_out
    source_id: arxiv:2407.21787
    url: https://arxiv.org/abs/2407.21787
  - citations:
      - compared_against: 'A single call to a stronger model'
        direction: hurt
        pattern: fan_out
    source_id: arxiv:2411.17501
    url: https://arxiv.org/abs/2411.17501
  - citations:
      - compared_against: 'Sequential Chain-of-Agents over the same chunks'
        direction: hurt
        pattern: map_reduce
    source_id: arxiv:2406.02818
    url: https://arxiv.org/abs/2406.02818
  - citations:
      - compared_against: 'Single-agent systems and centralized, decentralized and hybrid multi-agent architectures'
        direction: mixed
        pattern: independent_workers
      - compared_against: 'Single-agent systems and independent, decentralized and hybrid multi-agent architectures'
        direction: mixed
        pattern: supervisor
    source_id: arxiv:2512.08296
    url: https://arxiv.org/abs/2512.08296
  - citations:
      - compared_against: 'Oracle selection, random selection and single trajectories'
        direction: mixed
        pattern: independent_workers
    source_id: arxiv:2501.14723
    url: https://arxiv.org/abs/2501.14723
  - citations:
      - compared_against: 'Smaller agent networks and regular topologies such as chains and meshes'
        direction: mixed
        pattern: lane_swarm
    source_id: arxiv:2406.07155
    url: https://arxiv.org/abs/2406.07155
  - citations:
      - compared_against: 'Single-agent solo setups'
        direction: mixed
        pattern: lane_swarm
      - compared_against: 'A single agent'
        direction: mixed
        pattern: dynamic_spawning
    source_id: arxiv:2308.10848
    url: https://arxiv.org/abs/2308.10848
  - citations:
      - compared_against: 'Published state-of-the-art agent systems on each benchmark'
        direction: no_clear_gain
        pattern: supervisor
    source_id: arxiv:2411.04468
    url: https://arxiv.org/abs/2411.04468
  - citations:
      - compared_against: 'Expected task success of the same frameworks'
        direction: hurt
        pattern: supervisor
      - compared_against: 'The unmodified ChatDev configuration'
        direction: no_clear_gain
        pattern: hierarchical_delegation
      - compared_against: 'Expectations of benefit from multi-agent frameworks with reviewer or verifier roles'
        direction: no_clear_gain
        pattern: implement_review
      - compared_against: 'Single-agent and simpler baselines on popular benchmarks'
        direction: no_clear_gain
        pattern: null
    source_id: arxiv:2503.13657
    url: https://arxiv.org/abs/2503.13657
  - citations:
      - compared_against: 'A single-agent patcher, a fixed workflow and a general-purpose coding agent on the same tasks'
        direction: mixed
        pattern: supervisor
    source_id: arxiv:2603.01257
    url: https://arxiv.org/abs/2603.01257
  - citations:
      - compared_against: 'The same models editing code on their own'
        direction: helped
        pattern: planner_worker
    source_id: web:aider.chat/2024/09/26/architect.html
    url: https://aider.chat/2024/09/26/architect.html
  - citations:
      - compared_against: 'GPT-Engineer, a single-agent approach, and the MetaGPT multi-agent framework'
        direction: helped
        pattern: role_pipeline
    source_id: arxiv:2307.07924
    url: https://arxiv.org/abs/2307.07924
  - citations:
      - compared_against: 'Direct, chain-of-thought, self-planning, analogical and Reflexion prompting, and the Self-collaboration and AlphaCodium frameworks'
        direction: helped
        pattern: role_pipeline
    source_id: arxiv:2405.11403
    url: https://arxiv.org/abs/2405.11403
  - citations:
      - compared_against: 'Flat and hierarchical multi-agent structures under the same injected faults'
        direction: hurt
        pattern: role_pipeline
      - compared_against: 'Linear and flat multi-agent structures'
        direction: helped
        pattern: hierarchical_delegation
    source_id: arxiv:2408.00989
    url: https://arxiv.org/abs/2408.00989
  - citations:
      - compared_against: 'Chat-based multi-agent frameworks such as ChatDev and AgentVerse, and single models'
        direction: helped
        pattern: hierarchical_delegation
    source_id: arxiv:2308.00352
    url: https://arxiv.org/abs/2308.00352
  - citations:
      - compared_against: 'Agents working without persistent progress artifacts'
        direction: mixed
        pattern: shared_ledger
      - compared_against: 'A coding agent looping across context windows with compaction only'
        direction: helped
        pattern: successor_handoff
    source_id: web:anthropic.com/engineering/effective-harnesses-for-long-running-agents
    url: https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents
  - citations:
      - compared_against: 'Chain, tree, star, complete, layered and random topologies and frameworks such as AutoGen and GPTSwarm'
        direction: helped
        pattern: mailbox_network
    source_id: arxiv:2410.02506
    url: https://arxiv.org/abs/2410.02506
  - citations:
      - compared_against: 'Sampling more initial programs at equal budget'
        direction: mixed
        pattern: implement_review
    source_id: arxiv:2306.09896
    url: https://arxiv.org/abs/2306.09896
  - citations:
      - compared_against: 'Human contractors reviewing code without assistance'
        direction: helped
        pattern: implement_review
    source_id: arxiv:2407.00215
    url: https://arxiv.org/abs/2407.00215
  - citations:
      - compared_against: 'Single-model code generation and prompt-engineering enhancement methods'
        direction: helped
        pattern: implement_review
    source_id: arxiv:2312.13010
    url: https://arxiv.org/abs/2312.13010
  - citations:
      - compared_against: 'The same base model acting as a single agent'
        direction: helped
        pattern: implement_review
    source_id: arxiv:2304.07590
    url: https://arxiv.org/abs/2304.07590
  - citations:
      - compared_against: 'One-step generation with the same model'
        direction: helped
        pattern: critic_loop
    source_id: arxiv:2303.17651
    url: https://arxiv.org/abs/2303.17651
  - citations:
      - compared_against: 'The same agent without reflection'
        direction: helped
        pattern: critic_loop
    source_id: arxiv:2303.11366
    url: https://arxiv.org/abs/2303.11366
  - citations:
      - compared_against: 'Aggregating several outputs of the single best model'
        direction: mixed
        pattern: council
    source_id: arxiv:2502.00674
    url: https://arxiv.org/abs/2502.00674
  - citations:
      - compared_against: 'A single-threaded linear agent sharing full context'
        direction: hurt
        pattern: dynamic_spawning
    source_id: web:cognition.com/blog/dont-build-multi-agents
    url: https://cognition.com/blog/dont-build-multi-agents
  - citations:
      - compared_against: 'Prior multi-agent routing and system design methods'
        direction: helped
        pattern: adaptive_routing
    source_id: arxiv:2502.11133
    url: https://arxiv.org/abs/2502.11133
  - citations:
      - compared_against: 'ReAct, Reflexion, and the tree-search methods Tree of Thoughts and RAP'
        direction: helped
        pattern: tree_search
    source_id: arxiv:2310.04406
    url: https://arxiv.org/abs/2310.04406
  - citations:
      - compared_against: 'The same open-source software agent without tree search'
        direction: helped
        pattern: tree_search
    source_id: arxiv:2410.20285
    url: https://arxiv.org/abs/2410.20285
  - citations:
      - compared_against: 'State-of-the-art hand-designed agents'
        direction: helped
        pattern: architecture_search
    source_id: arxiv:2408.08435
    url: https://arxiv.org/abs/2408.08435
  - citations:
      - compared_against: 'Manually designed workflows and prior automated methods'
        direction: helped
        pattern: architecture_search
    source_id: arxiv:2410.10762
    url: https://arxiv.org/abs/2410.10762
  - citations:
      - compared_against: 'Handcrafted and automated multi-agent systems'
        direction: helped
        pattern: architecture_search
    source_id: arxiv:2502.04180
    url: https://arxiv.org/abs/2502.04180
  - citations:
      - compared_against: 'Unoptimized topologies and agent-scaling strategies such as self-consistency and debate'
        direction: mixed
        pattern: architecture_search
    source_id: arxiv:2502.02533
    url: https://arxiv.org/abs/2502.02533
path: /domains/coding/findings.md
pattern_index: /patterns/index.md
product_api_version: v1
schema_version: v0.1
title: 'Published findings on multi-agent patterns studied in coding'
---

# Published findings on multi-agent patterns studied in coding

The published findings of the [`coding` domain](/domains/coding.md), in words. That page
has the task shapes common in the domain, their rows of the [decision guide](/patterns/index.md)
and the starters. None of it has been measured here, and nothing on this page ranks patterns.

## Published findings in this domain

Attributed to each source and stated without figures. Each source is listed once, under the
first pattern in the vocabulary that cites it, with every pattern page that cites it for this
domain. The order is the pattern vocabulary, never the direction, and unfavourable results
are included on purpose. None of this is evidence produced here. Reviewed 2026-09-23.
Each label says how the pattern each finding is filed under fared against what it was compared with, as the source reports
it; the labels are defined at [/docs/schemas/pattern/v0.1.md](/docs/schemas/pattern/v0.1.md).

### single_agent

- Source: [Stop Overvaluing Multi-Agent Debate: We Must Rethink Evaluation and Embrace Model Heterogeneity](https://arxiv.org/abs/2502.08788), Zhang et al., 2025-02-12.
  - **No clear gain**, filed under [single_agent](/patterns/single_agent.md) — The authors report that multi-agent debate methods often failed to beat simple single-agent baselines such as chain-of-thought and self-consistency, even while using considerably more inference compute.
    Compared against: Multi-agent debate methods. Domain: Reasoning, knowledge and coding question answering.
    Caveat: The study covers a fixed set of debate methods and base models, and the authors find that mixing heterogeneous models can recover some gains.
  - **No clear gain**, filed under [debate](/patterns/debate.md) — The authors report that multi-agent debate often failed to outperform chain-of-thought and self-consistency despite using more inference compute, and that using heterogeneous models improved debate frameworks.
    Compared against: Single-agent chain-of-thought and self-consistency. Domain: Reasoning, knowledge and coding question answering.
    Caveat: Results are for the debate methods and base models evaluated, and heterogeneity was proposed as a remedy rather than exhaustively tested.
- Source: [Agentless: Demystifying LLM-based Software Engineering Agents](https://arxiv.org/abs/2407.01489), Xia et al., 2024-07-01.
  - **Helped**, filed under [single_agent](/patterns/single_agent.md) — The authors report that a simple fixed three-phase pipeline of localization, repair and patch validation outperformed all existing open-source autonomous software agents on SWE-bench Lite at low cost.
    Compared against: Open-source autonomous software engineering agents. Domain: Repository-level bug fixing. Benchmarks: SWE-bench Lite.
    Caveat: The comparison is against agents available at the time and on a single benchmark family, and later agent systems have moved the frontier.
- Source: [AI Agents That Matter](https://arxiv.org/abs/2407.01502), Kapoor et al., 2024-07-01.
  - **No clear gain**, filed under [single_agent](/patterns/single_agent.md) — The authors argue that state-of-the-art agents are often needlessly complex and costly, and that ignoring cost has led the community to mistaken conclusions about where accuracy gains come from.
    Compared against: Complex state-of-the-art agent designs. Domain: Agent benchmarking practice, including code generation. Benchmarks: HumanEval.
    Caveat: This is primarily a methodological critique of benchmarking rather than a broad controlled study of multi-agent topologies.

### fan_out

- Source: [More Agents Is All You Need (Li et al.)](https://arxiv.org/abs/2402.05120), Junyou Li et al., 2024-02-03.
  - **Helped**, filed under [fan_out](/patterns/fan_out.md) — The authors report that a simple sample-and-vote method lets performance scale with the number of agents instantiated, with larger gains on harder tasks relative to the model.
    Compared against: Single LLM call and more elaborate prompting or multi-agent methods. Domain: Reasoning and code generation. Benchmarks: GSM8K, MATH, MMLU, Chess State Tracking, HumanEval.
    Caveat: The authors note token cost grows in proportion to agent count and gains taper off at the highest difficulty levels they tested.
- Source: [Large Language Monkeys: scaling inference compute with repeated sampling (Brown et al.)](https://arxiv.org/abs/2407.21787), Bradley Brown et al., 2024-07-31.
  - **Mixed**, filed under [fan_out](/patterns/fan_out.md) — The authors report that coverage from repeated sampling keeps rising with sample count and converts into real gains where answers can be automatically verified, but majority voting and reward models plateau and fail to keep pace where no verifier exists.
    Compared against: Single-sample attempts. Domain: Math, formal proofs, competitive programming and software engineering. Benchmarks: GSM8K, MATH, MiniF2F, CodeContests, SWE-bench Lite.
    Caveat: Most of the upside is coverage, which only becomes accuracy when a reliable automatic verifier can pick the correct sample.
- Source: [The Limits of Inference Scaling Through Resampling (Stroebl et al.)](https://arxiv.org/abs/2411.17501), Benedikt Stroebl et al., 2024-11-26.
  - **Hurt**, filed under [fan_out](/patterns/fan_out.md) — The authors report that when verifiers such as unit tests are imperfect, resampling cannot remove false positives, so weaker models cannot reach a strong model's single-attempt accuracy and the optimal number of attempts is small.
    Compared against: A single call to a stronger model. Domain: Code generation with unit-test verification. Benchmarks: HumanEval, HumanEval+, MBPP, MBPP+.
    Caveat: The study is limited to coding benchmarks where the gap between weak and full test suites can be measured.

### map_reduce

- Source: [Chain of Agents: LLMs collaborating on long-context tasks (Zhang et al.)](https://arxiv.org/abs/2406.02818), Yusen Zhang et al., 2024-06-04.
  - **Hurt**, filed under [map_reduce](/patterns/map_reduce.md) — The authors report that their sequential chain of worker agents outperformed parallel merge-by-vote and hierarchical worker-to-manager baselines on every dataset tested, attributing the gap to parallel workers being unable to communicate.
    Compared against: Sequential Chain-of-Agents over the same chunks. Domain: Long-context question answering, summarization and code completion. Benchmarks: HotpotQA, MuSiQue, NarrativeQA, Qasper, QuALITY, QMSum, GovReport, RepoBench-P.
    Caveat: The parallel baselines were built by the authors of the competing sequential method, so they may not be tuned as strongly as dedicated map-reduce systems.

### independent_workers

- Source: [Towards a Science of Scaling Agent Systems](https://arxiv.org/abs/2512.08296), Kim et al., 2025-12-09.
  - **Mixed**, filed under [independent_workers](/patterns/independent_workers.md) — The authors report that independent multi-agent setups without centralized verification propagate and amplify errors far more than centralized coordination, and that multi-agent gains vanish or reverse once the single-agent baseline is already strong.
    Compared against: Single-agent systems and centralized, decentralized and hybrid multi-agent architectures. Domain: Agentic web browsing, finance, planning, workplace, software engineering and terminal tasks. Benchmarks: BrowseComp-Plus, Finance Agent, PlanCraft, WorkBench, SWE-bench Verified, Terminal-Bench.
    Caveat: The fitted predictive model explains only a modest share of performance variance, so the patterns are tendencies rather than guarantees.
  - **Mixed**, filed under [supervisor](/patterns/supervisor.md) — The authors report that centralized orchestration strongly helped on decomposable financial reasoning but degraded performance on sequential planning, while containing error amplification far better than independent agents at a large token overhead.
    Compared against: Single-agent systems and independent, decentralized and hybrid multi-agent architectures. Domain: Agentic benchmarks spanning finance, web browsing, planning, workplace, software and terminal tasks. Benchmarks: Finance-Agent, PlanCraft, BrowseComp-Plus, Workbench, SWE-bench Verified, Terminal-Bench.
    Caveat: Results depend on the chosen model families and benchmark set, and the fitted predictive model explains only part of the variance.
- Source: [CodeMonkeys: scaling test-time compute for software engineering (Ehrlich et al.)](https://arxiv.org/abs/2501.14723), Ryan Ehrlich et al., 2025-01-24.
  - **Mixed**, filed under [independent_workers](/patterns/independent_workers.md) — The authors report that running many independent multi-turn editing trajectories per issue and then selecting among them raises resolved-issue rates, but their selection step recovers only part of the gap between random pick and an oracle choice.
    Compared against: Oracle selection, random selection and single trajectories. Domain: Resolving real GitHub issues. Benchmarks: SWE-bench Verified.
    Caveat: The approach is expensive and its final score is capped by selection quality rather than by how many candidates are generated.

### lane_swarm

- Source: [Scaling LLM-based multi-agent collaboration, MacNet (Qian et al.)](https://arxiv.org/abs/2406.07155), Chen Qian et al., 2024-06-11.
  - **Mixed**, filed under [lane_swarm](/patterns/lane_swarm.md) — The authors report that organizing up to over a thousand agents in a directed acyclic graph yields performance that grows logistically with agent count, with irregular topologies outperforming regular ones.
    Compared against: Smaller agent networks and regular topologies such as chains and meshes. Domain: Reasoning, code generation, software development and constrained text generation. Benchmarks: MMLU, HumanEval, SRDD, CommonGen-Hard.
    Caveat: The authors report most topologies saturate at around a hundred agents, that dense interaction can overload agents, and that context cost grows quadratically without their memory control.
- Source: [AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors](https://arxiv.org/abs/2308.10848), Chen et al., 2023-08-21.
  - **Mixed**, filed under [lane_swarm](/patterns/lane_swarm.md) — The authors report that dynamically composed agent groups can outperform a single agent, but also document cases where group discussion hurt a weaker model and negative emergent behaviors such as destructive actions.
    Compared against: Single-agent solo setups. Domain: Reasoning, coding, tool use and embodied Minecraft tasks. Benchmarks: FED, Commongen-Challenge, MGSM, Logic Grid Puzzles, HumanEval.
    Caveat: Groups in these experiments are small, so the results say little about large swarms.
  - **Mixed**, filed under [dynamic_spawning](/patterns/dynamic_spawning.md) — AgentVerse reports that dynamically adjusting group composition by recruiting expert agents lets multi-agent groups outperform a single agent, while also documenting negative emergent social behaviors.
    Compared against: A single agent. Domain: Text understanding, reasoning, coding, tool use and embodied tasks.
    Caveat: Author-run evaluation of the proposed framework; the negative behaviors are described qualitatively.

### supervisor

- Source: [Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks](https://arxiv.org/abs/2411.04468), Fourney et al. (Microsoft Research), 2024-11-07.
  - **No clear gain**, filed under [supervisor](/patterns/supervisor.md) — The authors report that their orchestrator-led team of specialist agents achieved performance statistically comparable to, not better than, state-of-the-art systems on general agentic benchmarks, and trailed the top entries on one web benchmark.
    Compared against: Published state-of-the-art agent systems on each benchmark. Domain: Generalist web, file and coding tasks. Benchmarks: GAIA, AssistantBench, WebArena.
    Caveat: First-party evaluation by the system's builders against heterogeneous published baselines rather than matched single-agent controls.
- Source: [Why Do Multi-Agent LLM Systems Fail?](https://arxiv.org/abs/2503.13657), Cemri et al., 2025-03-17.
  - **Hurt**, filed under [supervisor](/patterns/supervisor.md) — The authors report that popular multi-agent frameworks, including orchestrator-led ones, fail often, with failures clustering into system design, inter-agent misalignment and task verification problems rather than only underlying model weakness.
    Compared against: Expected task success of the same frameworks. Domain: Coding, math and general agent tasks across several open-source multi-agent frameworks. Benchmarks: MAST-Data.
    Caveat: A failure taxonomy built from annotated traces, not a controlled comparison of topologies against a single-agent baseline.
  - **No clear gain**, filed under [hierarchical_delegation](/patterns/hierarchical_delegation.md) — The authors report that improving role specifications in the role-hierarchy framework ChatDev produced only a modest gain in task success, and conclude that many failures stem from organizational design and coordination rather than individual agent ability.
    Compared against: The unmodified ChatDev configuration. Domain: Software development tasks. Benchmarks: MAST-Data.
    Caveat: A single intervention case study rather than a systematic comparison of hierarchy depths.
  - **No clear gain**, filed under [implement_review](/patterns/implement_review.md) — Across annotated traces from popular multi-agent frameworks, the authors identify task verification failures, such as missing or incorrect checking, as one of the main categories of multi-agent breakdowns, alongside design issues and inter-agent misalignment.
    Compared against: Expectations of benefit from multi-agent frameworks with reviewer or verifier roles. Domain: Coding, math and general agent tasks.
    Caveat: This is a failure taxonomy rather than a controlled comparison, so it shows how reviewer roles fail but not how often they help.
  - **No clear gain**, filed under the general findings of [/patterns/index.md](/patterns/index.md) — Cemri and colleagues note that multi-agent performance gains on popular benchmarks are often minimal and build a failure taxonomy spanning system design, inter-agent misalignment and task verification.
    Compared against: Single-agent and simpler baselines on popular benchmarks. Domain: Multi-agent frameworks across coding, math and general tasks. Benchmarks: MAST-Data.
    Caveat: The taxonomy characterizes failures rather than measuring a single head-to-head gain.
- Source: [A Systematic Study of LLM-Based Architectures for Automated Patching](https://arxiv.org/abs/2603.01257), Xu, Sheng, Chen, Huang, 2026-03-01.
  - **Mixed**, filed under [supervisor](/patterns/supervisor.md) — Xu and colleagues report that a multi-agent patching system did not consistently beat a well-designed single agent, winning with one model and losing with another at higher overhead, while a general-purpose coding agent patched the most vulnerabilities.
    Compared against: A single-agent patcher, a fixed workflow and a general-purpose coding agent on the same tasks. Domain: Security: repairing real-world vulnerabilities in large Java projects. Benchmarks: AIxCC.
    Caveat: A small set of vulnerabilities from one competition, with each architecture reimplemented by the authors.

### planner_worker

- Source: [Aider blog: Separating code reasoning and editing](https://aider.chat/2024/09/26/architect.html), Aider (Paul Gauthier), 2024-09-26.
  - **Helped**, filed under [planner_worker](/patterns/planner_worker.md) — Aider reports that splitting work between an architect model that proposes a solution and an editor model that writes the edits scored at or above each model working alone on its code editing benchmark, with the best pairing setting a new high score.
    Compared against: The same models editing code on their own. Domain: Code editing. Benchmarks: aider code editing benchmark.
    Caveat: Self-published benchmark by the tool's developer, and the top pairings were described as too slow for interactive use.

### role_pipeline

- Source: [ChatDev: Communicative Agents for Software Development](https://arxiv.org/abs/2307.07924), Qian et al., 2023-07-16.
  - **Helped**, filed under [role_pipeline](/patterns/role_pipeline.md) — The authors report that a chat chain moving specialised agents through design, coding and testing phases produced more complete, executable and consistent software than a single-agent tool and than MetaGPT, and that removing the agents’ roles caused the largest drop in their ablation.
    Compared against: GPT-Engineer, a single-agent approach, and the MetaGPT multi-agent framework. Domain: Software development from natural-language requirements. Benchmarks: SRDD.
    Caveat: Evaluated on the authors’ own requirement dataset with proxy metrics and pairwise preference judgements rather than executed test suites.
- Source: [MapCoder: Multi-Agent Code Generation for Competitive Problem Solving](https://arxiv.org/abs/2405.11403), Islam et al., 2024-05-18.
  - **Helped**, filed under [role_pipeline](/patterns/role_pipeline.md) — The authors report that a pipeline of agents that recall similar problems, plan, write code and debug it outperformed direct, chain-of-thought, planning and reflection prompting and earlier multi-agent frameworks across several models, and that removing any one agent lowered the result, the debugging agent most.
    Compared against: Direct, chain-of-thought, self-planning, analogical and Reflexion prompting, and the Self-collaboration and AlphaCodium frameworks. Domain: Competitive programming and program synthesis. Benchmarks: HumanEval, MBPP, APPS, CodeContests, xCodeEval.
    Caveat: The authors note that it generates many tokens, the comparison is not matched for compute, and a failed debugging stage loops back to the next plan, so the chain is not strictly one-way.
- Source: [On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents](https://arxiv.org/abs/2408.00989), Huang et al., 2024-08-02.
  - **Hurt**, filed under [role_pipeline](/patterns/role_pipeline.md) — The authors report that one-way linear chains of agents lost the most performance when faulty agents were injected, while a structure in which a leader directs two peers that talk to each other lost the least.
    Compared against: Flat and hierarchical multi-agent structures under the same injected faults. Domain: Code generation, math, translation and text evaluation. Benchmarks: HumanEval, CIAR, CommonMT, FairEval.
    Caveat: Structures were represented by a handful of existing systems that differ in more than topology, and robustness under injected errors is not the same as baseline accuracy.
  - **Helped**, filed under [hierarchical_delegation](/patterns/hierarchical_delegation.md) — The authors report that a mixed hierarchical structure lost the least performance when faulty agents were injected, while one-way linear pipelines such as MetaGPT-style chains lost the most.
    Compared against: Linear and flat multi-agent structures. Domain: Code generation, math, translation and text evaluation. Benchmarks: HumanEval, CIAR, CommonMT, FairEval.
    Caveat: Structures were represented by a handful of existing systems that differ in more than topology, and robustness under injected errors is not the same as baseline accuracy.

### hierarchical_delegation

- Source: [MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework](https://arxiv.org/abs/2308.00352), Hong et al., 2023-08-01.
  - **Helped**, filed under [hierarchical_delegation](/patterns/hierarchical_delegation.md) — The authors report that encoding standardized operating procedures into an assembly line of role agents produced more coherent software solutions than earlier chat-based multi-agent systems and used fewer tokens per line of code than ChatDev, though more tokens in total.
    Compared against: Chat-based multi-agent frameworks such as ChatDev and AgentVerse, and single models. Domain: Software engineering and code generation. Benchmarks: HumanEval, MBPP, SoftwareDev.
    Caveat: Evaluated largely by the framework's authors, including on a self-constructed software task set.

### shared_ledger

- Source: [Anthropic Engineering: Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents), Justin Young (Anthropic), 2025-11-26.
  - **Mixed**, filed under [shared_ledger](/patterns/shared_ledger.md) — Anthropic reports that a progress log, a feature list and git history shared across successive agent sessions helped long-running coding agents resume work, but agents still declared victory early or marked features complete without end-to-end testing.
    Compared against: Agents working without persistent progress artifacts. Domain: Long-running software development.
    Caveat: Vendor engineering report with qualitative observations only and no measured comparison.
  - **Helped**, filed under [successor_handoff](/patterns/successor_handoff.md) — Anthropic reports that compaction alone was not enough for a frontier coding agent working across many context windows, and that an initializer agent plus structured progress notes for each new session addressed failures like premature completion.
    Compared against: A coding agent looping across context windows with compaction only. Domain: Long-running software development.
    Caveat: Engineering guidance from qualitative experience without a controlled benchmark.

### mailbox_network

- Source: [Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems](https://arxiv.org/abs/2410.02506), Zhang et al., 2024-10-03.
  - **Helped**, filed under [mailbox_network](/patterns/mailbox_network.md) — The authors report substantial communication redundancy in multi-agent message-passing graphs and that pruning it kept performance comparable while greatly reducing token cost.
    Compared against: Chain, tree, star, complete, layered and random topologies and frameworks such as AutoGen and GPTSwarm. Domain: General, math reasoning and code generation. Benchmarks: MMLU, GSM8K, MultiArith, SVAMP, AQuA, HumanEval.
    Caveat: Evaluated on short-form reasoning and coding benchmarks by the proposing authors.

### implement_review

- Source: [Is Self-Repair a Silver Bullet for Code Generation?](https://arxiv.org/abs/2306.09896), Olausson et al., 2023-06-16.
  - **Mixed**, filed under [implement_review](/patterns/implement_review.md) — The authors report that self-repair gains for code generation were often modest or absent once repair cost was accounted for, but became substantially larger when feedback came from a stronger model or from humans.
    Compared against: Sampling more initial programs at equal budget. Domain: Code generation. Benchmarks: HumanEval, APPS.
    Caveat: Model generations tested are now dated, and the stronger-reviewer and human-feedback conditions were small-scale.
- Source: [LLM Critics Help Catch LLM Bugs](https://arxiv.org/abs/2407.00215), McAleese et al. (OpenAI), 2024-06-28.
  - **Helped**, filed under [implement_review](/patterns/implement_review.md) — The authors report that trained LLM critics caught more bugs in model-written code than paid human contractors and their critiques were usually preferred, but critics also hallucinated bugs, and human-plus-critic teams hallucinated less.
    Compared against: Human contractors reviewing code without assistance. Domain: Code review of model-written code.
    Caveat: The critics were specially trained with reinforcement learning from human feedback rather than prompted, and results are self-reported by the vendor.
- Source: [AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation](https://arxiv.org/abs/2312.13010), Huang et al., 2023-12-20.
  - **Helped**, filed under [implement_review](/patterns/implement_review.md) — The authors report that a programmer agent refined by feedback from separate test-designer and test-executor agents outperformed single-model code generation and prior enhancement methods while using fewer tokens.
    Compared against: Single-model code generation and prompt-engineering enhancement methods. Domain: Function-level code generation. Benchmarks: HumanEval, MBPP.
    Caveat: Feedback comes from executing generated tests, so gains may reflect execution feedback more than LLM review, and benchmarks are short standalone functions.
- Source: [Self-collaboration Code Generation via ChatGPT](https://arxiv.org/abs/2304.07590), Dong et al., 2023-04-15.
  - **Helped**, filed under [implement_review](/patterns/implement_review.md) — The authors report that a virtual team of analyst, coder and tester roles substantially improved pass rates over the same base model acting alone.
    Compared against: The same base model acting as a single agent. Domain: Code generation. Benchmarks: HumanEval, MBPP.
    Caveat: Evaluated mainly with an early ChatGPT model on function-level benchmarks, without compute-matched single-agent baselines.

### critic_loop

- Source: [Self-Refine: Iterative Refinement with Self-Feedback](https://arxiv.org/abs/2303.17651), Madaan et al., 2023-03-30.
  - **Helped**, filed under [critic_loop](/patterns/critic_loop.md) — The authors report that iterative self-feedback and refinement by the same model improved outputs over one-step generation across a diverse set of tasks, as judged by humans and automatic metrics.
    Compared against: One-step generation with the same model. Domain: Dialogue, code optimization, math reasoning and other generation tasks. Benchmarks: GSM8K.
    Caveat: Gains were concentrated in open-ended generation tasks, and later work found intrinsic self-correction on reasoning tasks to be much weaker.
- Source: [Reflexion: Language Agents with Verbal Reinforcement Learning](https://arxiv.org/abs/2303.11366), Shinn et al., 2023-03-20.
  - **Helped**, filed under [critic_loop](/patterns/critic_loop.md) — The authors report that agents reflecting verbally on task feedback and storing reflections in memory improved significantly over a baseline agent on sequential decision-making, coding and reasoning tasks.
    Compared against: The same agent without reflection. Domain: Sequential decision-making, coding and language reasoning. Benchmarks: HumanEval, ALFWorld, HotpotQA.
    Caveat: Improvements rely on external feedback signals such as unit tests or environment rewards across multiple trials, not on self-critique alone.

### council

- Source: [Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?](https://arxiv.org/abs/2502.00674), Li et al., 2025-02-02.
  - **Mixed**, filed under [council](/patterns/council.md) — The authors report that aggregating multiple outputs from only the single best model outperformed the standard mixture of different models in many scenarios, because mixing lowered average quality.
    Compared against: Aggregating several outputs of the single best model. Domain: Instruction following, knowledge, code reasoning and math. Benchmarks: AlpacaEval 2.0, MMLU, CRUX, MATH.
    Caveat: The authors also identify scenarios where mixing different models helps, so the result is about diversity versus quality rather than ensembling in general.

### dynamic_spawning

- Source: [Cognition blog: Don't Build Multi-Agents](https://cognition.com/blog/dont-build-multi-agents), Walden Yan (Cognition), 2025-06-12.
  - **Hurt**, filed under [dynamic_spawning](/patterns/dynamic_spawning.md) — Cognition argues against parallel subagents because subagents that cannot see each other's work make conflicting implicit decisions, recommending single-threaded agents by default.
    Compared against: A single-threaded linear agent sharing full context. Domain: Software engineering agents.
    Caveat: Vendor guidance based on an illustrative example, not a controlled experiment.

### adaptive_routing

- Source: [MasRouter: Learning to Route LLMs for Multi-Agent Systems](https://arxiv.org/abs/2502.11133), Yue et al., 2025-02-16.
  - **Helped**, filed under [adaptive_routing](/patterns/adaptive_routing.md) — MasRouter reports that jointly routing collaboration mode, roles and LLMs in a multi-agent system improved accuracy and reduced overhead versus prior methods.
    Compared against: Prior multi-agent routing and system design methods. Domain: Code generation, math and reasoning. Benchmarks: MBPP, HumanEval.
    Caveat: Author-run comparison with modest accuracy gains.

### tree_search

- Source: [Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models](https://arxiv.org/abs/2310.04406), Zhou et al., 2023-10-06.
  - **Helped**, filed under [tree_search](/patterns/tree_search.md) — The authors report that Monte Carlo tree search over an agent’s actions, with a language-model value function and self-reflection, outperformed acting and reflection baselines and earlier tree-search methods on programming, question answering and web shopping.
    Compared against: ReAct, Reflexion, and the tree-search methods Tree of Thoughts and RAP. Domain: Programming, interactive question answering, web shopping and math puzzles. Benchmarks: HumanEval, MBPP, HotpotQA, WebShop, Game of 24.
    Caveat: It assumes the environment can be reverted to an earlier state, and it costs more than a single acting agent.
- Source: [SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement](https://arxiv.org/abs/2410.20285), Antoniades et al., 2024-10-26.
  - **Helped**, filed under [tree_search](/patterns/tree_search.md) — The authors report that adding Monte Carlo tree search, a value agent and a discriminator agent to a software agent resolved more issues than the same agent without search across five models, and that results improved with deeper search.
    Compared against: The same open-source software agent without tree search. Domain: Repository-level software engineering: resolving GitHub issues. Benchmarks: SWE-bench Lite.
    Caveat: Author-run comparison under a capped search budget, and the value function often failed to pick the correct solution, which the authors name as the main room for improvement.

### architecture_search

- Source: [Automated Design of Agentic Systems](https://arxiv.org/abs/2408.08435), Hu, Lu, Clune, 2024-08-15.
  - **Helped**, filed under [architecture_search](/patterns/architecture_search.md) — ADAS reports that agents discovered by a meta agent programming new designs in code outperformed hand-designed agents and kept their advantage when transferred across domains and models.
    Compared against: State-of-the-art hand-designed agents. Domain: Reading comprehension, math, science and coding. Benchmarks: DROP, MGSM.
    Caveat: Search cost is not weighed against the gains in the abstract, and later work questions its cost-effectiveness.
- Source: [AFlow: Automating Agentic Workflow Generation](https://arxiv.org/abs/2410.10762), Zhang et al., 2024-10-14.
  - **Helped**, filed under [architecture_search](/patterns/architecture_search.md) — AFlow reports that Monte Carlo tree search over code-represented workflows beat prior baselines and let smaller models outperform a much larger model on some tasks at a fraction of its inference cost.
    Compared against: Manually designed workflows and prior automated methods. Domain: Question answering, code generation and math.
    Caveat: Author-run evaluation; the cheaper-model result holds only on specific tasks.
- Source: [Multi-agent Architecture Search via Agentic Supernet](https://arxiv.org/abs/2502.04180), Zhang et al., 2025-02-06.
  - **Helped**, filed under [architecture_search](/patterns/architecture_search.md) — MaAS reports that sampling query-dependent agent architectures from a learned supernet matched or beat existing systems while using only a fraction of their inference cost, with cross-dataset and cross-model transfer.
    Compared against: Handcrafted and automated multi-agent systems. Domain: Math, coding and tool use.
    Caveat: Author-run comparison; some reported accuracy gains are small.
- Source: [Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies](https://arxiv.org/abs/2502.02533), Zhou et al., 2025-02-04.
  - **Mixed**, filed under [architecture_search](/patterns/architecture_search.md) — MASS reports that only a small fraction of topologies improved performance, with most failing to help or degrading it, and that prompt optimization was more token-effective than scaling agent count.
    Compared against: Unoptimized topologies and agent-scaling strategies such as self-consistency and debate. Domain: Reasoning, multi-hop question answering and coding. Benchmarks: MATH, DROP, HotpotQA, MuSiQue, 2WikiMQA, MBPP, HumanEval, LiveCodeBench.
    Caveat: Findings come from one search framework and a limited set of model backbones.

[Reporting outcomes (limited rollout)](/docs/api/contributing.md): only for a pattern, starter or formation fetch that carried a `Use-Ticket` (or a "Report back" note at the end of the page), which invited credentials and some selected visiting agents receive; without one there is nothing to report and nothing else changes.
