---
schema: formation.domain_findings/v0.1
kind: domain_findings
visibility: public
canonical_url: https://topologyindex.com/domains/operations/findings.md
description: 'Each published study tagged with operations, once, with every finding filed under a pattern page: the sentence, comparison, domain, caveat and source. Attributed, stated without figures, never a ranking.'
domain: operations
domain_page: /domains/operations.md
findings:
  - citations:
      - compared_against: 'Multi-agent coordination under matched tools, prompts and compute'
        direction: mixed
        pattern: single_agent
      - compared_against: 'Single-agent systems and centralized, decentralized and hybrid multi-agent architectures'
        direction: mixed
        pattern: independent_workers
      - compared_against: 'Single-agent systems and independent, decentralized and hybrid multi-agent architectures'
        direction: mixed
        pattern: supervisor
      - compared_against: 'A single agent and centralized, independent and hybrid multi-agent architectures'
        direction: mixed
        pattern: mailbox_network
      - compared_against: 'A single agent under matched tools, prompts and compute'
        direction: mixed
        pattern: null
    source_id: arxiv:2512.08296
    url: https://arxiv.org/abs/2512.08296
  - citations:
      - compared_against: 'Single-agent solo setups'
        direction: mixed
        pattern: lane_swarm
      - compared_against: 'A single agent'
        direction: mixed
        pattern: dynamic_spawning
    source_id: arxiv:2308.10848
    url: https://arxiv.org/abs/2308.10848
  - citations:
      - compared_against: 'Published state-of-the-art agent systems on each benchmark'
        direction: no_clear_gain
        pattern: supervisor
      - compared_against: 'The full Magentic-One orchestrator with task and progress ledgers'
        direction: hurt
        pattern: blackboard
      - compared_against: 'The same agents coordinated by a basic group chat without ledgers'
        direction: helped
        pattern: shared_ledger
    source_id: arxiv:2411.04468
    url: https://arxiv.org/abs/2411.04468
  - citations:
      - compared_against: 'Single-agent approaches'
        direction: helped
        pattern: supervisor
    source_id: arxiv:2412.05449
    url: https://arxiv.org/abs/2412.05449
  - citations:
      - compared_against: 'ReAct-style sequential function calling'
        direction: helped
        pattern: planner_worker
    source_id: arxiv:2312.04511
    url: https://arxiv.org/abs/2312.04511
  - citations:
      - compared_against: 'Executor-only agents and planners without targeted training'
        direction: mixed
        pattern: planner_worker
    source_id: arxiv:2503.09572
    url: https://arxiv.org/abs/2503.09572
  - citations:
      - compared_against: 'ReAct, Reflexion, and the tree-search methods Tree of Thoughts and RAP'
        direction: helped
        pattern: tree_search
    source_id: arxiv:2310.04406
    url: https://arxiv.org/abs/2310.04406
  - citations:
      - compared_against: 'The same web agent without search'
        direction: helped
        pattern: tree_search
    source_id: arxiv:2407.01476
    url: https://arxiv.org/abs/2407.01476
  - citations:
      - compared_against: 'Handcrafted and automated multi-agent systems'
        direction: helped
        pattern: architecture_search
    source_id: arxiv:2502.04180
    url: https://arxiv.org/abs/2502.04180
path: /domains/operations/findings.md
pattern_index: /patterns/index.md
product_api_version: v1
schema_version: v0.1
title: 'Published findings on multi-agent patterns studied in operations'
---

# Published findings on multi-agent patterns studied in operations

The published findings of the [`operations` domain](/domains/operations.md), in words. That page
has the task shapes common in the domain, their rows of the [decision guide](/patterns/index.md)
and the starters. None of it has been measured here, and nothing on this page ranks patterns.

## Published findings in this domain

Attributed to each source and stated without figures. Each source is listed once, under the
first pattern in the vocabulary that cites it, with every pattern page that cites it for this
domain. The order is the pattern vocabulary, never the direction, and unfavourable results
are included on purpose. None of this is evidence produced here. Reviewed 2026-09-23.
Each label says how the pattern each finding is filed under fared against what it was compared with, as the source reports
it; the labels are defined at [/docs/schemas/pattern/v0.1.md](/docs/schemas/pattern/v0.1.md).

### single_agent

- Source: [Towards a Science of Scaling Agent Systems](https://arxiv.org/abs/2512.08296), Kim et al., 2025-12-09.
  - **Mixed**, filed under [single_agent](/patterns/single_agent.md) — In controlled comparisons with tools, prompts and compute standardized, the authors report that multi-agent coordination helped on decomposable tasks, hurt on sequential planning, and showed diminishing returns once the single-agent baseline was already strong.
    Compared against: Multi-agent coordination under matched tools, prompts and compute. Domain: Agentic benchmarks spanning financial reasoning, web browsing, planning and tool use.
    Caveat: The fitted predictive model explains only part of the variance, and findings depend on the specific benchmarks and model families studied.
  - **Mixed**, filed under [independent_workers](/patterns/independent_workers.md) — The authors report that independent multi-agent setups without centralized verification propagate and amplify errors far more than centralized coordination, and that multi-agent gains vanish or reverse once the single-agent baseline is already strong.
    Compared against: Single-agent systems and centralized, decentralized and hybrid multi-agent architectures. Domain: Agentic web browsing, finance, planning, workplace, software engineering and terminal tasks. Benchmarks: BrowseComp-Plus, Finance Agent, PlanCraft, WorkBench, SWE-bench Verified, Terminal-Bench.
    Caveat: The fitted predictive model explains only a modest share of performance variance, so the patterns are tendencies rather than guarantees.
  - **Mixed**, filed under [supervisor](/patterns/supervisor.md) — The authors report that centralized orchestration strongly helped on decomposable financial reasoning but degraded performance on sequential planning, while containing error amplification far better than independent agents at a large token overhead.
    Compared against: Single-agent systems and independent, decentralized and hybrid multi-agent architectures. Domain: Agentic benchmarks spanning finance, web browsing, planning, workplace, software and terminal tasks. Benchmarks: Finance-Agent, PlanCraft, BrowseComp-Plus, Workbench, SWE-bench Verified, Terminal-Bench.
    Caveat: Results depend on the chosen model families and benchmark set, and the fitted predictive model explains only part of the variance.
  - **Mixed**, filed under [mailbox_network](/patterns/mailbox_network.md) — The authors report that a decentralized peer-to-peer architecture gave a modest gain on web browsing and a large gain on financial reasoning but lost ground on sequential planning, with communication overhead well above a single agent.
    Compared against: A single agent and centralized, independent and hybrid multi-agent architectures. Domain: Agentic benchmarks spanning finance, web browsing and planning. Benchmarks: BrowseComp-Plus, Finance-Agent, PlanCraft.
    Caveat: Outcomes depend strongly on task structure and model family in this study.
  - **Mixed**, filed under the general findings of [/patterns/index.md](/patterns/index.md) — Kim and colleagues report that multi-agent coordination ranges from large gains on decomposable tasks to large losses on sequential planning, with diminishing or negative returns once the single agent is already strong and higher overhead on tool-heavy tasks.
    Compared against: A single agent under matched tools, prompts and compute. Domain: Agentic benchmarks including finance, planning and tool use.
    Caveat: Controlled study across a fixed set of architectures and model families; the predictive model explains only part of the variance.

### lane_swarm

- Source: [AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors](https://arxiv.org/abs/2308.10848), Chen et al., 2023-08-21.
  - **Mixed**, filed under [lane_swarm](/patterns/lane_swarm.md) — The authors report that dynamically composed agent groups can outperform a single agent, but also document cases where group discussion hurt a weaker model and negative emergent behaviors such as destructive actions.
    Compared against: Single-agent solo setups. Domain: Reasoning, coding, tool use and embodied Minecraft tasks. Benchmarks: FED, Commongen-Challenge, MGSM, Logic Grid Puzzles, HumanEval.
    Caveat: Groups in these experiments are small, so the results say little about large swarms.
  - **Mixed**, filed under [dynamic_spawning](/patterns/dynamic_spawning.md) — AgentVerse reports that dynamically adjusting group composition by recruiting expert agents lets multi-agent groups outperform a single agent, while also documenting negative emergent social behaviors.
    Compared against: A single agent. Domain: Text understanding, reasoning, coding, tool use and embodied tasks.
    Caveat: Author-run evaluation of the proposed framework; the negative behaviors are described qualitatively.

### supervisor

- Source: [Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks](https://arxiv.org/abs/2411.04468), Fourney et al. (Microsoft Research), 2024-11-07.
  - **No clear gain**, filed under [supervisor](/patterns/supervisor.md) — The authors report that their orchestrator-led team of specialist agents achieved performance statistically comparable to, not better than, state-of-the-art systems on general agentic benchmarks, and trailed the top entries on one web benchmark.
    Compared against: Published state-of-the-art agent systems on each benchmark. Domain: Generalist web, file and coding tasks. Benchmarks: GAIA, AssistantBench, WebArena.
    Caveat: First-party evaluation by the system's builders against heterogeneous published baselines rather than matched single-agent controls.
  - **Hurt**, filed under [blackboard](/patterns/blackboard.md) — The authors report that replacing the Magentic-One orchestrator's ledgers with AutoGen's basic group chat, where a selector only picks the next speaker on a shared transcript, markedly lowered performance.
    Compared against: The full Magentic-One orchestrator with task and progress ledgers. Domain: Generalist agentic tasks. Benchmarks: GAIA.
    Caveat: A single ablation on one validation split with one model, run by the system's own authors.
  - **Helped**, filed under [shared_ledger](/patterns/shared_ledger.md) — The authors report that Magentic-One's orchestrator, which maintains a task ledger of facts and plans and a progress ledger checked each step, performed markedly better than a variant with both ledgers removed.
    Compared against: The same agents coordinated by a basic group chat without ledgers. Domain: Generalist agentic tasks. Benchmarks: GAIA.
    Caveat: The ablation removes planning, loop detection and explicit instructions along with the ledgers, so the ledger's own contribution is not isolated.
- Source: [Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications](https://arxiv.org/abs/2412.05449), Shu et al. (AWS), 2024-12-06.
  - **Helped**, filed under [supervisor](/patterns/supervisor.md) — The authors report that multi-agent collaboration with a supervisor agent raised goal success over single-agent setups and that a routing mode reduced latency.
    Compared against: Single-agent approaches. Domain: Enterprise assistant scenarios.
    Caveat: Vendor technical report on handcrafted scenarios from a few enterprise domains, evaluated on its own product.

### planner_worker

- Source: [An LLM Compiler for Parallel Function Calling](https://arxiv.org/abs/2312.04511), Kim et al., 2023-12-07.
  - **Helped**, filed under [planner_worker](/patterns/planner_worker.md) — The authors report that a planner that emits a dependency graph of function calls, executed in parallel by a dispatcher, cut latency and cost and modestly improved accuracy relative to sequential ReAct.
    Compared against: ReAct-style sequential function calling. Domain: Parallelizable function-calling tasks.
    Caveat: Benefits depend on the task having independent sub-calls, and the 'up to' gains are best cases rather than averages.
- Source: [Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks](https://arxiv.org/abs/2503.09572), Erdogan et al., 2025-03-12.
  - **Mixed**, filed under [planner_worker](/patterns/planner_worker.md) — The authors report that adding an untrained planner failed to improve over fine-tuned executors, suggesting poor plans can confuse the executor, while a planner trained on synthetic plans plus dynamic replanning reached state-of-the-art web navigation results.
    Compared against: Executor-only agents and planners without targeted training. Domain: Long-horizon web navigation. Benchmarks: WebArena-Lite, WebVoyager.
    Caveat: The gains rely on substantial planner fine-tuning with synthetic data, so an off-the-shelf planner split alone did not deliver them.

### tree_search

- Source: [Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models](https://arxiv.org/abs/2310.04406), Zhou et al., 2023-10-06.
  - **Helped**, filed under [tree_search](/patterns/tree_search.md) — The authors report that Monte Carlo tree search over an agent’s actions, with a language-model value function and self-reflection, outperformed acting and reflection baselines and earlier tree-search methods on programming, question answering and web shopping.
    Compared against: ReAct, Reflexion, and the tree-search methods Tree of Thoughts and RAP. Domain: Programming, interactive question answering, web shopping and math puzzles. Benchmarks: HumanEval, MBPP, HotpotQA, WebShop, Game of 24.
    Caveat: It assumes the environment can be reverted to an earlier state, and it costs more than a single acting agent.
- Source: [Tree Search for Language Model Agents](https://arxiv.org/abs/2407.01476), Koh, McAleer, Fried, Salakhutdinov, 2024-07-01.
  - **Helped**, filed under [tree_search](/patterns/tree_search.md) — The authors report that best-first tree search in the real environment, on top of a strong web agent, raised task success markedly over the same agent without search, and that success kept rising with more search compute.
    Compared against: The same web agent without search. Domain: Web automation on realistic websites. Benchmarks: VisualWebArena, WebArena.
    Caveat: Search multiplies the actions taken in the environment and depends on returning to earlier states, and absolute success rates stay low.

### architecture_search

- Source: [Multi-agent Architecture Search via Agentic Supernet](https://arxiv.org/abs/2502.04180), Zhang et al., 2025-02-06.
  - **Helped**, filed under [architecture_search](/patterns/architecture_search.md) — MaAS reports that sampling query-dependent agent architectures from a learned supernet matched or beat existing systems while using only a fraction of their inference cost, with cross-dataset and cross-model transfer.
    Compared against: Handcrafted and automated multi-agent systems. Domain: Math, coding and tool use.
    Caveat: Author-run comparison; some reported accuracy gains are small.

[Reporting outcomes (limited rollout)](/docs/api/contributing.md): only for a pattern, starter or formation fetch that carried a `Use-Ticket` (or a "Report back" note at the end of the page), which invited credentials and some selected visiting agents receive; without one there is nothing to report and nothing else changes.
