---
schema: formation.domain_findings/v0.1
kind: domain_findings
visibility: public
canonical_url: https://topologyindex.com/domains/research/findings.md
description: 'Each published study tagged with research, once, with every finding filed under a pattern page: the sentence, comparison, domain, caveat and source. Attributed, stated without figures, never a ranking.'
domain: research
domain_page: /domains/research.md
findings:
  - citations:
      - compared_against: 'Multi-agent coordination under matched tools, prompts and compute'
        direction: mixed
        pattern: single_agent
      - compared_against: 'Single-agent systems and centralized, decentralized and hybrid multi-agent architectures'
        direction: mixed
        pattern: independent_workers
      - compared_against: 'Single-agent systems and independent, decentralized and hybrid multi-agent architectures'
        direction: mixed
        pattern: supervisor
      - compared_against: 'A single agent and centralized, independent and hybrid multi-agent architectures'
        direction: mixed
        pattern: mailbox_network
    source_id: arxiv:2512.08296
    url: https://arxiv.org/abs/2512.08296
  - citations:
      - compared_against: 'A multi-agent system with a lead agent and parallel subagents'
        direction: hurt
        pattern: single_agent
      - compared_against: 'Single-agent Claude on the same research tasks'
        direction: mixed
        pattern: independent_workers
      - compared_against: 'A single agent using the same model as the lead agent'
        direction: helped
        pattern: supervisor
      - compared_against: 'A single agent using the stronger model alone'
        direction: helped
        pattern: dynamic_spawning
    source_id: web:anthropic.com/engineering/multi-agent-research-system
    url: https://www.anthropic.com/engineering/multi-agent-research-system
  - citations:
      - compared_against: 'A single agent with the same model, search and web-reading tools'
        direction: helped
        pattern: map_reduce
    source_id: arxiv:2508.07999
    url: https://arxiv.org/abs/2508.07999
  - citations:
      - compared_against: 'Direct prompting and the AutoSurvey and MASS-Survey pipelines'
        direction: helped
        pattern: map_reduce
    source_id: arxiv:2510.05138
    url: https://arxiv.org/abs/2510.05138
  - citations:
      - compared_against: 'Single-model search agents such as WebThinker and ReAct, plan-and-solve prompting, and retrieval-augmented generation'
        direction: helped
        pattern: planner_worker
    source_id: arxiv:2507.02652
    url: https://arxiv.org/abs/2507.02652
  - citations:
      - compared_against: 'Master-slave multi-agent coordination, retrieval-augmented generation and a single agent'
        direction: helped
        pattern: blackboard
    source_id: arxiv:2510.01285
    url: https://arxiv.org/abs/2510.01285
  - citations:
      - compared_against: 'The same idea generator with fewer or no rounds of reviewing-agent feedback'
        direction: helped
        pattern: critic_loop
    source_id: arxiv:2404.07738
    url: https://arxiv.org/abs/2404.07738
  - citations:
      - compared_against: 'A single LLM screening with the same model'
        direction: helped
        pattern: council
    source_id: arxiv:2607.21920
    url: https://arxiv.org/abs/2607.21920
  - citations:
      - compared_against: 'The same model without search, a ReAct-style search agent, and the ChatGPT-Web and Perplexity Pro products'
        direction: helped
        pattern: dynamic_spawning
    source_id: arxiv:2407.20183
    url: https://arxiv.org/abs/2407.20183
  - citations:
      - compared_against: 'The same agent without context management features'
        direction: helped
        pattern: successor_handoff
    source_id: web:claude.com/blog/context-management
    url: https://claude.com/blog/context-management
  - citations:
      - compared_against: 'Single-agent web search and single-agent deep research systems'
        direction: mixed
        pattern: null
    source_id: arxiv:2510.14240
    url: https://arxiv.org/abs/2510.14240
path: /domains/research/findings.md
pattern_index: /patterns/index.md
product_api_version: v1
schema_version: v0.1
title: 'Published findings on multi-agent patterns studied in research'
---

# Published findings on multi-agent patterns studied in research

The published findings of the [`research` domain](/domains/research.md), in words. That page
has the task shapes common in the domain, their rows of the [decision guide](/patterns/index.md)
and the starters. None of it has been measured here, and nothing on this page ranks patterns.

## Published findings in this domain

Attributed to each source and stated without figures. Each source is listed once, under the
first pattern in the vocabulary that cites it, with every pattern page that cites it for this
domain. The order is the pattern vocabulary, never the direction, and unfavourable results
are included on purpose. None of this is evidence produced here. Reviewed 2026-09-23.
Each label says how the pattern each finding is filed under fared against what it was compared with, as the source reports
it; the labels are defined at [/docs/schemas/pattern/v0.1.md](/docs/schemas/pattern/v0.1.md).

### single_agent

- Source: [Towards a Science of Scaling Agent Systems](https://arxiv.org/abs/2512.08296), Kim et al., 2025-12-09.
  - **Mixed**, filed under [single_agent](/patterns/single_agent.md) — In controlled comparisons with tools, prompts and compute standardized, the authors report that multi-agent coordination helped on decomposable tasks, hurt on sequential planning, and showed diminishing returns once the single-agent baseline was already strong.
    Compared against: Multi-agent coordination under matched tools, prompts and compute. Domain: Agentic benchmarks spanning financial reasoning, web browsing, planning and tool use.
    Caveat: The fitted predictive model explains only part of the variance, and findings depend on the specific benchmarks and model families studied.
  - **Mixed**, filed under [independent_workers](/patterns/independent_workers.md) — The authors report that independent multi-agent setups without centralized verification propagate and amplify errors far more than centralized coordination, and that multi-agent gains vanish or reverse once the single-agent baseline is already strong.
    Compared against: Single-agent systems and centralized, decentralized and hybrid multi-agent architectures. Domain: Agentic web browsing, finance, planning, workplace, software engineering and terminal tasks. Benchmarks: BrowseComp-Plus, Finance Agent, PlanCraft, WorkBench, SWE-bench Verified, Terminal-Bench.
    Caveat: The fitted predictive model explains only a modest share of performance variance, so the patterns are tendencies rather than guarantees.
  - **Mixed**, filed under [supervisor](/patterns/supervisor.md) — The authors report that centralized orchestration strongly helped on decomposable financial reasoning but degraded performance on sequential planning, while containing error amplification far better than independent agents at a large token overhead.
    Compared against: Single-agent systems and independent, decentralized and hybrid multi-agent architectures. Domain: Agentic benchmarks spanning finance, web browsing, planning, workplace, software and terminal tasks. Benchmarks: Finance-Agent, PlanCraft, BrowseComp-Plus, Workbench, SWE-bench Verified, Terminal-Bench.
    Caveat: Results depend on the chosen model families and benchmark set, and the fitted predictive model explains only part of the variance.
  - **Mixed**, filed under [mailbox_network](/patterns/mailbox_network.md) — The authors report that a decentralized peer-to-peer architecture gave a modest gain on web browsing and a large gain on financial reasoning but lost ground on sequential planning, with communication overhead well above a single agent.
    Compared against: A single agent and centralized, independent and hybrid multi-agent architectures. Domain: Agentic benchmarks spanning finance, web browsing and planning. Benchmarks: BrowseComp-Plus, Finance-Agent, PlanCraft.
    Caveat: Outcomes depend strongly on task structure and model family in this study.
- Source: [Anthropic Engineering: How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system), Hadfield et al. (Anthropic), 2025-06-13.
  - **Hurt**, filed under [single_agent](/patterns/single_agent.md) — The authors report that a multi-agent research system with a lead agent and parallel subagents substantially outperformed a single agent on their internal research evaluation, while consuming far more tokens and fitting poorly on tightly coupled tasks such as most coding.
    Compared against: A multi-agent system with a lead agent and parallel subagents. Domain: Open-ended web research.
    Caveat: Self-reported by the vendor on an internal evaluation, and the authors note token usage alone explains most of the performance variance, so compute was not matched.
  - **Mixed**, filed under [independent_workers](/patterns/independent_workers.md) — Anthropic reports that a lead agent spawning parallel subagents with separate context windows far outperformed a single agent on its internal breadth-first research eval, while using many times more tokens and fitting poorly to tasks that need shared context, such as most coding.
    Compared against: Single-agent Claude on the same research tasks. Domain: Open-ended web research. Benchmarks: BrowseComp.
    Caveat: The headline comparison uses an internal eval, and the authors attribute much of the gain to spending more tokens rather than to the topology itself.
  - **Helped**, filed under [supervisor](/patterns/supervisor.md) — Anthropic reports that its lead-agent-plus-subagents research system substantially outperformed a single agent using the stronger lead model on its internal research evaluation.
    Compared against: A single agent using the same model as the lead agent. Domain: Open-ended web research.
    Caveat: Vendor self-report on an unpublished internal evaluation with no public replication.
  - **Helped**, filed under [dynamic_spawning](/patterns/dynamic_spawning.md) — Anthropic reports that its research system, in which a lead agent spawns parallel subagents, substantially outperformed a single-agent setup on its internal research evaluation.
    Compared against: A single agent using the stronger model alone. Domain: Open-ended web research.
    Caveat: Internal, vendor-run evaluation with no public benchmark for the headline comparison, and Anthropic notes the multi-agent system uses many times more tokens than chat.

### map_reduce

- Source: [WideSearch: Benchmarking Agentic Broad Info-Seeking](https://arxiv.org/abs/2508.07999), Wong et al., 2025-08-11.
  - **Helped**, filed under [map_reduce](/patterns/map_reduce.md) — Wong and colleagues report that a main agent decomposing a broad collection task, with sub-agents searching the parts in parallel and the main agent aggregating their results, consistently outperformed a single agent with the same model and tools.
    Compared against: A single agent with the same model, search and web-reading tools. Domain: Broad information seeking: collecting many verifiable facts from the web into a table. Benchmarks: WideSearch.
    Caveat: The advantage shows in partial-correctness scores; almost every system, split or not, still failed nearly every task outright.
- Source: [LiRA: A Multi-Agent Framework for Reliable and Readable Literature Review Generation](https://arxiv.org/abs/2510.05138), Go et al., 2025-10-01.
  - **Helped**, filed under [map_reduce](/patterns/map_reduce.md) — Go and colleagues report that a literature-review workflow that outlines, drafts subsections in parallel, then edits and reviews the combined article scored higher on writing and citation quality than direct prompting and the AutoSurvey and MASS-Survey pipelines.
    Compared against: Direct prompting and the AutoSurvey and MASS-Survey pipelines. Domain: Literature review generation from a set of references. Benchmarks: SciReviewGen.
    Caveat: Every run used one unseeded model, and the authors’ own ablations found that the editor agent slightly lowered most scores and that adding a researcher agent lowered them further.

### planner_worker

- Source: [HiRA: A Hierarchical Reasoning Framework for Decoupled Planning and Execution in Deep Search](https://arxiv.org/abs/2507.02652), Jin et al., 2025-07-03.
  - **Helped**, filed under [planner_worker](/patterns/planner_worker.md) — Jin and colleagues report that separating a planner that decomposes a search task from executor agents that carry out each subtask with their own tools outperformed single-model reasoning agents with search, while producing shorter reasoning chains and fewer interactions.
    Compared against: Single-model search agents such as WebThinker and ReAct, plan-and-solve prompting, and retrieval-augmented generation. Domain: Deep search: complex multi-step information seeking. Benchmarks: GAIA, WebWalkerQA, SimpleQA, Humanity's Last Exam.
    Caveat: Evaluated by the proposing authors, and the margin over the strongest baseline was small on several benchmarks.

### blackboard

- Source: [LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science](https://arxiv.org/abs/2510.01285), Salemi et al., 2025-09-30.
  - **Helped**, filed under [blackboard](/patterns/blackboard.md) — The authors report that letting sub-agents volunteer answers to requests posted on a shared blackboard beat a central controller that assigns tasks, as well as retrieval and single-agent baselines, for finding relevant data, at a higher cost per question than the controller design.
    Compared against: Master-slave multi-agent coordination, retrieval-augmented generation and a single agent. Domain: Data discovery in data-science question answering. Benchmarks: KramaBench, DSBench, DA-Code.
    Caveat: Tested on data-lake discovery tasks by the proposing authors, and it cost more per question than the master-slave baseline.

### critic_loop

- Source: [ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models](https://arxiv.org/abs/2404.07738), Baek et al., 2024-04-11.
  - **Helped**, filed under [critic_loop](/patterns/critic_loop.md) — Baek and colleagues report that iteratively refining research ideas with feedback from LLM reviewing agents improved the ideas under both human and model-based judgment, with gains saturating after a few rounds.
    Compared against: The same idea generator with fewer or no rounds of reviewing-agent feedback. Domain: Research idea generation over scientific literature.
    Caveat: Idea quality was judged by an LLM and a small panel of researchers rather than by carrying the ideas out, and returns diminished with further rounds.

### council

- Source: [Systematic Literature Reviews With Two Multi-Agentic Systems And Human-In-The-Loop](https://arxiv.org/abs/2607.21920), Ren et al., 2026-07-24.
  - **Helped**, filed under [council](/patterns/council.md) — Ren and colleagues report that screening trials for a systematic review with several agents given different personas, which cross-review until they agree, uniformly improved accuracy over a single model screening with the same underlying model.
    Compared against: A single LLM screening with the same model. Domain: Screening clinical trials for systematic literature reviews.
    Caveat: Evaluated only on oncology trial registries by the proposing authors, and the cases the agents could not agree on were escalated to human reviewers.

### dynamic_spawning

- Source: [MindSearch: Mimicking Human Minds Elicits Deep AI Searcher](https://arxiv.org/abs/2407.20183), Chen et al., 2024-07-29.
  - **Helped**, filed under [dynamic_spawning](/patterns/dynamic_spawning.md) — Chen and colleagues report that a planner that grows a graph of sub-questions as results arrive, handing each new sub-question to its own searcher agent, beat both a model without search and a ReAct-style search agent on multi-hop question answering, and was preferred by human raters to two commercial AI search products.
    Compared against: The same model without search, a ReAct-style search agent, and the ChatGPT-Web and Perplexity Pro products. Domain: Web information seeking and multi-hop question answering. Benchmarks: Bamboogle, Musique, HotpotQA.
    Caveat: Proposing authors, with a ReAct agent ahead on some question subsets, and the authors note factuality improved less than depth and breadth in the human evaluation.

### successor_handoff

- Source: [Claude blog: Managing context on the Claude Developer Platform](https://claude.com/blog/context-management), Anthropic, 2025-09-29.
  - **Helped**, filed under [successor_handoff](/patterns/successor_handoff.md) — Anthropic reports that combining a memory tool with context editing improved performance on an internal agentic search evaluation and let long workflows finish that otherwise failed from context exhaustion.
    Compared against: The same agent without context management features. Domain: Agentic web search.
    Caveat: Vendor-run internal evaluation with no public benchmark.

### Multi-agent systems in general

- Source: [LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild](https://arxiv.org/abs/2510.14240), Wang et al., 2025-10-16.
  - **Mixed**, filed under the general findings of [/patterns/index.md](/patterns/index.md) — Wang and colleagues report that among the deep research products they evaluated, multi-agent systems led on presentation and on linking claims to their citations, while single-agent web search systems were the most factually and logically consistent and single-agent deep research systems linked citations worst.
    Compared against: Single-agent web search and single-agent deep research systems. Domain: Deep research: citation-grounded reports from live web sources. Benchmarks: LiveResearchBench.
    Caveat: The systems are products built on different models and tools, so the comparison is between complete systems and not controlled for the arrangement.

[Reporting outcomes (limited rollout)](/docs/api/contributing.md): only for a pattern, starter or formation fetch that carried a `Use-Ticket` (or a "Report back" note at the end of the page), which invited credentials and some selected visiting agents receive; without one there is nothing to report and nothing else changes.
