---
schema: formation.domain_findings/v0.1
kind: domain_findings
visibility: public
canonical_url: https://topologyindex.com/domains/reasoning/findings.md
description: 'Each published study tagged with reasoning, once, with every finding filed under a pattern page: the sentence, comparison, domain, caveat and source. Attributed, stated without figures, never a ranking.'
domain: reasoning
domain_page: /domains/reasoning.md
findings:
  - citations:
      - compared_against: 'Multi-agent discussion frameworks using the same backbone models'
        direction: no_clear_gain
        pattern: single_agent
      - compared_against: 'A single agent with strong prompts and demonstrations'
        direction: no_clear_gain
        pattern: mailbox_network
      - compared_against: 'A single agent with strong prompts and demonstrations'
        direction: no_clear_gain
        pattern: null
    source_id: arxiv:2402.18272
    url: https://arxiv.org/abs/2402.18272
  - citations:
      - compared_against: 'Multi-agent debate methods'
        direction: no_clear_gain
        pattern: single_agent
      - compared_against: 'Single-agent chain-of-thought and self-consistency'
        direction: no_clear_gain
        pattern: debate
    source_id: arxiv:2502.08788
    url: https://arxiv.org/abs/2502.08788
  - citations:
      - compared_against: 'Multi-agent coordination under matched tools, prompts and compute'
        direction: mixed
        pattern: single_agent
      - compared_against: 'Single-agent systems and centralized, decentralized and hybrid multi-agent architectures'
        direction: mixed
        pattern: independent_workers
      - compared_against: 'Single-agent systems and independent, decentralized and hybrid multi-agent architectures'
        direction: mixed
        pattern: supervisor
      - compared_against: 'A single agent and centralized, independent and hybrid multi-agent architectures'
        direction: mixed
        pattern: mailbox_network
      - compared_against: 'A single agent under matched tools, prompts and compute'
        direction: mixed
        pattern: null
    source_id: arxiv:2512.08296
    url: https://arxiv.org/abs/2512.08296
  - citations:
      - compared_against: 'Single greedy-decoded chain-of-thought'
        direction: helped
        pattern: fan_out
    source_id: arxiv:2203.11171
    url: https://arxiv.org/abs/2203.11171
  - citations:
      - compared_against: 'Single LLM call and more elaborate prompting or multi-agent methods'
        direction: helped
        pattern: fan_out
    source_id: arxiv:2402.05120
    url: https://arxiv.org/abs/2402.05120
  - citations:
      - compared_against: 'Single-sample attempts'
        direction: mixed
        pattern: fan_out
    source_id: arxiv:2407.21787
    url: https://arxiv.org/abs/2407.21787
  - citations:
      - compared_against: 'Vote and Filter-Vote systems at smaller call counts'
        direction: mixed
        pattern: fan_out
      - compared_against: 'Voting systems with fewer model calls'
        direction: mixed
        pattern: council
    source_id: arxiv:2403.02419
    url: https://arxiv.org/abs/2403.02419
  - citations:
      - compared_against: 'Fine-tuned model producing a single answer'
        direction: helped
        pattern: fan_out
    source_id: arxiv:2110.14168
    url: https://arxiv.org/abs/2110.14168
  - citations:
      - compared_against: 'Standard sequential decoding'
        direction: mixed
        pattern: map_reduce
    source_id: arxiv:2307.15337
    url: https://arxiv.org/abs/2307.15337
  - citations:
      - compared_against: 'Smaller agent networks and regular topologies such as chains and meshes'
        direction: mixed
        pattern: lane_swarm
    source_id: arxiv:2406.07155
    url: https://arxiv.org/abs/2406.07155
  - citations:
      - compared_against: 'Single-agent solo setups'
        direction: mixed
        pattern: lane_swarm
      - compared_against: 'A single agent'
        direction: mixed
        pattern: dynamic_spawning
    source_id: arxiv:2308.10848
    url: https://arxiv.org/abs/2308.10848
  - citations:
      - compared_against: 'Expected task success of the same frameworks'
        direction: hurt
        pattern: supervisor
      - compared_against: 'Expectations of benefit from multi-agent frameworks with reviewer or verifier roles'
        direction: no_clear_gain
        pattern: implement_review
      - compared_against: 'Single-agent and simpler baselines on popular benchmarks'
        direction: no_clear_gain
        pattern: null
    source_id: arxiv:2503.13657
    url: https://arxiv.org/abs/2503.13657
  - citations:
      - compared_against: 'Interleaved observation-dependent reasoning such as ReAct'
        direction: helped
        pattern: planner_worker
    source_id: arxiv:2305.18323
    url: https://arxiv.org/abs/2305.18323
  - citations:
      - compared_against: 'Zero-shot and few-shot chain-of-thought prompting'
        direction: helped
        pattern: planner_worker
    source_id: arxiv:2305.04091
    url: https://arxiv.org/abs/2305.04091
  - citations:
      - compared_against: 'Flat and hierarchical multi-agent structures under the same injected faults'
        direction: hurt
        pattern: role_pipeline
      - compared_against: 'Linear and flat multi-agent structures'
        direction: helped
        pattern: hierarchical_delegation
    source_id: arxiv:2408.00989
    url: https://arxiv.org/abs/2408.00989
  - citations:
      - compared_against: 'Strong reasoning models and multi-agent frameworks such as AgentVerse'
        direction: helped
        pattern: hierarchical_delegation
    source_id: arxiv:2502.11098
    url: https://arxiv.org/abs/2502.11098
  - citations:
      - compared_against: 'A single centralized decision-maker with equivalent information access'
        direction: hurt
        pattern: hierarchical_delegation
    source_id: arxiv:2603.26993
    url: https://arxiv.org/abs/2603.26993
  - citations:
      - compared_against: 'Chain-of-thought, static multi-agent systems and autonomous multi-agent systems such as GPTSwarm and AFlow'
        direction: helped
        pattern: blackboard
    source_id: arxiv:2507.01701
    url: https://arxiv.org/abs/2507.01701
  - citations:
      - compared_against: 'Fully connected multi-agent debate'
        direction: helped
        pattern: mailbox_network
    source_id: arxiv:2406.11776
    url: https://arxiv.org/abs/2406.11776
  - citations:
      - compared_against: 'Chain, tree, star, complete, layered and random topologies and frameworks such as AutoGen and GPTSwarm'
        direction: helped
        pattern: mailbox_network
    source_id: arxiv:2410.02506
    url: https://arxiv.org/abs/2410.02506
  - citations:
      - compared_against: 'Self-consistency and ensembling over multiple reasoning paths'
        direction: no_clear_gain
        pattern: mailbox_network
      - compared_against: 'Self-consistency and ensembling over multiple reasoning paths'
        direction: no_clear_gain
        pattern: debate
      - compared_against: 'Self-consistency and ensembling prompting strategies'
        direction: no_clear_gain
        pattern: null
    source_id: arxiv:2311.17371
    url: https://arxiv.org/abs/2311.17371
  - citations:
      - compared_against: 'Different numbers of agents, rounds and debate or reflection strategies'
        direction: mixed
        pattern: mailbox_network
    source_id: arxiv:2310.02124
    url: https://arxiv.org/abs/2310.02124
  - citations:
      - compared_against: 'One-step generation with the same model'
        direction: helped
        pattern: critic_loop
    source_id: arxiv:2303.17651
    url: https://arxiv.org/abs/2303.17651
  - citations:
      - compared_against: 'The same agent without reflection'
        direction: helped
        pattern: critic_loop
    source_id: arxiv:2303.11366
    url: https://arxiv.org/abs/2303.11366
  - citations:
      - compared_against: 'The same model without tool-interactive critiquing'
        direction: helped
        pattern: critic_loop
    source_id: arxiv:2305.11738
    url: https://arxiv.org/abs/2305.11738
  - citations:
      - compared_against: 'The model''s initial answers before intrinsic self-correction'
        direction: hurt
        pattern: critic_loop
    source_id: arxiv:2310.01798
    url: https://arxiv.org/abs/2310.01798
  - citations:
      - compared_against: 'Iterative prompting with a sound external verifier and one-shot generation'
        direction: hurt
        pattern: critic_loop
    source_id: arxiv:2402.08115
    url: https://arxiv.org/abs/2402.08115
  - citations:
      - compared_against: 'Aggregating several outputs of the single best model'
        direction: mixed
        pattern: council
    source_id: arxiv:2502.00674
    url: https://arxiv.org/abs/2502.00674
  - citations:
      - compared_against: 'A single model instance and single-model reflection'
        direction: helped
        pattern: debate
    source_id: arxiv:2305.14325
    url: https://arxiv.org/abs/2305.14325
  - citations:
      - compared_against: 'Self-reflection with a single model'
        direction: helped
        pattern: debate
    source_id: arxiv:2305.19118
    url: https://arxiv.org/abs/2305.19118
  - citations:
      - compared_against: 'Naive baselines without debate, such as a single consultant'
        direction: helped
        pattern: debate
    source_id: arxiv:2402.06782
    url: https://arxiv.org/abs/2402.06782
  - citations:
      - compared_against: 'Majority voting over independent agent answers'
        direction: no_clear_gain
        pattern: debate
    source_id: arxiv:2508.17536
    url: https://arxiv.org/abs/2508.17536
  - citations:
      - compared_against: 'Existing multi-agent methods with predefined agents'
        direction: helped
        pattern: dynamic_spawning
    source_id: arxiv:2309.17288
    url: https://arxiv.org/abs/2309.17288
  - citations:
      - compared_against: 'Different model families and network sizes'
        direction: mixed
        pattern: coordinator_election
    source_id: arxiv:2507.08616
    url: https://arxiv.org/abs/2507.08616
  - citations:
      - compared_against: 'A shared initial majority vote without a leader'
        direction: no_clear_gain
        pattern: coordinator_election
    source_id: arxiv:2606.19111
    url: https://arxiv.org/abs/2606.19111
  - citations:
      - compared_against: 'Dictatorial and plurality collective decision rules'
        direction: mixed
        pattern: coordinator_election
    source_id: arxiv:2410.15168
    url: https://arxiv.org/abs/2410.15168
  - citations:
      - compared_against: 'Always using the strong model'
        direction: helped
        pattern: adaptive_routing
    source_id: arxiv:2406.18665
    url: https://arxiv.org/abs/2406.18665
  - citations:
      - compared_against: 'The best individual LLM API'
        direction: helped
        pattern: adaptive_routing
    source_id: arxiv:2305.05176
    url: https://arxiv.org/abs/2305.05176
  - citations:
      - compared_against: 'Prior multi-agent routing and system design methods'
        direction: helped
        pattern: adaptive_routing
    source_id: arxiv:2502.11133
    url: https://arxiv.org/abs/2502.11133
  - citations:
      - compared_against: 'ReAct, Reflexion, and the tree-search methods Tree of Thoughts and RAP'
        direction: helped
        pattern: tree_search
    source_id: arxiv:2310.04406
    url: https://arxiv.org/abs/2310.04406
  - citations:
      - compared_against: 'Input-output and chain-of-thought prompting of the same model'
        direction: helped
        pattern: tree_search
    source_id: arxiv:2305.10601
    url: https://arxiv.org/abs/2305.10601
  - citations:
      - compared_against: 'Best-of-N sampling scored by the same verifier'
        direction: mixed
        pattern: tree_search
    source_id: arxiv:2408.03314
    url: https://arxiv.org/abs/2408.03314
  - citations:
      - compared_against: 'State-of-the-art hand-designed agents'
        direction: helped
        pattern: architecture_search
    source_id: arxiv:2408.08435
    url: https://arxiv.org/abs/2408.08435
  - citations:
      - compared_against: 'Manually designed workflows and prior automated methods'
        direction: helped
        pattern: architecture_search
    source_id: arxiv:2410.10762
    url: https://arxiv.org/abs/2410.10762
  - citations:
      - compared_against: 'Handcrafted and automated multi-agent systems'
        direction: helped
        pattern: architecture_search
    source_id: arxiv:2502.04180
    url: https://arxiv.org/abs/2502.04180
  - citations:
      - compared_against: 'Unoptimized topologies and agent-scaling strategies such as self-consistency and debate'
        direction: mixed
        pattern: architecture_search
    source_id: arxiv:2502.02533
    url: https://arxiv.org/abs/2502.02533
path: /domains/reasoning/findings.md
pattern_index: /patterns/index.md
product_api_version: v1
schema_version: v0.1
title: 'Published findings on multi-agent patterns studied in reasoning'
---

# Published findings on multi-agent patterns studied in reasoning

The published findings of the [`reasoning` domain](/domains/reasoning.md), in words. That page
has the task shapes common in the domain, their rows of the [decision guide](/patterns/index.md)
and the starters. None of it has been measured here, and nothing on this page ranks patterns.

## Published findings in this domain

Attributed to each source and stated without figures. Each source is listed once, under the
first pattern in the vocabulary that cites it, with every pattern page that cites it for this
domain. The order is the pattern vocabulary, never the direction, and unfavourable results
are included on purpose. None of this is evidence produced here. Reviewed 2026-09-23.
Each label says how the pattern each finding is filed under fared against what it was compared with, as the source reports
it; the labels are defined at [/docs/schemas/pattern/v0.1.md](/docs/schemas/pattern/v0.1.md).

### single_agent

- Source: [Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?](https://arxiv.org/abs/2402.18272), Wang et al., 2024-02-28.
  - **No clear gain**, filed under [single_agent](/patterns/single_agent.md) — The authors report that a single agent with strong prompts reached nearly the same performance as the best existing multi-agent discussion method, with discussion only pulling ahead when no demonstrations were in the prompt.
    Compared against: Multi-agent discussion frameworks using the same backbone models. Domain: General reasoning tasks.
    Caveat: Results come from earlier model generations and prompt-based discussion setups, so they may not transfer to tool-using agents.
  - **No clear gain**, filed under [mailbox_network](/patterns/mailbox_network.md) — The authors report that a single agent with strong prompts nearly matched the best multi-agent discussion framework, with discussion helping mainly when no demonstrations were given and sometimes spreading wrong answers.
    Compared against: A single agent with strong prompts and demonstrations. Domain: Commonsense, math and deductive reasoning. Benchmarks: ECQA, GSM8K, FOLIO-wiki.
    Caveat: Limited to reasoning benchmarks with discussion-style exchanges rather than tool-using agents.
  - **No clear gain**, filed under the general findings of [/patterns/index.md](/patterns/index.md) — Wang and colleagues report that a single agent with strong prompts achieves almost the same performance as the best multi-agent discussion, which wins only when no demonstrations are given.
    Compared against: A single agent with strong prompts and demonstrations. Domain: Reasoning tasks.
    Caveat: Covers discussion-style multi-agent setups, not tool-using or parallel agent systems.
- Source: [Stop Overvaluing Multi-Agent Debate: We Must Rethink Evaluation and Embrace Model Heterogeneity](https://arxiv.org/abs/2502.08788), Zhang et al., 2025-02-12.
  - **No clear gain**, filed under [single_agent](/patterns/single_agent.md) — The authors report that multi-agent debate methods often failed to beat simple single-agent baselines such as chain-of-thought and self-consistency, even while using considerably more inference compute.
    Compared against: Multi-agent debate methods. Domain: Reasoning, knowledge and coding question answering.
    Caveat: The study covers a fixed set of debate methods and base models, and the authors find that mixing heterogeneous models can recover some gains.
  - **No clear gain**, filed under [debate](/patterns/debate.md) — The authors report that multi-agent debate often failed to outperform chain-of-thought and self-consistency despite using more inference compute, and that using heterogeneous models improved debate frameworks.
    Compared against: Single-agent chain-of-thought and self-consistency. Domain: Reasoning, knowledge and coding question answering.
    Caveat: Results are for the debate methods and base models evaluated, and heterogeneity was proposed as a remedy rather than exhaustively tested.
- Source: [Towards a Science of Scaling Agent Systems](https://arxiv.org/abs/2512.08296), Kim et al., 2025-12-09.
  - **Mixed**, filed under [single_agent](/patterns/single_agent.md) — In controlled comparisons with tools, prompts and compute standardized, the authors report that multi-agent coordination helped on decomposable tasks, hurt on sequential planning, and showed diminishing returns once the single-agent baseline was already strong.
    Compared against: Multi-agent coordination under matched tools, prompts and compute. Domain: Agentic benchmarks spanning financial reasoning, web browsing, planning and tool use.
    Caveat: The fitted predictive model explains only part of the variance, and findings depend on the specific benchmarks and model families studied.
  - **Mixed**, filed under [independent_workers](/patterns/independent_workers.md) — The authors report that independent multi-agent setups without centralized verification propagate and amplify errors far more than centralized coordination, and that multi-agent gains vanish or reverse once the single-agent baseline is already strong.
    Compared against: Single-agent systems and centralized, decentralized and hybrid multi-agent architectures. Domain: Agentic web browsing, finance, planning, workplace, software engineering and terminal tasks. Benchmarks: BrowseComp-Plus, Finance Agent, PlanCraft, WorkBench, SWE-bench Verified, Terminal-Bench.
    Caveat: The fitted predictive model explains only a modest share of performance variance, so the patterns are tendencies rather than guarantees.
  - **Mixed**, filed under [supervisor](/patterns/supervisor.md) — The authors report that centralized orchestration strongly helped on decomposable financial reasoning but degraded performance on sequential planning, while containing error amplification far better than independent agents at a large token overhead.
    Compared against: Single-agent systems and independent, decentralized and hybrid multi-agent architectures. Domain: Agentic benchmarks spanning finance, web browsing, planning, workplace, software and terminal tasks. Benchmarks: Finance-Agent, PlanCraft, BrowseComp-Plus, Workbench, SWE-bench Verified, Terminal-Bench.
    Caveat: Results depend on the chosen model families and benchmark set, and the fitted predictive model explains only part of the variance.
  - **Mixed**, filed under [mailbox_network](/patterns/mailbox_network.md) — The authors report that a decentralized peer-to-peer architecture gave a modest gain on web browsing and a large gain on financial reasoning but lost ground on sequential planning, with communication overhead well above a single agent.
    Compared against: A single agent and centralized, independent and hybrid multi-agent architectures. Domain: Agentic benchmarks spanning finance, web browsing and planning. Benchmarks: BrowseComp-Plus, Finance-Agent, PlanCraft.
    Caveat: Outcomes depend strongly on task structure and model family in this study.
  - **Mixed**, filed under the general findings of [/patterns/index.md](/patterns/index.md) — Kim and colleagues report that multi-agent coordination ranges from large gains on decomposable tasks to large losses on sequential planning, with diminishing or negative returns once the single agent is already strong and higher overhead on tool-heavy tasks.
    Compared against: A single agent under matched tools, prompts and compute. Domain: Agentic benchmarks including finance, planning and tool use.
    Caveat: Controlled study across a fixed set of architectures and model families; the predictive model explains only part of the variance.

### fan_out

- Source: [Self-Consistency Improves Chain of Thought Reasoning in Language Models (Wang et al.)](https://arxiv.org/abs/2203.11171), Xuezhi Wang et al., 2022-03-21.
  - **Helped**, filed under [fan_out](/patterns/fan_out.md) — The authors report that sampling diverse chain-of-thought reasoning paths and taking the most consistent answer by vote substantially improves accuracy over a single greedy chain.
    Compared against: Single greedy-decoded chain-of-thought. Domain: Arithmetic and commonsense reasoning. Benchmarks: GSM8K, SVAMP, AQuA, StrategyQA, ARC-Challenge.
    Caveat: Voting needs a short answer that can be matched across samples, so it does not directly apply to open-ended outputs.
- Source: [More Agents Is All You Need (Li et al.)](https://arxiv.org/abs/2402.05120), Junyou Li et al., 2024-02-03.
  - **Helped**, filed under [fan_out](/patterns/fan_out.md) — The authors report that a simple sample-and-vote method lets performance scale with the number of agents instantiated, with larger gains on harder tasks relative to the model.
    Compared against: Single LLM call and more elaborate prompting or multi-agent methods. Domain: Reasoning and code generation. Benchmarks: GSM8K, MATH, MMLU, Chess State Tracking, HumanEval.
    Caveat: The authors note token cost grows in proportion to agent count and gains taper off at the highest difficulty levels they tested.
- Source: [Large Language Monkeys: scaling inference compute with repeated sampling (Brown et al.)](https://arxiv.org/abs/2407.21787), Bradley Brown et al., 2024-07-31.
  - **Mixed**, filed under [fan_out](/patterns/fan_out.md) — The authors report that coverage from repeated sampling keeps rising with sample count and converts into real gains where answers can be automatically verified, but majority voting and reward models plateau and fail to keep pace where no verifier exists.
    Compared against: Single-sample attempts. Domain: Math, formal proofs, competitive programming and software engineering. Benchmarks: GSM8K, MATH, MiniF2F, CodeContests, SWE-bench Lite.
    Caveat: Most of the upside is coverage, which only becomes accuracy when a reliable automatic verifier can pick the correct sample.
- Source: [Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems](https://arxiv.org/abs/2403.02419), Chen et al., 2024-03-04.
  - **Mixed**, filed under [fan_out](/patterns/fan_out.md) — The authors report that accuracy of majority-vote and filter-then-vote systems can first increase and then decrease as the number of LLM calls grows, because extra calls help easy queries but hurt hard ones.
    Compared against: Vote and Filter-Vote systems at smaller call counts. Domain: Multiple-choice question answering and fact verification. Benchmarks: MMLU Physics, TruthfulQA, GPQA, AVeriTeC.
    Caveat: The analysis covers only simple voting aggregators on multiple-choice style tasks, not richer selection or verification schemes.
  - **Mixed**, filed under [council](/patterns/council.md) — The authors report that majority-vote systems can first improve and then degrade as more model calls are added, because extra calls help on easy queries but hurt on hard ones.
    Compared against: Voting systems with fewer model calls. Domain: Language tasks aggregated by majority vote.
    Caveat: Analyzes simple vote and filter-vote designs rather than richer judge panels or deliberating councils.
- Source: [Training Verifiers to Solve Math Word Problems (Cobbe et al., OpenAI)](https://arxiv.org/abs/2110.14168), Karl Cobbe et al., 2021-10-27.
  - **Helped**, filed under [fan_out](/patterns/fan_out.md) — The authors report that generating many candidate solutions and selecting the one ranked highest by a trained verifier substantially improves accuracy and scales better with data than fine-tuning alone.
    Compared against: Fine-tuned model producing a single answer. Domain: Grade-school math word problems. Benchmarks: GSM8K.
    Caveat: The gain depends on training a task-specific verifier, which requires labelled correct and incorrect solutions.

### map_reduce

- Source: [Skeleton-of-Thought: prompting LLMs for efficient parallel generation (Ning et al.)](https://arxiv.org/abs/2307.15337), Xuefei Ning et al., 2023-07-28.
  - **Mixed**, filed under [map_reduce](/patterns/map_reduce.md) — The authors report that first drafting an answer skeleton and then expanding each point in parallel gives considerable speed-ups across many LLMs and can improve quality on some question categories.
    Compared against: Standard sequential decoding. Domain: Open-ended question answering. Benchmarks: Vicuna-80, WizardLM.
    Caveat: The authors report weak results on math, coding, writing and estimation questions where later points depend on earlier ones.

### lane_swarm

- Source: [Scaling LLM-based multi-agent collaboration, MacNet (Qian et al.)](https://arxiv.org/abs/2406.07155), Chen Qian et al., 2024-06-11.
  - **Mixed**, filed under [lane_swarm](/patterns/lane_swarm.md) — The authors report that organizing up to over a thousand agents in a directed acyclic graph yields performance that grows logistically with agent count, with irregular topologies outperforming regular ones.
    Compared against: Smaller agent networks and regular topologies such as chains and meshes. Domain: Reasoning, code generation, software development and constrained text generation. Benchmarks: MMLU, HumanEval, SRDD, CommonGen-Hard.
    Caveat: The authors report most topologies saturate at around a hundred agents, that dense interaction can overload agents, and that context cost grows quadratically without their memory control.
- Source: [AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors](https://arxiv.org/abs/2308.10848), Chen et al., 2023-08-21.
  - **Mixed**, filed under [lane_swarm](/patterns/lane_swarm.md) — The authors report that dynamically composed agent groups can outperform a single agent, but also document cases where group discussion hurt a weaker model and negative emergent behaviors such as destructive actions.
    Compared against: Single-agent solo setups. Domain: Reasoning, coding, tool use and embodied Minecraft tasks. Benchmarks: FED, Commongen-Challenge, MGSM, Logic Grid Puzzles, HumanEval.
    Caveat: Groups in these experiments are small, so the results say little about large swarms.
  - **Mixed**, filed under [dynamic_spawning](/patterns/dynamic_spawning.md) — AgentVerse reports that dynamically adjusting group composition by recruiting expert agents lets multi-agent groups outperform a single agent, while also documenting negative emergent social behaviors.
    Compared against: A single agent. Domain: Text understanding, reasoning, coding, tool use and embodied tasks.
    Caveat: Author-run evaluation of the proposed framework; the negative behaviors are described qualitatively.

### supervisor

- Source: [Why Do Multi-Agent LLM Systems Fail?](https://arxiv.org/abs/2503.13657), Cemri et al., 2025-03-17.
  - **Hurt**, filed under [supervisor](/patterns/supervisor.md) — The authors report that popular multi-agent frameworks, including orchestrator-led ones, fail often, with failures clustering into system design, inter-agent misalignment and task verification problems rather than only underlying model weakness.
    Compared against: Expected task success of the same frameworks. Domain: Coding, math and general agent tasks across several open-source multi-agent frameworks. Benchmarks: MAST-Data.
    Caveat: A failure taxonomy built from annotated traces, not a controlled comparison of topologies against a single-agent baseline.
  - **No clear gain**, filed under [implement_review](/patterns/implement_review.md) — Across annotated traces from popular multi-agent frameworks, the authors identify task verification failures, such as missing or incorrect checking, as one of the main categories of multi-agent breakdowns, alongside design issues and inter-agent misalignment.
    Compared against: Expectations of benefit from multi-agent frameworks with reviewer or verifier roles. Domain: Coding, math and general agent tasks.
    Caveat: This is a failure taxonomy rather than a controlled comparison, so it shows how reviewer roles fail but not how often they help.
  - **No clear gain**, filed under the general findings of [/patterns/index.md](/patterns/index.md) — Cemri and colleagues note that multi-agent performance gains on popular benchmarks are often minimal and build a failure taxonomy spanning system design, inter-agent misalignment and task verification.
    Compared against: Single-agent and simpler baselines on popular benchmarks. Domain: Multi-agent frameworks across coding, math and general tasks. Benchmarks: MAST-Data.
    Caveat: The taxonomy characterizes failures rather than measuring a single head-to-head gain.

### planner_worker

- Source: [ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models](https://arxiv.org/abs/2305.18323), Xu et al., 2023-05-23.
  - **Helped**, filed under [planner_worker](/patterns/planner_worker.md) — The authors report that decoupling up-front planning from tool observations, with separate planner, worker and solver roles, used several times fewer tokens and slightly improved accuracy over interleaved reason-and-act prompting.
    Compared against: Interleaved observation-dependent reasoning such as ReAct. Domain: Multi-step tool-augmented question answering. Benchmarks: HotpotQA.
    Caveat: Gains were shown with older models on a small set of QA benchmarks, and a fixed up-front plan cannot adapt to surprising observations.
- Source: [Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models](https://arxiv.org/abs/2305.04091), Wang et al., 2023-05-06.
  - **Helped**, filed under [planner_worker](/patterns/planner_worker.md) — The authors report that prompting a model to first devise a plan and then carry out subtasks substantially outperformed zero-shot chain-of-thought and approached few-shot chain-of-thought on math reasoning.
    Compared against: Zero-shot and few-shot chain-of-thought prompting. Domain: Arithmetic, commonsense and symbolic reasoning.
    Caveat: This is a single-model prompting technique, not separate planner and executor agents, evaluated on an older model.

### role_pipeline

- Source: [On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents](https://arxiv.org/abs/2408.00989), Huang et al., 2024-08-02.
  - **Hurt**, filed under [role_pipeline](/patterns/role_pipeline.md) — The authors report that one-way linear chains of agents lost the most performance when faulty agents were injected, while a structure in which a leader directs two peers that talk to each other lost the least.
    Compared against: Flat and hierarchical multi-agent structures under the same injected faults. Domain: Code generation, math, translation and text evaluation. Benchmarks: HumanEval, CIAR, CommonMT, FairEval.
    Caveat: Structures were represented by a handful of existing systems that differ in more than topology, and robustness under injected errors is not the same as baseline accuracy.
  - **Helped**, filed under [hierarchical_delegation](/patterns/hierarchical_delegation.md) — The authors report that a mixed hierarchical structure lost the least performance when faulty agents were injected, while one-way linear pipelines such as MetaGPT-style chains lost the most.
    Compared against: Linear and flat multi-agent structures. Domain: Code generation, math, translation and text evaluation. Benchmarks: HumanEval, CIAR, CommonMT, FairEval.
    Caveat: Structures were represented by a handful of existing systems that differ in more than topology, and robustness under injected errors is not the same as baseline accuracy.

### hierarchical_delegation

- Source: [Talk Structurally, Act Hierarchically: A Collaborative Framework for LLM Multi-Agent Systems](https://arxiv.org/abs/2502.11098), Wang et al., 2025-02-16.
  - **Helped**, filed under [hierarchical_delegation](/patterns/hierarchical_delegation.md) — The authors report that a supervisor-led hierarchy with a nested evaluation team and structured messages outperformed strong single-model and multi-agent baselines across their tasks.
    Compared against: Strong reasoning models and multi-agent frameworks such as AgentVerse. Domain: Question answering and advertisement text generation. Benchmarks: MMLU, WikiQA, Camera.
    Caveat: The authors themselves flag a very high API cost for the experiments.
- Source: [On the Reliability Limits of LLM-Based Multi-Agent Planning](https://arxiv.org/abs/2603.26993), Ao, Gao, Simchi-Levi, 2026-03-27.
  - **Hurt**, filed under [hierarchical_delegation](/patterns/hierarchical_delegation.md) — The authors argue that multi-stage LLM planning networks that pass limited language messages lose information at each hand-off and, absent new external signals, cannot beat a single centralized decision-maker with the same information.
    Compared against: A single centralized decision-maker with equivalent information access. Domain: Multi-stage planning and decision making.
    Caveat: Primarily a theoretical technical note with small controlled experiments rather than a large empirical benchmark study.

### blackboard

- Source: [Exploring Advanced LLM Multi-Agent Systems Based on Blackboard Architecture](https://arxiv.org/abs/2507.01701), Han and Zhang, 2025-07-02.
  - **Helped**, filed under [blackboard](/patterns/blackboard.md) — The authors report that a blackboard system in which agents read and write a shared public space, with agents selected based on its contents, matched or beat static and autonomous multi-agent systems on average while using comparatively few tokens.
    Compared against: Chain-of-thought, static multi-agent systems and autonomous multi-agent systems such as GPTSwarm and AFlow. Domain: Knowledge, scientific, symbolic and math reasoning. Benchmarks: MMLU, ARC-Challenge, GPQA-Diamond, BBH, MATH, GSM8K.
    Caveat: The authors note limited agent types without tool use and few benchmarks, and 'competitive' is not a consistent win.

### mailbox_network

- Source: [Improving Multi-Agent Debate with Sparse Communication Topology](https://arxiv.org/abs/2406.11776), Li et al., 2024-06-17.
  - **Helped**, filed under [mailbox_network](/patterns/mailbox_network.md) — The authors report that multi-agent debate over sparse, ring-like communication graphs matched or beat fully connected debate while substantially cutting cost.
    Compared against: Fully connected multi-agent debate. Domain: Math reasoning, multimodal reasoning and alignment labeling. Benchmarks: MATH, GSM8K, MathVista, Anthropic-HH.
    Caveat: Studied only within debate-style answer refinement, with few agents and a small set of benchmarks.
- Source: [Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems](https://arxiv.org/abs/2410.02506), Zhang et al., 2024-10-03.
  - **Helped**, filed under [mailbox_network](/patterns/mailbox_network.md) — The authors report substantial communication redundancy in multi-agent message-passing graphs and that pruning it kept performance comparable while greatly reducing token cost.
    Compared against: Chain, tree, star, complete, layered and random topologies and frameworks such as AutoGen and GPTSwarm. Domain: General, math reasoning and code generation. Benchmarks: MMLU, GSM8K, MultiArith, SVAMP, AQuA, HumanEval.
    Caveat: Evaluated on short-form reasoning and coding benchmarks by the proposing authors.
- Source: [Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs](https://arxiv.org/abs/2311.17371), Smit et al., 2023-11-29.
  - **No clear gain**, filed under [mailbox_network](/patterns/mailbox_network.md) — The authors report that multi-agent debate systems in their current form did not reliably outperform simpler strategies such as self-consistency and ensembling, although some improved after hyperparameter tuning.
    Compared against: Self-consistency and ensembling over multiple reasoning paths. Domain: Medical and general reasoning question answering. Benchmarks: MedQA, PubMedQA, MMLU, CosmosQA, CIAR, GPQA.
    Caveat: Results are sensitive to hyperparameters such as agent agreement, so conclusions may shift with tuning.
  - **No clear gain**, filed under [debate](/patterns/debate.md) — The authors report that multi-agent debate systems in their current form did not reliably outperform self-consistency and multi-path ensembling, though some became competitive after hyperparameter tuning.
    Compared against: Self-consistency and ensembling over multiple reasoning paths. Domain: Question answering across popular research datasets.
    Caveat: Debate methods proved sensitive to hyperparameters, so conclusions may shift with tuning effort.
  - **No clear gain**, filed under the general findings of [/patterns/index.md](/patterns/index.md) — Smit and colleagues report that multi-agent debate does not reliably outperform self-consistency and ensembling and is more sensitive to hyperparameters.
    Compared against: Self-consistency and ensembling prompting strategies. Domain: Question answering including medical.
    Caveat: Tuning agent agreement levels let some debate systems surpass other protocols, so results depend on configuration.
- Source: [Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View](https://arxiv.org/abs/2310.02124), Zhang et al., 2023-10-03.
  - **Mixed**, filed under [mailbox_network](/patterns/mailbox_network.md) — The authors report that peer collaboration gains plateaued beyond a small number of agents and rounds, and that agents increasingly conformed to each other over rounds, sometimes converging on wrong answers.
    Compared against: Different numbers of agents, rounds and debate or reflection strategies. Domain: Knowledge, math and state-tracking reasoning. Benchmarks: MMLU, MATH, Chess Move Validity.
    Caveat: Small question samples per dataset and a single model family.

### critic_loop

- Source: [Self-Refine: Iterative Refinement with Self-Feedback](https://arxiv.org/abs/2303.17651), Madaan et al., 2023-03-30.
  - **Helped**, filed under [critic_loop](/patterns/critic_loop.md) — The authors report that iterative self-feedback and refinement by the same model improved outputs over one-step generation across a diverse set of tasks, as judged by humans and automatic metrics.
    Compared against: One-step generation with the same model. Domain: Dialogue, code optimization, math reasoning and other generation tasks. Benchmarks: GSM8K.
    Caveat: Gains were concentrated in open-ended generation tasks, and later work found intrinsic self-correction on reasoning tasks to be much weaker.
- Source: [Reflexion: Language Agents with Verbal Reinforcement Learning](https://arxiv.org/abs/2303.11366), Shinn et al., 2023-03-20.
  - **Helped**, filed under [critic_loop](/patterns/critic_loop.md) — The authors report that agents reflecting verbally on task feedback and storing reflections in memory improved significantly over a baseline agent on sequential decision-making, coding and reasoning tasks.
    Compared against: The same agent without reflection. Domain: Sequential decision-making, coding and language reasoning. Benchmarks: HumanEval, ALFWorld, HotpotQA.
    Caveat: Improvements rely on external feedback signals such as unit tests or environment rewards across multiple trials, not on self-critique alone.
- Source: [CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing](https://arxiv.org/abs/2305.11738), Gou et al., 2023-05-19.
  - **Helped**, filed under [critic_loop](/patterns/critic_loop.md) — The authors report that letting a model verify and revise its output using external tools such as search and code interpreters consistently improved performance, and they stress that external feedback is crucial for self-improvement.
    Compared against: The same model without tool-interactive critiquing. Domain: Free-form question answering, math program synthesis and toxicity reduction. Benchmarks: TriviaQA, HotpotQA, GSM8K.
    Caveat: The benefit depends on reliable tool feedback, so it does not show that critique without external grounding works.
- Source: [Large Language Models Cannot Self-Correct Reasoning Yet](https://arxiv.org/abs/2310.01798), Huang et al., 2023-10-03.
  - **Hurt**, filed under [critic_loop](/patterns/critic_loop.md) — The authors report that models struggled to self-correct their reasoning without external feedback, and that performance sometimes degraded after self-correction.
    Compared against: The model's initial answers before intrinsic self-correction. Domain: Math, commonsense and multi-hop question answering. Benchmarks: GSM8K, CommonSenseQA, HotpotQA.
    Caveat: Covers intrinsic self-correction with earlier models on reasoning tasks, and does not address correction driven by external feedback.
- Source: [On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks](https://arxiv.org/abs/2402.08115), Stechly et al., 2024-02-12.
  - **Hurt**, filed under [critic_loop](/patterns/critic_loop.md) — The authors report significant performance collapse when a model critiqued its own answers, contrasted with significant gains when a sound external verifier checked solutions.
    Compared against: Iterative prompting with a sound external verifier and one-shot generation. Domain: Reasoning and planning puzzles. Benchmarks: Game of 24, Graph Coloring, STRIPS planning.
    Caveat: Evaluated a single model on puzzle-like domains where exact verifiers exist, which may not represent open-ended tasks.

### council

- Source: [Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?](https://arxiv.org/abs/2502.00674), Li et al., 2025-02-02.
  - **Mixed**, filed under [council](/patterns/council.md) — The authors report that aggregating multiple outputs from only the single best model outperformed the standard mixture of different models in many scenarios, because mixing lowered average quality.
    Compared against: Aggregating several outputs of the single best model. Domain: Instruction following, knowledge, code reasoning and math. Benchmarks: AlpacaEval 2.0, MMLU, CRUX, MATH.
    Caveat: The authors also identify scenarios where mixing different models helps, so the result is about diversity versus quality rather than ensembling in general.

### debate

- Source: [Improving Factuality and Reasoning in Language Models through Multiagent Debate](https://arxiv.org/abs/2305.14325), Du et al., 2023-05-23.
  - **Helped**, filed under [debate](/patterns/debate.md) — The authors report that multiple model instances debating over several rounds improved mathematical and strategic reasoning and reduced hallucinated facts compared with a single model.
    Compared against: A single model instance and single-model reflection. Domain: Arithmetic, grade-school math, strategy and factual biographies. Benchmarks: GSM8K, MMLU.
    Caveat: Uses earlier models and does not match compute against self-consistency sampling with the same number of calls.
- Source: [Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate](https://arxiv.org/abs/2305.19118), Liang et al., 2023-05-30.
  - **Helped**, filed under [debate](/patterns/debate.md) — The authors report that a judged tit-for-tat debate between agents countered the degeneration of thought seen in self-reflection and improved results, while noting a judge may be unfair when agents use different models.
    Compared against: Self-reflection with a single model. Domain: Commonsense machine translation and counter-intuitive arithmetic reasoning. Benchmarks: CIAR.
    Caveat: Evaluated on two narrow challenge datasets and sensitive to debate settings such as the level of disagreement.
- Source: [Debating with More Persuasive LLMs Leads to More Truthful Answers](https://arxiv.org/abs/2402.06782), Khan et al., 2024-02-09.
  - **Helped**, filed under [debate](/patterns/debate.md) — The authors report that debate between expert models consistently helped weaker model judges and humans pick correct answers, and that more persuasive debaters further improved judge accuracy.
    Compared against: Naive baselines without debate, such as a single consultant. Domain: Reading comprehension with information asymmetry. Benchmarks: QuALITY.
    Caveat: Tests a scalable-oversight setting with hidden information, not general accuracy improvement for a solver.
- Source: [Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?](https://arxiv.org/abs/2508.17536), Choi et al., 2025-08-24.
  - **No clear gain**, filed under [debate](/patterns/debate.md) — The authors report that majority voting alone accounts for most of the gains usually attributed to multi-agent debate, and argue theoretically that debate by itself does not improve expected correctness.
    Compared against: Majority voting over independent agent answers. Domain: Natural language processing benchmarks.
    Caveat: The theoretical model idealizes the debate process, and the authors show targeted interventions can make debate more effective.

### dynamic_spawning

- Source: [AutoAgents: A Framework for Automatic Agent Generation](https://arxiv.org/abs/2309.17288), Chen et al., 2023-09-29.
  - **Helped**, filed under [dynamic_spawning](/patterns/dynamic_spawning.md) — AutoAgents reports that generating task-specific agents at runtime, with an observer reviewing plans, yields more coherent and accurate solutions than existing multi-agent methods.
    Compared against: Existing multi-agent methods with predefined agents. Domain: Open-ended question answering and creative writing.
    Caveat: Author-run comparison; the abstract gives no cost comparison against simpler single-agent baselines.

### coordinator_election

- Source: [AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs](https://arxiv.org/abs/2507.08616), Grötschla et al., 2025-07-11.
  - **Mixed**, filed under [coordinator_election](/patterns/coordinator_election.md) — AgentsNet reports that frontier models can solve coordination problems including leader election in small networks, but performance falls off sharply as network size grows.
    Compared against: Different model families and network sizes. Domain: Distributed coordination problems drawn from graph theory. Benchmarks: AgentsNet.
    Caveat: Synthetic graph tasks may not reflect leader selection in practical agent systems.
- Source: [Leadership as Coordination Control in Multi-Agent LLM Teams](https://arxiv.org/abs/2606.19111), Haewoon Kwak, 2026-06-17.
  - **No clear gain**, filed under [coordinator_election](/patterns/coordinator_election.md) — Kwak reports that no leadership controller dominated on accuracy, with transactional control roughly matching a simple initial vote, and advantages appearing only under narrow conditions.
    Compared against: A shared initial majority vote without a leader. Domain: Multi-agent LLM team reasoning.
    Caveat: A recent single-author preprint; the conditions under which leaders help are specific and measured on limited task regimes.
- Source: [An Electoral Approach to Diversify LLM-based Multi-Agent Collective Decision-Making](https://arxiv.org/abs/2410.15168), Zhao, Wang, Peng, 2024-10-19.
  - **Mixed**, filed under [coordinator_election](/patterns/coordinator_election.md) — Zhao and colleagues report that most surveyed LLM multi-agent systems rely on dictatorial or plurality decision rules, and that alternative ordinal voting methods can improve reasoning and robustness for some leading models.
    Compared against: Dictatorial and plurality collective decision rules. Domain: Reasoning benchmarks.
    Caveat: Studies collective voting rather than leader election, and gains held only for some models.

### adaptive_routing

- Source: [RouteLLM: Learning to Route LLMs with Preference Data](https://arxiv.org/abs/2406.18665), Ong et al., 2024-06-26.
  - **Helped**, filed under [adaptive_routing](/patterns/adaptive_routing.md) — RouteLLM reports that learned routers choosing between a strong and a weak model substantially reduced cost without compromising response quality and transferred when the model pair changed.
    Compared against: Always using the strong model. Domain: General chat, knowledge and math questions. Benchmarks: MT Bench, MMLU, GSM8K.
    Caveat: Routes between models for single queries rather than between agents, and gains vary by benchmark.
- Source: [FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance](https://arxiv.org/abs/2305.05176), Chen, Zaharia, Zou, 2023-05-09.
  - **Helped**, filed under [adaptive_routing](/patterns/adaptive_routing.md) — FrugalGPT reports that a learned cascade over LLM APIs could match the best single model at a small fraction of the cost or improve accuracy at equal cost.
    Compared against: The best individual LLM API. Domain: Classification and question answering tasks.
    Caveat: Evaluated on older model APIs and pricing, and cascades are a sequential form of routing.
- Source: [MasRouter: Learning to Route LLMs for Multi-Agent Systems](https://arxiv.org/abs/2502.11133), Yue et al., 2025-02-16.
  - **Helped**, filed under [adaptive_routing](/patterns/adaptive_routing.md) — MasRouter reports that jointly routing collaboration mode, roles and LLMs in a multi-agent system improved accuracy and reduced overhead versus prior methods.
    Compared against: Prior multi-agent routing and system design methods. Domain: Code generation, math and reasoning. Benchmarks: MBPP, HumanEval.
    Caveat: Author-run comparison with modest accuracy gains.

### tree_search

- Source: [Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models](https://arxiv.org/abs/2310.04406), Zhou et al., 2023-10-06.
  - **Helped**, filed under [tree_search](/patterns/tree_search.md) — The authors report that Monte Carlo tree search over an agent’s actions, with a language-model value function and self-reflection, outperformed acting and reflection baselines and earlier tree-search methods on programming, question answering and web shopping.
    Compared against: ReAct, Reflexion, and the tree-search methods Tree of Thoughts and RAP. Domain: Programming, interactive question answering, web shopping and math puzzles. Benchmarks: HumanEval, MBPP, HotpotQA, WebShop, Game of 24.
    Caveat: It assumes the environment can be reverted to an earlier state, and it costs more than a single acting agent.
- Source: [Tree of Thoughts: Deliberate Problem Solving with Large Language Models](https://arxiv.org/abs/2305.10601), Yao et al., 2023-05-17.
  - **Helped**, filed under [tree_search](/patterns/tree_search.md) — The authors report that letting a model propose, evaluate and backtrack over intermediate thoughts solved far more planning puzzles than chain-of-thought prompting of the same model.
    Compared against: Input-output and chain-of-thought prompting of the same model. Domain: Puzzles and constrained creative writing that need planning or search. Benchmarks: Game of 24, Creative Writing, Mini Crosswords.
    Caveat: One model plays every role, on a few small tasks built by the authors, at many times the cost of a single chain of thought.
- Source: [Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](https://arxiv.org/abs/2408.03314), Snell, Lee, Xu, Kumar, 2024-08-06.
  - **Mixed**, filed under [tree_search](/patterns/tree_search.md) — The authors report that beam search guided by a process reward model beat best-of-N sampling on harder questions and at small budgets, but degraded on easy questions as the budget grew, because the search exploited the verifier’s errors.
    Compared against: Best-of-N sampling scored by the same verifier. Domain: Competition mathematics. Benchmarks: MATH.
    Caveat: One model family on one math benchmark, and a search over reasoning steps within one model rather than between agents.

### architecture_search

- Source: [Automated Design of Agentic Systems](https://arxiv.org/abs/2408.08435), Hu, Lu, Clune, 2024-08-15.
  - **Helped**, filed under [architecture_search](/patterns/architecture_search.md) — ADAS reports that agents discovered by a meta agent programming new designs in code outperformed hand-designed agents and kept their advantage when transferred across domains and models.
    Compared against: State-of-the-art hand-designed agents. Domain: Reading comprehension, math, science and coding. Benchmarks: DROP, MGSM.
    Caveat: Search cost is not weighed against the gains in the abstract, and later work questions its cost-effectiveness.
- Source: [AFlow: Automating Agentic Workflow Generation](https://arxiv.org/abs/2410.10762), Zhang et al., 2024-10-14.
  - **Helped**, filed under [architecture_search](/patterns/architecture_search.md) — AFlow reports that Monte Carlo tree search over code-represented workflows beat prior baselines and let smaller models outperform a much larger model on some tasks at a fraction of its inference cost.
    Compared against: Manually designed workflows and prior automated methods. Domain: Question answering, code generation and math.
    Caveat: Author-run evaluation; the cheaper-model result holds only on specific tasks.
- Source: [Multi-agent Architecture Search via Agentic Supernet](https://arxiv.org/abs/2502.04180), Zhang et al., 2025-02-06.
  - **Helped**, filed under [architecture_search](/patterns/architecture_search.md) — MaAS reports that sampling query-dependent agent architectures from a learned supernet matched or beat existing systems while using only a fraction of their inference cost, with cross-dataset and cross-model transfer.
    Compared against: Handcrafted and automated multi-agent systems. Domain: Math, coding and tool use.
    Caveat: Author-run comparison; some reported accuracy gains are small.
- Source: [Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies](https://arxiv.org/abs/2502.02533), Zhou et al., 2025-02-04.
  - **Mixed**, filed under [architecture_search](/patterns/architecture_search.md) — MASS reports that only a small fraction of topologies improved performance, with most failing to help or degrading it, and that prompt optimization was more token-effective than scaling agent count.
    Compared against: Unoptimized topologies and agent-scaling strategies such as self-consistency and debate. Domain: Reasoning, multi-hop question answering and coding. Benchmarks: MATH, DROP, HotpotQA, MuSiQue, 2WikiMQA, MBPP, HumanEval, LiveCodeBench.
    Caveat: Findings come from one search framework and a limited set of model backbones.

[Reporting outcomes (limited rollout)](/docs/api/contributing.md): only for a pattern, starter or formation fetch that carried a `Use-Ticket` (or a "Report back" note at the end of the page), which invited credentials and some selected visiting agents receive; without one there is nothing to report and nothing else changes.
