---
schema: formation.domain/v0.1
kind: domain
visibility: public
canonical_url: https://topologyindex.com/domains/reasoning.md
community_ranking: null
description: 'Task shapes common in reasoning, the published findings tagged with the domain and the starters in it. Hypotheses and attributed findings, never a ranking.'
domain: reasoning
domain_index: /domains/index.md
evaluable_here: []
findings:
  - citations:
      - compared_against: 'Multi-agent discussion frameworks using the same backbone models'
        direction: no_clear_gain
        pattern: single_agent
      - compared_against: 'A single agent with strong prompts and demonstrations'
        direction: no_clear_gain
        pattern: mailbox_network
      - compared_against: 'A single agent with strong prompts and demonstrations'
        direction: no_clear_gain
        pattern: null
    source_id: arxiv:2402.18272
    url: https://arxiv.org/abs/2402.18272
  - citations:
      - compared_against: 'Multi-agent debate methods'
        direction: no_clear_gain
        pattern: single_agent
      - compared_against: 'Single-agent chain-of-thought and self-consistency'
        direction: no_clear_gain
        pattern: debate
    source_id: arxiv:2502.08788
    url: https://arxiv.org/abs/2502.08788
  - citations:
      - compared_against: 'Multi-agent coordination under matched tools, prompts and compute'
        direction: mixed
        pattern: single_agent
      - compared_against: 'Single-agent systems and centralized, decentralized and hybrid multi-agent architectures'
        direction: mixed
        pattern: independent_workers
      - compared_against: 'Single-agent systems and independent, decentralized and hybrid multi-agent architectures'
        direction: mixed
        pattern: supervisor
      - compared_against: 'A single agent and centralized, independent and hybrid multi-agent architectures'
        direction: mixed
        pattern: mailbox_network
      - compared_against: 'A single agent under matched tools, prompts and compute'
        direction: mixed
        pattern: null
    source_id: arxiv:2512.08296
    url: https://arxiv.org/abs/2512.08296
  - citations:
      - compared_against: 'Single greedy-decoded chain-of-thought'
        direction: helped
        pattern: fan_out
    source_id: arxiv:2203.11171
    url: https://arxiv.org/abs/2203.11171
  - citations:
      - compared_against: 'Single LLM call and more elaborate prompting or multi-agent methods'
        direction: helped
        pattern: fan_out
    source_id: arxiv:2402.05120
    url: https://arxiv.org/abs/2402.05120
  - citations:
      - compared_against: 'Single-sample attempts'
        direction: mixed
        pattern: fan_out
    source_id: arxiv:2407.21787
    url: https://arxiv.org/abs/2407.21787
  - citations:
      - compared_against: 'Vote and Filter-Vote systems at smaller call counts'
        direction: mixed
        pattern: fan_out
      - compared_against: 'Voting systems with fewer model calls'
        direction: mixed
        pattern: council
    source_id: arxiv:2403.02419
    url: https://arxiv.org/abs/2403.02419
  - citations:
      - compared_against: 'Fine-tuned model producing a single answer'
        direction: helped
        pattern: fan_out
    source_id: arxiv:2110.14168
    url: https://arxiv.org/abs/2110.14168
  - citations:
      - compared_against: 'Standard sequential decoding'
        direction: mixed
        pattern: map_reduce
    source_id: arxiv:2307.15337
    url: https://arxiv.org/abs/2307.15337
  - citations:
      - compared_against: 'Smaller agent networks and regular topologies such as chains and meshes'
        direction: mixed
        pattern: lane_swarm
    source_id: arxiv:2406.07155
    url: https://arxiv.org/abs/2406.07155
  - citations:
      - compared_against: 'Single-agent solo setups'
        direction: mixed
        pattern: lane_swarm
      - compared_against: 'A single agent'
        direction: mixed
        pattern: dynamic_spawning
    source_id: arxiv:2308.10848
    url: https://arxiv.org/abs/2308.10848
  - citations:
      - compared_against: 'Expected task success of the same frameworks'
        direction: hurt
        pattern: supervisor
      - compared_against: 'Expectations of benefit from multi-agent frameworks with reviewer or verifier roles'
        direction: no_clear_gain
        pattern: implement_review
      - compared_against: 'Single-agent and simpler baselines on popular benchmarks'
        direction: no_clear_gain
        pattern: null
    source_id: arxiv:2503.13657
    url: https://arxiv.org/abs/2503.13657
  - citations:
      - compared_against: 'Interleaved observation-dependent reasoning such as ReAct'
        direction: helped
        pattern: planner_worker
    source_id: arxiv:2305.18323
    url: https://arxiv.org/abs/2305.18323
  - citations:
      - compared_against: 'Zero-shot and few-shot chain-of-thought prompting'
        direction: helped
        pattern: planner_worker
    source_id: arxiv:2305.04091
    url: https://arxiv.org/abs/2305.04091
  - citations:
      - compared_against: 'Flat and hierarchical multi-agent structures under the same injected faults'
        direction: hurt
        pattern: role_pipeline
      - compared_against: 'Linear and flat multi-agent structures'
        direction: helped
        pattern: hierarchical_delegation
    source_id: arxiv:2408.00989
    url: https://arxiv.org/abs/2408.00989
  - citations:
      - compared_against: 'Strong reasoning models and multi-agent frameworks such as AgentVerse'
        direction: helped
        pattern: hierarchical_delegation
    source_id: arxiv:2502.11098
    url: https://arxiv.org/abs/2502.11098
  - citations:
      - compared_against: 'A single centralized decision-maker with equivalent information access'
        direction: hurt
        pattern: hierarchical_delegation
    source_id: arxiv:2603.26993
    url: https://arxiv.org/abs/2603.26993
  - citations:
      - compared_against: 'Chain-of-thought, static multi-agent systems and autonomous multi-agent systems such as GPTSwarm and AFlow'
        direction: helped
        pattern: blackboard
    source_id: arxiv:2507.01701
    url: https://arxiv.org/abs/2507.01701
  - citations:
      - compared_against: 'Fully connected multi-agent debate'
        direction: helped
        pattern: mailbox_network
    source_id: arxiv:2406.11776
    url: https://arxiv.org/abs/2406.11776
  - citations:
      - compared_against: 'Chain, tree, star, complete, layered and random topologies and frameworks such as AutoGen and GPTSwarm'
        direction: helped
        pattern: mailbox_network
    source_id: arxiv:2410.02506
    url: https://arxiv.org/abs/2410.02506
  - citations:
      - compared_against: 'Self-consistency and ensembling over multiple reasoning paths'
        direction: no_clear_gain
        pattern: mailbox_network
      - compared_against: 'Self-consistency and ensembling over multiple reasoning paths'
        direction: no_clear_gain
        pattern: debate
      - compared_against: 'Self-consistency and ensembling prompting strategies'
        direction: no_clear_gain
        pattern: null
    source_id: arxiv:2311.17371
    url: https://arxiv.org/abs/2311.17371
  - citations:
      - compared_against: 'Different numbers of agents, rounds and debate or reflection strategies'
        direction: mixed
        pattern: mailbox_network
    source_id: arxiv:2310.02124
    url: https://arxiv.org/abs/2310.02124
  - citations:
      - compared_against: 'One-step generation with the same model'
        direction: helped
        pattern: critic_loop
    source_id: arxiv:2303.17651
    url: https://arxiv.org/abs/2303.17651
  - citations:
      - compared_against: 'The same agent without reflection'
        direction: helped
        pattern: critic_loop
    source_id: arxiv:2303.11366
    url: https://arxiv.org/abs/2303.11366
  - citations:
      - compared_against: 'The same model without tool-interactive critiquing'
        direction: helped
        pattern: critic_loop
    source_id: arxiv:2305.11738
    url: https://arxiv.org/abs/2305.11738
  - citations:
      - compared_against: 'The model''s initial answers before intrinsic self-correction'
        direction: hurt
        pattern: critic_loop
    source_id: arxiv:2310.01798
    url: https://arxiv.org/abs/2310.01798
  - citations:
      - compared_against: 'Iterative prompting with a sound external verifier and one-shot generation'
        direction: hurt
        pattern: critic_loop
    source_id: arxiv:2402.08115
    url: https://arxiv.org/abs/2402.08115
  - citations:
      - compared_against: 'Aggregating several outputs of the single best model'
        direction: mixed
        pattern: council
    source_id: arxiv:2502.00674
    url: https://arxiv.org/abs/2502.00674
  - citations:
      - compared_against: 'A single model instance and single-model reflection'
        direction: helped
        pattern: debate
    source_id: arxiv:2305.14325
    url: https://arxiv.org/abs/2305.14325
  - citations:
      - compared_against: 'Self-reflection with a single model'
        direction: helped
        pattern: debate
    source_id: arxiv:2305.19118
    url: https://arxiv.org/abs/2305.19118
  - citations:
      - compared_against: 'Naive baselines without debate, such as a single consultant'
        direction: helped
        pattern: debate
    source_id: arxiv:2402.06782
    url: https://arxiv.org/abs/2402.06782
  - citations:
      - compared_against: 'Majority voting over independent agent answers'
        direction: no_clear_gain
        pattern: debate
    source_id: arxiv:2508.17536
    url: https://arxiv.org/abs/2508.17536
  - citations:
      - compared_against: 'Existing multi-agent methods with predefined agents'
        direction: helped
        pattern: dynamic_spawning
    source_id: arxiv:2309.17288
    url: https://arxiv.org/abs/2309.17288
  - citations:
      - compared_against: 'Different model families and network sizes'
        direction: mixed
        pattern: coordinator_election
    source_id: arxiv:2507.08616
    url: https://arxiv.org/abs/2507.08616
  - citations:
      - compared_against: 'A shared initial majority vote without a leader'
        direction: no_clear_gain
        pattern: coordinator_election
    source_id: arxiv:2606.19111
    url: https://arxiv.org/abs/2606.19111
  - citations:
      - compared_against: 'Dictatorial and plurality collective decision rules'
        direction: mixed
        pattern: coordinator_election
    source_id: arxiv:2410.15168
    url: https://arxiv.org/abs/2410.15168
  - citations:
      - compared_against: 'Always using the strong model'
        direction: helped
        pattern: adaptive_routing
    source_id: arxiv:2406.18665
    url: https://arxiv.org/abs/2406.18665
  - citations:
      - compared_against: 'The best individual LLM API'
        direction: helped
        pattern: adaptive_routing
    source_id: arxiv:2305.05176
    url: https://arxiv.org/abs/2305.05176
  - citations:
      - compared_against: 'Prior multi-agent routing and system design methods'
        direction: helped
        pattern: adaptive_routing
    source_id: arxiv:2502.11133
    url: https://arxiv.org/abs/2502.11133
  - citations:
      - compared_against: 'ReAct, Reflexion, and the tree-search methods Tree of Thoughts and RAP'
        direction: helped
        pattern: tree_search
    source_id: arxiv:2310.04406
    url: https://arxiv.org/abs/2310.04406
  - citations:
      - compared_against: 'Input-output and chain-of-thought prompting of the same model'
        direction: helped
        pattern: tree_search
    source_id: arxiv:2305.10601
    url: https://arxiv.org/abs/2305.10601
  - citations:
      - compared_against: 'Best-of-N sampling scored by the same verifier'
        direction: mixed
        pattern: tree_search
    source_id: arxiv:2408.03314
    url: https://arxiv.org/abs/2408.03314
  - citations:
      - compared_against: 'State-of-the-art hand-designed agents'
        direction: helped
        pattern: architecture_search
    source_id: arxiv:2408.08435
    url: https://arxiv.org/abs/2408.08435
  - citations:
      - compared_against: 'Manually designed workflows and prior automated methods'
        direction: helped
        pattern: architecture_search
    source_id: arxiv:2410.10762
    url: https://arxiv.org/abs/2410.10762
  - citations:
      - compared_against: 'Handcrafted and automated multi-agent systems'
        direction: helped
        pattern: architecture_search
    source_id: arxiv:2502.04180
    url: https://arxiv.org/abs/2502.04180
  - citations:
      - compared_against: 'Unoptimized topologies and agent-scaling strategies such as self-consistency and debate'
        direction: mixed
        pattern: architecture_search
    source_id: arxiv:2502.02533
    url: https://arxiv.org/abs/2502.02533
findings_page: /domains/reasoning/findings.md
path: /domains/reasoning.md
pattern_index: /patterns/index.md
product_api_version: v1
schema_version: v0.1
shapes:
  - example: 'a question one careful pass can answer'
    shape: small_or_single_owner
  - example: 'a proof or computation that can be verified'
    shape: easier_to_check_than_do
  - example: 'a multi-step problem where a critique can catch an error'
    shape: needs_rounds_of_criticism
  - example: 'a problem whose sampled answers differ, where agreement or a checker picks one'
    shape: fails_often_attempts_vary
starters:
  - name: reasoning-critic-loop
    path: /starters/reasoning-critic-loop/0.1.0.md
    task_classes:
      - reasoning.proof_review
    version: '0.1.0'
task_class_prefix: reasoning.
title: 'Which multi-agent pattern for reasoning tasks?'
---

# Which multi-agent pattern for reasoning tasks?

**Short answer:** for a question one careful pass can answer, start with [single_agent](/patterns/single_agent.md); avoid it when the task clearly exceeds one context window.
Other shapes of reasoning start as the table below says.
Hypotheses from the [decision guide](/patterns/index.md), not a ranking, and nothing here is
measured; [published findings](/domains/reasoning/findings.md) keep the unfavourable ones.

Math, logic, knowledge and question answering, where the answer comes from thinking rather than from acting on an environment.

The `reasoning` domain of the [task domains](/domains/index.md): task classes that start
with `reasoning.`.

## Task shapes common in this domain

Hypotheses about the work, each a row of the decision guide. Choose by the shape of your task,
not by the domain.

- `small_or_single_owner`: a question one careful pass can answer
- `easier_to_check_than_do`: a proof or computation that can be verified
- `needs_rounds_of_criticism`: a multi-step problem where a critique can catch an error
- `fails_often_attempts_vary`: a problem whose sampled answers differ, where agreement or a checker picks one

Where to start by the shape of the task. Every row is a hypothesis to test against a strong
single-agent configuration, not a ranking: nothing in this table has been measured here.

| If the task… | Start with | Consider next | Avoid when |
| --- | --- | --- | --- |
| is small, or has one clear owner | [single_agent](/patterns/single_agent.md) | [implement_review](/patterns/implement_review.md) | the task clearly exceeds one context window |
| is easier to check than to do | [implement_review](/patterns/implement_review.md) | [critic_loop](/patterns/critic_loop.md) | nothing outside the roles can validate the result |
| needs several rounds of criticism | [critic_loop](/patterns/critic_loop.md) | [council](/patterns/council.md) | rounds stop converging |
| often fails, but attempts vary | [fan_out](/patterns/fan_out.md) | [council](/patterns/council.md) | nothing can cheaply pick the winning attempt |

## Published findings in this domain

Typed in the `findings` frontmatter: each study tagged with this domain once, in pattern
vocabulary order (never by direction), with every pattern page that cites it, how that
pattern fared ([`direction`](/docs/schemas/pattern/v0.1.md)) and what it was compared with.
Unfavourable results are included on purpose; none of this is evidence produced here. Each
finding in words, with its caveat and source: [/domains/reasoning/findings.md](/domains/reasoning/findings.md).

## Starters

Unvalidated starting points that declare a task class in this domain; nobody has run them
here.

- [reasoning-critic-loop](/starters/reasoning-critic-loop/0.1.0.md): An author and a critic pass a proof or derivation back and forth until the bounded review cycle produces a checked draft. Task classes: `reasoning.proof_review`.

## What can be evaluated here

Nothing in this domain yet. This deployment evaluates only `coding.bugfix`; an empty domain is a valid state, not a gap to fill with claims.

## To find out for your workload

Nothing on this page says which arrangement will work for your task. The private
recommendation and evaluation routes compare complete configurations on your own workload;
the [integration guide](/docs/api/integration.md) says how to reach them.

[Reporting outcomes (limited rollout)](/docs/api/contributing.md): only for a pattern, starter or formation fetch that carried a `Use-Ticket` (or a "Report back" note at the end of the page), which invited credentials and some selected visiting agents receive; without one there is nothing to report and nothing else changes.
