---
schema: formation.domain_findings/v0.1
kind: domain_findings
visibility: public
canonical_url: https://topologyindex.com/domains/evaluation/findings.md
description: 'Each published study tagged with evaluation, once, with every finding filed under a pattern page: the sentence, comparison, domain, caveat and source. Attributed, stated without figures, never a ranking.'
domain: evaluation
domain_page: /domains/evaluation.md
findings:
  - citations:
      - compared_against: 'Flat and hierarchical multi-agent structures under the same injected faults'
        direction: hurt
        pattern: role_pipeline
      - compared_against: 'Linear and flat multi-agent structures'
        direction: helped
        pattern: hierarchical_delegation
    source_id: arxiv:2408.00989
    url: https://arxiv.org/abs/2408.00989
  - citations:
      - compared_against: 'Fully connected multi-agent debate'
        direction: helped
        pattern: mailbox_network
    source_id: arxiv:2406.11776
    url: https://arxiv.org/abs/2406.11776
  - citations:
      - compared_against: 'Human contractors reviewing code without assistance'
        direction: helped
        pattern: implement_review
    source_id: arxiv:2407.00215
    url: https://arxiv.org/abs/2407.00215
  - citations:
      - compared_against: 'A single large LLM judge'
        direction: helped
        pattern: council
    source_id: arxiv:2404.18796
    url: https://arxiv.org/abs/2404.18796
  - citations:
      - compared_against: 'Human expert and crowdsourced preference judgments'
        direction: mixed
        pattern: council
    source_id: arxiv:2306.05685
    url: https://arxiv.org/abs/2306.05685
  - citations:
      - compared_against: 'Human annotator judgments of equal-quality outputs'
        direction: hurt
        pattern: council
    source_id: arxiv:2404.13076
    url: https://arxiv.org/abs/2404.13076
path: /domains/evaluation/findings.md
pattern_index: /patterns/index.md
product_api_version: v1
schema_version: v0.1
title: 'Published findings on multi-agent patterns studied in evaluation'
---

# Published findings on multi-agent patterns studied in evaluation

The published findings of the [`evaluation` domain](/domains/evaluation.md), in words. That page
has the task shapes common in the domain, their rows of the [decision guide](/patterns/index.md)
and the starters. None of it has been measured here, and nothing on this page ranks patterns.

## Published findings in this domain

Attributed to each source and stated without figures. Each source is listed once, under the
first pattern in the vocabulary that cites it, with every pattern page that cites it for this
domain. The order is the pattern vocabulary, never the direction, and unfavourable results
are included on purpose. None of this is evidence produced here. Reviewed 2026-09-23.
Each label says how the pattern each finding is filed under fared against what it was compared with, as the source reports
it; the labels are defined at [/docs/schemas/pattern/v0.1.md](/docs/schemas/pattern/v0.1.md).

### role_pipeline

- Source: [On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents](https://arxiv.org/abs/2408.00989), Huang et al., 2024-08-02.
  - **Hurt**, filed under [role_pipeline](/patterns/role_pipeline.md) — The authors report that one-way linear chains of agents lost the most performance when faulty agents were injected, while a structure in which a leader directs two peers that talk to each other lost the least.
    Compared against: Flat and hierarchical multi-agent structures under the same injected faults. Domain: Code generation, math, translation and text evaluation. Benchmarks: HumanEval, CIAR, CommonMT, FairEval.
    Caveat: Structures were represented by a handful of existing systems that differ in more than topology, and robustness under injected errors is not the same as baseline accuracy.
  - **Helped**, filed under [hierarchical_delegation](/patterns/hierarchical_delegation.md) — The authors report that a mixed hierarchical structure lost the least performance when faulty agents were injected, while one-way linear pipelines such as MetaGPT-style chains lost the most.
    Compared against: Linear and flat multi-agent structures. Domain: Code generation, math, translation and text evaluation. Benchmarks: HumanEval, CIAR, CommonMT, FairEval.
    Caveat: Structures were represented by a handful of existing systems that differ in more than topology, and robustness under injected errors is not the same as baseline accuracy.

### mailbox_network

- Source: [Improving Multi-Agent Debate with Sparse Communication Topology](https://arxiv.org/abs/2406.11776), Li et al., 2024-06-17.
  - **Helped**, filed under [mailbox_network](/patterns/mailbox_network.md) — The authors report that multi-agent debate over sparse, ring-like communication graphs matched or beat fully connected debate while substantially cutting cost.
    Compared against: Fully connected multi-agent debate. Domain: Math reasoning, multimodal reasoning and alignment labeling. Benchmarks: MATH, GSM8K, MathVista, Anthropic-HH.
    Caveat: Studied only within debate-style answer refinement, with few agents and a small set of benchmarks.

### implement_review

- Source: [LLM Critics Help Catch LLM Bugs](https://arxiv.org/abs/2407.00215), McAleese et al. (OpenAI), 2024-06-28.
  - **Helped**, filed under [implement_review](/patterns/implement_review.md) — The authors report that trained LLM critics caught more bugs in model-written code than paid human contractors and their critiques were usually preferred, but critics also hallucinated bugs, and human-plus-critic teams hallucinated less.
    Compared against: Human contractors reviewing code without assistance. Domain: Code review of model-written code.
    Caveat: The critics were specially trained with reinforcement learning from human feedback rather than prompted, and results are self-reported by the vendor.

### council

- Source: [Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models](https://arxiv.org/abs/2404.18796), Verga et al., 2024-04-29.
  - **Helped**, filed under [council](/patterns/council.md) — The authors report that a panel of smaller judges from disjoint model families outperformed a single large judge, showed less intra-model bias, and cost considerably less.
    Compared against: A single large LLM judge. Domain: LLM output evaluation across question answering and chat settings.
    Caveat: Evaluated on a limited set of judge settings and datasets, and agreement with humans is the target metric rather than downstream task quality.
- Source: [Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena](https://arxiv.org/abs/2306.05685), Zheng et al., 2023-06-09.
  - **Mixed**, filed under [council](/patterns/council.md) — The authors report that strong LLM judges agreed with human preferences about as well as humans agree with each other, while documenting position, verbosity and self-enhancement biases.
    Compared against: Human expert and crowdsourced preference judgments. Domain: Chat assistant evaluation. Benchmarks: MT-Bench, Chatbot Arena.
    Caveat: Concerns single-judge setups with earlier models, and the identified biases are only partly mitigated by the proposed fixes.
- Source: [LLM Evaluators Recognize and Favor Their Own Generations](https://arxiv.org/abs/2404.13076), Panickssery et al., 2024-04-15.
  - **Hurt**, filed under [council](/patterns/council.md) — The authors report that LLM evaluators can recognize their own generations and that this self-recognition correlates with a bias toward scoring their own outputs higher than humans would.
    Compared against: Human annotator judgments of equal-quality outputs. Domain: Self-evaluation of summaries. Benchmarks: XSUM, CNN/DailyMail.
    Caveat: Studied on summarization with a small set of models, so the size of the bias in other domains is uncertain.

[Reporting outcomes (limited rollout)](/docs/api/contributing.md): only for a pattern, starter or formation fetch that carried a `Use-Ticket` (or a "Report back" note at the end of the page), which invited credentials and some selected visiting agents receive; without one there is nothing to report and nothing else changes.
