---
schema: formation.domain/v0.1
kind: domain
visibility: public
canonical_url: https://topologyindex.com/domains/evaluation.md
community_ranking: null
description: 'Task shapes common in evaluation, the published findings tagged with the domain and the starters in it. Hypotheses and attributed findings, never a ranking.'
domain: evaluation
domain_index: /domains/index.md
evaluable_here: []
findings:
  - citations:
      - compared_against: 'Flat and hierarchical multi-agent structures under the same injected faults'
        direction: hurt
        pattern: role_pipeline
      - compared_against: 'Linear and flat multi-agent structures'
        direction: helped
        pattern: hierarchical_delegation
    source_id: arxiv:2408.00989
    url: https://arxiv.org/abs/2408.00989
  - citations:
      - compared_against: 'Fully connected multi-agent debate'
        direction: helped
        pattern: mailbox_network
    source_id: arxiv:2406.11776
    url: https://arxiv.org/abs/2406.11776
  - citations:
      - compared_against: 'Human contractors reviewing code without assistance'
        direction: helped
        pattern: implement_review
    source_id: arxiv:2407.00215
    url: https://arxiv.org/abs/2407.00215
  - citations:
      - compared_against: 'A single large LLM judge'
        direction: helped
        pattern: council
    source_id: arxiv:2404.18796
    url: https://arxiv.org/abs/2404.18796
  - citations:
      - compared_against: 'Human expert and crowdsourced preference judgments'
        direction: mixed
        pattern: council
    source_id: arxiv:2306.05685
    url: https://arxiv.org/abs/2306.05685
  - citations:
      - compared_against: 'Human annotator judgments of equal-quality outputs'
        direction: hurt
        pattern: council
    source_id: arxiv:2404.13076
    url: https://arxiv.org/abs/2404.13076
findings_page: /domains/evaluation/findings.md
path: /domains/evaluation.md
pattern_index: /patterns/index.md
product_api_version: v1
schema_version: v0.1
shapes:
  - example: 'one judgment against a clear rubric'
    shape: small_or_single_owner
  - example: 'grading against a reference answer or a rubric'
    shape: easier_to_check_than_do
  - example: 'verdicts that vary between judges, where agreement across several picks one'
    shape: fails_often_attempts_vary
starters:
  - name: evaluation-fan-out
    path: /starters/evaluation-fan-out/0.1.0.md
    task_classes:
      - evaluation.rubric_judgment
    version: '0.1.0'
task_class_prefix: evaluation.
title: 'Which multi-agent pattern for evaluation and LLM-as-judge?'
---

# Which multi-agent pattern for evaluation and LLM-as-judge?

**Short answer:** for one judgment against a clear rubric, start with [single_agent](/patterns/single_agent.md); avoid it when the task clearly exceeds one context window.
Other shapes of evaluation start as the table below says.
Hypotheses from the [decision guide](/patterns/index.md), not a ranking, and nothing here is
measured; [published findings](/domains/evaluation/findings.md) keep the unfavourable ones.

Judging the output of models and agents: LLM-as-judge, graders and review panels.

The `evaluation` domain of the [task domains](/domains/index.md): task classes that start
with `evaluation.`.

## Task shapes common in this domain

Hypotheses about the work, each a row of the decision guide. Choose by the shape of your task,
not by the domain.

- `small_or_single_owner`: one judgment against a clear rubric
- `easier_to_check_than_do`: grading against a reference answer or a rubric
- `fails_often_attempts_vary`: verdicts that vary between judges, where agreement across several picks one

Where to start by the shape of the task. Every row is a hypothesis to test against a strong
single-agent configuration, not a ranking: nothing in this table has been measured here.

| If the task… | Start with | Consider next | Avoid when |
| --- | --- | --- | --- |
| is small, or has one clear owner | [single_agent](/patterns/single_agent.md) | [implement_review](/patterns/implement_review.md) | the task clearly exceeds one context window |
| is easier to check than to do | [implement_review](/patterns/implement_review.md) | [critic_loop](/patterns/critic_loop.md) | nothing outside the roles can validate the result |
| often fails, but attempts vary | [fan_out](/patterns/fan_out.md) | [council](/patterns/council.md) | nothing can cheaply pick the winning attempt |

## Published findings in this domain

Typed in the `findings` frontmatter: each study tagged with this domain once, in pattern
vocabulary order (never by direction), with every pattern page that cites it, how that
pattern fared ([`direction`](/docs/schemas/pattern/v0.1.md)) and what it was compared with.
Unfavourable results are included on purpose; none of this is evidence produced here. Each
finding in words, with its caveat and source: [/domains/evaluation/findings.md](/domains/evaluation/findings.md).

## Starters

Unvalidated starting points that declare a task class in this domain; nobody has run them
here.

- [evaluation-fan-out](/starters/evaluation-fan-out/0.1.0.md): Independent attempts answer a difficult evaluation question, and a selector applies the declared rubric to choose one result. Task classes: `evaluation.rubric_judgment`.

## What can be evaluated here

Nothing in this domain yet. This deployment evaluates only `coding.bugfix`; an empty domain is a valid state, not a gap to fill with claims.

## To find out for your workload

Nothing on this page says which arrangement will work for your task. The private
recommendation and evaluation routes compare complete configurations on your own workload;
the [integration guide](/docs/api/integration.md) says how to reach them.

[Reporting outcomes (limited rollout)](/docs/api/contributing.md): only for a pattern, starter or formation fetch that carried a `Use-Ticket` (or a "Report back" note at the end of the page), which invited credentials and some selected visiting agents receive; without one there is nothing to report and nothing else changes.
