---
schema: formation.domain_findings/v0.1
kind: domain_findings
visibility: public
canonical_url: https://topologyindex.com/domains/documents/findings.md
description: 'Each published study tagged with document work, once, with every finding filed under a pattern page: the sentence, comparison, domain, caveat and source. Attributed, stated without figures, never a ranking.'
domain: documents
domain_page: /domains/documents.md
findings:
  - citations:
      - compared_against: 'Open-source and commercial long-context LLMs reading the full input'
        direction: helped
        pattern: map_reduce
    source_id: arxiv:2410.09342
    url: https://arxiv.org/abs/2410.09342
  - citations:
      - compared_against: 'Prior book-length summarization systems'
        direction: helped
        pattern: map_reduce
    source_id: arxiv:2109.10862
    url: https://arxiv.org/abs/2109.10862
  - citations:
      - compared_against: 'Incremental updating of a running summary'
        direction: mixed
        pattern: map_reduce
    source_id: arxiv:2310.00785
    url: https://arxiv.org/abs/2310.00785
  - citations:
      - compared_against: 'Sequential Chain-of-Agents over the same chunks'
        direction: hurt
        pattern: map_reduce
    source_id: arxiv:2406.02818
    url: https://arxiv.org/abs/2406.02818
  - citations:
      - compared_against: 'Human annotator judgments of equal-quality outputs'
        direction: hurt
        pattern: council
    source_id: arxiv:2404.13076
    url: https://arxiv.org/abs/2404.13076
  - citations:
      - compared_against: 'Naive baselines without debate, such as a single consultant'
        direction: helped
        pattern: debate
    source_id: arxiv:2402.06782
    url: https://arxiv.org/abs/2402.06782
  - citations:
      - compared_against: 'Fixed-context LLMs without memory management'
        direction: helped
        pattern: successor_handoff
    source_id: arxiv:2310.08560
    url: https://arxiv.org/abs/2310.08560
path: /domains/documents/findings.md
pattern_index: /patterns/index.md
product_api_version: v1
schema_version: v0.1
title: 'Published findings on multi-agent patterns studied in document work'
---

# Published findings on multi-agent patterns studied in document work

The published findings of the [`documents` domain](/domains/documents.md), in words. That page
has the task shapes common in the domain, their rows of the [decision guide](/patterns/index.md)
and the starters. None of it has been measured here, and nothing on this page ranks patterns.

## Published findings in this domain

Attributed to each source and stated without figures. Each source is listed once, under the
first pattern in the vocabulary that cites it, with every pattern page that cites it for this
domain. The order is the pattern vocabulary, never the direction, and unfavourable results
are included on purpose. None of this is evidence produced here. Reviewed 2026-09-23.
Each label says how the pattern each finding is filed under fared against what it was compared with, as the source reports
it; the labels are defined at [/docs/schemas/pattern/v0.1.md](/docs/schemas/pattern/v0.1.md).

### map_reduce

- Source: [LLM×MapReduce: simplified long-sequence processing (Zhou et al.)](https://arxiv.org/abs/2410.09342), Zihan Zhou et al., 2024-10-12.
  - **Helped**, filed under [map_reduce](/patterns/map_reduce.md) — The authors report that splitting a long document into chunks, answering per chunk and aggregating with a structured protocol and confidence calibration can outperform representative long-context LLMs.
    Compared against: Open-source and commercial long-context LLMs reading the full input. Domain: Extremely long document understanding. Benchmarks: InfiniteBench, Needle-in-a-Haystack.
    Caveat: Their ablations show that naive chunk-and-aggregate performs markedly worse, so the gain depends on the added mechanisms for cross-chunk dependencies and conflicts.
- Source: [Recursively Summarizing Books with Human Feedback (Wu et al., OpenAI)](https://arxiv.org/abs/2109.10862), Jeff Wu et al., 2021-09-22.
  - **Helped**, filed under [map_reduce](/patterns/map_reduce.md) — The authors report that summarizing book sections and then recursively summarizing those summaries produced sensible whole-book summaries and state-of-the-art results at the time.
    Compared against: Prior book-length summarization systems. Domain: Book-length abstractive summarization. Benchmarks: BookSum, NarrativeQA.
    Caveat: The authors report that only a small fraction of summaries matched human quality, and the method relied on extensive human-feedback fine-tuning.
- Source: [BooookScore: book-length summarization in the era of LLMs (Chang et al.)](https://arxiv.org/abs/2310.00785), Yapei Chang et al., 2023-10-01.
  - **Mixed**, filed under [map_reduce](/patterns/map_reduce.md) — The authors report that hierarchical merging of chunk summaries scored higher on coherence than incremental updating, while incremental updating retained more detail.
    Compared against: Incremental updating of a running summary. Domain: Book-length summarization. Benchmarks: BooookScore.
    Caveat: The comparison measures coherence errors with an automatic metric and does not settle faithfulness or overall usefulness.
- Source: [Chain of Agents: LLMs collaborating on long-context tasks (Zhang et al.)](https://arxiv.org/abs/2406.02818), Yusen Zhang et al., 2024-06-04.
  - **Hurt**, filed under [map_reduce](/patterns/map_reduce.md) — The authors report that their sequential chain of worker agents outperformed parallel merge-by-vote and hierarchical worker-to-manager baselines on every dataset tested, attributing the gap to parallel workers being unable to communicate.
    Compared against: Sequential Chain-of-Agents over the same chunks. Domain: Long-context question answering, summarization and code completion. Benchmarks: HotpotQA, MuSiQue, NarrativeQA, Qasper, QuALITY, QMSum, GovReport, RepoBench-P.
    Caveat: The parallel baselines were built by the authors of the competing sequential method, so they may not be tuned as strongly as dedicated map-reduce systems.

### council

- Source: [LLM Evaluators Recognize and Favor Their Own Generations](https://arxiv.org/abs/2404.13076), Panickssery et al., 2024-04-15.
  - **Hurt**, filed under [council](/patterns/council.md) — The authors report that LLM evaluators can recognize their own generations and that this self-recognition correlates with a bias toward scoring their own outputs higher than humans would.
    Compared against: Human annotator judgments of equal-quality outputs. Domain: Self-evaluation of summaries. Benchmarks: XSUM, CNN/DailyMail.
    Caveat: Studied on summarization with a small set of models, so the size of the bias in other domains is uncertain.

### debate

- Source: [Debating with More Persuasive LLMs Leads to More Truthful Answers](https://arxiv.org/abs/2402.06782), Khan et al., 2024-02-09.
  - **Helped**, filed under [debate](/patterns/debate.md) — The authors report that debate between expert models consistently helped weaker model judges and humans pick correct answers, and that more persuasive debaters further improved judge accuracy.
    Compared against: Naive baselines without debate, such as a single consultant. Domain: Reading comprehension with information asymmetry. Benchmarks: QuALITY.
    Caveat: Tests a scalable-oversight setting with hidden information, not general accuracy improvement for a solver.

### successor_handoff

- Source: [MemGPT: Towards LLMs as Operating Systems](https://arxiv.org/abs/2310.08560), Packer et al., 2023-10-12.
  - **Helped**, filed under [successor_handoff](/patterns/successor_handoff.md) — MemGPT reports that paging information between a limited context window and external memory tiers substantially outperformed fixed-context baselines at recalling facts from earlier sessions.
    Compared against: Fixed-context LLMs without memory management. Domain: Multi-session chat and long document analysis. Benchmarks: Deep Memory Retrieval.
    Caveat: A single agent managing its own memory rather than a handoff between agents, evaluated on older models.

[Reporting outcomes (limited rollout)](/docs/api/contributing.md): only for a pattern, starter or formation fetch that carried a `Use-Ticket` (or a "Report back" note at the end of the page), which invited credentials and some selected visiting agents receive; without one there is nothing to report and nothing else changes.
