Self-Adapting Group of Experts for
Multi-Agent Reasoning

1MBZUAI   2Michigan State University   3RIKEN AIP

Abstract

Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent's response depends on its model's capabilities, the reasoning strategy defined by its system prompt, and the information in its input context. Existing frameworks often adapt communication by changing this context while leaving individual prompts fixed, even when a problem calls for different skills. We study whether agents' initial responses can identify a strategy better suited to the current problem and guide its transfer to other agents. To address this, we introduce SAGE (Self-Adapting Group of Experts), a training-free framework that uses answer agreement, prefix consistency, and reciprocal peer review to select a strategy donor. SAGE transfers the selected donor's reasoning strategy to the other agents while preserving their original roles. This transfer uses only the agents' original system prompts, without access to the problem or generated solutions. After strategy adaptation, agents exchange responses through a dynamic, sparse directed acyclic graph that routes information from higher-scoring agents to lower-scoring agents. Experiments across multiple agent backbones and reasoning benchmarks show that SAGE achieves higher average accuracy than the evaluated baselines.

Most multi-agent systems change what agents read. SAGE also changes how they reason.

Pick a donor, share its strategy, then collaborate.

Overview of SAGE in four panels. 1, Select a donor: four agents answer the query, each answer is scored by agreement plus prefix consistency, and peer review selects agent A as donor. 2, Transfer strategy: the donor's system prompt guides rewrites of the other agents' prompts, which keep their own roles; the rewriter sees prompts only. 3, Collaborate: over T rounds, agents revise along a sparse DAG from higher- to lower-scoring agents, and the DAG is rebuilt after each round. 4, Pool and vote: a weighted vote over initial answers, reviews and round answers, with stage leaders weighted 1.5, selects the final answer.

Illustrative run with N = 4 agents and at most K = 2 parents per agent. SAGE allows up to T = 3 rounds; this run stops after two. Training-free, no external judge. Static figure

The rewriter never sees the question.

A borrowed habit still fixes the mistake.

MMLU · Ministral-3-3B answer = D
Select the best translation into predicate logic: Kevin is introduced to José by Wilma. (j: José, k: Kevin, w: Wilma, Ixyz: x introduces y to z) (A) Iwjk(B) Ijkw(C) Ikjw(D) Iwkj
AlgebraicSolver, original prompt
You are an expert in algebraic problem solving.
[specialty paragraph, unchanged]

When solving, prioritize:

1. Defining variables clearly
2. Creating equations from the problem statement
3. Solving equations step by step
4. Simplifying expressions carefully
5. Checking that the solution satisfies the original conditions
Rewritten with donor SkepticalVerificationSolver
You are an expert in algebraic problem solving.
[specialty paragraph, unchanged]

When solving, prioritize:
1. Defining variables clearly and ensuring they align with the problem's constraints.
2. Forming equations that accurately represent the relationships described.
3. Solving equations step-by-step with explicit intermediate calculations.
4. Simplifying expressions methodically to avoid hidden errors.
5. Verifying each step for consistency and checking the final answer against original conditions.

**General habits to adopt:**
- Double-check variable assignments for logical consistency.
- Validate each equation's validity before proceeding.
- Test edge cases or boundary conditions where applicable.
- Compare intermediate results with expected outcomes to detect discrepancies.
- Present solutions with clear, numbered steps and explicit reasoning.
  1. Initial answers A · D · C · C
  2. Rewritten AlgebraicSolver A → D
  3. Pooled vote D 12.0 vs C 8.5

Without rewriting, the same run answers C. Prompts are from the paper's appendix; highlights mark text added or reworded by the rewrite.

SAGE has the best average accuracy on both backbones.

Accuracy (%) of SAGE and baselines on six benchmarks, mean ± sample standard deviation over three runs
MethodMATHGSM8KAQuAGSM-HMMLUGPQAAVG
Qwen2.5-1.5B-Instruct
Single72.07±2.0069.80±0.7261.80±1.7334.27±0.5054.93±0.1228.96±2.7853.64
CoT70.00±0.7271.40±2.2359.20±0.9231.80±0.4052.87±1.1028.11±3.2152.23
MOC70.20±1.2571.20±0.9244.67±0.8334.27±1.2250.20±1.7326.94±2.0449.58
MAD-M272.00±2.4273.80±0.8062.33±0.6134.80±0.5352.20±1.0425.25±2.6753.40
G-Designer71.93±0.1272.47±0.6460.13±1.8034.13±0.9552.67±1.2227.78±5.1353.19
SelfOrg72.87±1.3071.53±1.7562.53±1.3034.00±1.4049.87±2.0427.95±1.1753.12
SAGE78.47±1.0177.33±1.1769.07±0.6439.00±1.0654.33±1.6331.65±3.0458.31
Ministral-3-3B-Instruct-2512
Single89.93±0.1290.53±0.5872.13±1.2245.13±0.3171.00±1.3936.53±4.5867.54
CoT89.47±1.6390.93±1.0374.27±2.3946.73±0.9072.13±0.2336.03±2.5468.26
MOC91.20±0.5390.73±0.6185.40±1.0651.53±0.8171.80±0.7244.44±2.8172.52
MAD-M291.53±0.4292.13±0.3185.47±0.8151.93±0.4273.67±1.3347.98±2.8173.79
G-Designer95.27±0.6491.40±0.6081.67±1.3651.80±0.9271.00±1.7147.64±1.4673.13
SelfOrg95.27±0.1290.80±0.7286.47±0.3148.47±0.6173.47±0.3143.94±0.8773.07
SAGE95.93±0.1293.07±0.5087.93±0.2353.53±1.2175.73±0.5044.28±3.2575.08

BibTeX

@misc{quamar2026selfadaptinggroupexpertsmultiagent,
      title={Self-Adapting Group of Experts for Multi-Agent Reasoning},
      author={Mohammad Atif Quamar and Nurbek Tastan and Karthik Nandakumar and Junpei Komiyama},
      year={2026},
      eprint={2609.35412},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2609.35412},
}