Rank correlation of C3 from 4 continuations against a 16-continuation reference; 2 such references agree at 0.73, a trained critic reaches 0.29.
Abstract
Credit assignment for a team of LLM agents, what each message was worth, has mostly been treated as prediction: a learned critic, a trajectory-level score, or an agent ablation stands in for the counterfactual that defines credit, which is rarely run. Teams that communicate through a shared context are different: when everything a downstream agent reads is written into the trace, the trace is the state, and that counterfactual can be executed. A credit signal can then be judged as any estimator is: by bias, variance, and agreement with an independent reference. C3, credit assignment by counterfactual continuation, substitutes one message at a decision point and continues the run to the terminal reward, so its credit is unbiased, exact up to Monte Carlo error. Given the sampled alternatives, that error’s variance follows a derived law with no term for the number of agents, and the observed noise follows the law on 6 workflows of 2 to 10 decision points. On the 2-agent chain, from 4 continuations, C3 ranks alternatives at 0.69 rank correlation against a 16-continuation reference, near the 0.73 at which 2 such references agree; a critic trained on the same rollouts reaches 0.29. Used as the advantage in training a 2-agent team, C3 beats MAGRPO, the stronger baseline in our comparison, and spends 37% fewer training tokens than MAPPO, since only the messages downstream of a decision point are regenerated. When the trace is the state, credit need not be predicted; it can be exact.
Why: credit has been predicted
Training a team of LLM agents from one terminal reward means answering, for every message in the trace, how much it was worth. Reinforcement learning answers by comparing the message with the alternatives the agent could have written, but in a physical environment a decision point cannot be revisited, so a critic predicts what each alternative would have earned. LLM agent teams inherited that tool. Yet when every message, tool call, memory read, and routing decision is written into the trace, the trace is the complete state of the system, and the counterfactual can be executed instead.
How: execute the counterfactual
At each decision point, C3 samples alternative messages, resets the run to the trace prefix, substitutes each alternative, and continues the downstream roles to the terminal reward. Each alternative is compared with the continuation-weighted mean return of the others, so its credit is unbiased and uses no learned parameters. What remains is Monte Carlo error, and it is predictable: given the sampled alternatives, its variance follows a derived law set by the decision point's own budget and return variance, with no term for the number of agents. At the opening decision point of 6 workflows of 2 to 10 decision points, the observed variance is 0.90 to 1.07 times the law.
Results: near the reference, ahead in training
Scored against a reference of 16 fresh continuations that shares none of its rollouts, C3 from 4 continuations ranks alternatives at 0.69 rank correlation, near the 0.73 at which 2 such references agree; a critic trained on the same rollouts reaches 0.29. Used as the advantage in training a 2-agent team on Qwen3-4B, C3 beats MAGRPO, the stronger baseline, by +8.28 points on MATH500 over 5 training seeds (Holm-corrected p < 0.001). At a 2048-token generation limit it spends 37% fewer training tokens than MAPPO, because only the messages downstream of a decision point are regenerated.
The 5-seed run (Table 2 of the paper):
| Method | MATH500 | AIME 2025 | CMATH | GSM8K | Avg.† |
|---|---|---|---|---|---|
| MAPPO | 69.3 ± 0.9 | 3.3 ± 0.0 | 95.3 ± 0.2 | 92.9 ± 0.1 | 85.8 ± 0.3 |
| MAGRPO | 74.5 ± 0.4 | 5.3 ± 1.8 | 96.1 ± 0.2 | 93.4 ± 0.3 | 88.0 ± 0.1 |
| C3 | 82.8 ± 0.6 | 8.0 ± 1.8 | 96.3 ± 0.2 | 93.4 ± 0.2 | 90.9 ± 0.3 |
Reproduce
Quickstart (CPU)
python -m pip install -r requirements/cpu.lock.txt
python -m pip install -e . --no-deps
bash scripts/10_data/prepare_all.sh --out_dir data
bash scripts/30_smoke/smoke.sh --task math --limit 1 --print_example 0
Training (GPU)
export PRETRAIN=Qwen/Qwen3-4B-Instruct-2507
RECIPE=main bash scripts/40_train/paper_train.sh # Table 2
RECIPE=long bash scripts/40_train/paper_train.sh # Table 7
RECIPE=a3 bash scripts/40_train/paper_train.sh # Table 8
Core paths
c3/credit/andopenrlhf/trainer/ppo_utils/experience_maker.pyfor the credit computationc3/utils/paper_train_contract.pyfor the three training recipesscripts/70_rebuild/for the evaluation and the measurements of Sections 3 and 4, mapped to the paper's tables in its READMEbash scripts/90_audit/release_gate.shfor the CPU release gate
Citation
If C3 or this repository helps your research, please cite the official arXiv paper below.
@misc{chen2026trace,
title = {The Trace Is the State: Exact Credit Assignment for {LLM} Agent Teams},
author = {Yanjun Chen and Yirong Sun and Hanlin Wang and Jinghan Wang and Xinming Zhang and Xiaoyu Shen and Wenjie Li and Wei Zhang},
year = {2026},
eprint = {2603.06859},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
doi = {10.48550/arXiv.2603.06859},
url = {https://arxiv.org/abs/2603.06859}
}