Project Page

The Trace Is the State: Exact Credit Assignment for LLM Agent Teams

Yanjun Chen1,2*, Yirong Sun1, Hanlin Wang2, Jinghan Wang3, Xinming Zhang1, Xiaoyu Shen1, Wenjie Li2, Wei Zhang1†

1Eastern Institute of Technology   2The Hong Kong Polytechnic University   3Harbin Institute of Technology

*Correspondence: yan-jun.chen@connect.polyu.hk   †Corresponding author: zhw@eitech.edu.cn

When the trace is the state, credit need not be predicted; it can be exact.

Status Preprint on arXiv as 2603.06859, under review

Read the official abstract on arXiv or download the paper here.

C3, credit assignment by counterfactual continuation
Agreement with a reference 0.69

Rank correlation of C3 from 4 continuations against a 16-continuation reference; 2 such references agree at 0.73, a trained critic reaches 0.29.

Noise law 2 to 10

Decision points across 6 workflows on which the observed noise follows a derived law with no term for the number of agents.

Training, MATH500 +8.28

Points over MAGRPO, the stronger baseline, with C3 as the advantage for a 2-agent team; 5 training seeds.

Training cost 37%

Fewer training tokens than MAPPO at a 2048-token generation limit: only messages downstream of a decision point are regenerated.

Abstract

Credit assignment for a team of LLM agents, what each message was worth, has mostly been treated as prediction: a learned critic, a trajectory-level score, or an agent ablation stands in for the counterfactual that defines credit, which is rarely run. Teams that communicate through a shared context are different: when everything a downstream agent reads is written into the trace, the trace is the state, and that counterfactual can be executed. A credit signal can then be judged as any estimator is: by bias, variance, and agreement with an independent reference. C3, credit assignment by counterfactual continuation, substitutes one message at a decision point and continues the run to the terminal reward, so its credit is unbiased, exact up to Monte Carlo error. Given the sampled alternatives, that error’s variance follows a derived law with no term for the number of agents, and the observed noise follows the law on 6 workflows of 2 to 10 decision points. On the 2-agent chain, from 4 continuations, C3 ranks alternatives at 0.69 rank correlation against a 16-continuation reference, near the 0.73 at which 2 such references agree; a critic trained on the same rollouts reaches 0.29. Used as the advantage in training a 2-agent team, C3 beats MAGRPO, the stronger baseline in our comparison, and spends 37% fewer training tokens than MAPPO, since only the messages downstream of a decision point are regenerated. When the trace is the state, credit need not be predicted; it can be exact.

Why: credit has been predicted

Training a team of LLM agents from one terminal reward means answering, for every message in the trace, how much it was worth. Reinforcement learning answers by comparing the message with the alternatives the agent could have written, but in a physical environment a decision point cannot be revisited, so a critic predicts what each alternative would have earned. LLM agent teams inherited that tool. Yet when every message, tool call, memory read, and routing decision is written into the trace, the trace is the complete state of the system, and the counterfactual can be executed instead.

Two ways to value an alternative message: predicted by a critic, or executed by C3 when the trace is the state
Figure 1: Two ways to value an alternative message. (a) Credit from predicted counterfactuals: the decision point is not revisited, so a critic fitted on past runs predicts the value of each alternative. (b) Credit from executed counterfactuals, when the trace is the state: C3 resets the run to the trace prefix, substitutes an alternative message, runs the downstream roles to the terminal reward, and compares the mean reward of each alternative with the continuation-weighted mean of the other alternatives. Against a 16-continuation reference, C3 from 4 continuations reaches 0.69 rank correlation, and the critic reaches 0.29.

How: execute the counterfactual

At each decision point, C3 samples alternative messages, resets the run to the trace prefix, substitutes each alternative, and continues the downstream roles to the terminal reward. Each alternative is compared with the continuation-weighted mean return of the others, so its credit is unbiased and uses no learned parameters. What remains is Monte Carlo error, and it is predictable: given the sampled alternatives, its variance follows a derived law set by the decision point's own budget and return variance, with no term for the number of agents. At the opening decision point of 6 workflows of 2 to 10 decision points, the observed variance is 0.90 to 1.07 times the law.

Left: split-half reliability against the implied reliability S/(S+2N) for 6 workflows. Right: observed noise over the derived law, within a band of 10 percent
Figure 2: Reliability follows the signal share of the variance, and the noise follows the law. Left: split-half reliability of 6 workflows on a policy held fixed, against the share S/(S+2N) of a 2-continuation half's variance that is signal; the dashed line is equality. Right: the ratio of the observed noise on an advantage to the derived law, over a grid of branching factors and continuation counts on the 2-agent chain and over the 6 workflows; the band marks ±10%.

Results: near the reference, ahead in training

Scored against a reference of 16 fresh continuations that shares none of its rollouts, C3 from 4 continuations ranks alternatives at 0.69 rank correlation, near the 0.73 at which 2 such references agree; a critic trained on the same rollouts reaches 0.29. Used as the advantage in training a 2-agent team on Qwen3-4B, C3 beats MAGRPO, the stronger baseline, by +8.28 points on MATH500 over 5 training seeds (Holm-corrected p < 0.001). At a 2048-token generation limit it spends 37% fewer training tokens than MAPPO, because only the messages downstream of a decision point are regenerated.

The 5-seed run (Table 2 of the paper):

MethodMATH500AIME 2025CMATHGSM8KAvg.†
MAPPO69.3 ± 0.93.3 ± 0.095.3 ± 0.292.9 ± 0.185.8 ± 0.3
MAGRPO74.5 ± 0.45.3 ± 1.896.1 ± 0.293.4 ± 0.388.0 ± 0.1
C382.8 ± 0.68.0 ± 1.896.3 ± 0.293.4 ± 0.290.9 ± 0.3
Qwen3-4B: greedy accuracy (%) under a 512-token generation limit, mean ± std over 5 training seeds on the 2-agent workflow; best value per column in bold, including ties at the printed precision. †Avg. is the mean of MATH500, CMATH and GSM8K; AIME 2025 (30 problems) is reported beside it, not inside it.

Reproduce

Quickstart (CPU)

python -m pip install -r requirements/cpu.lock.txt
python -m pip install -e . --no-deps
bash scripts/10_data/prepare_all.sh --out_dir data
bash scripts/30_smoke/smoke.sh --task math --limit 1 --print_example 0

Training (GPU)

export PRETRAIN=Qwen/Qwen3-4B-Instruct-2507
RECIPE=main bash scripts/40_train/paper_train.sh  # Table 2
RECIPE=long bash scripts/40_train/paper_train.sh  # Table 7
RECIPE=a3   bash scripts/40_train/paper_train.sh  # Table 8

Core paths

  • c3/credit/ and openrlhf/trainer/ppo_utils/experience_maker.py for the credit computation
  • c3/utils/paper_train_contract.py for the three training recipes
  • scripts/70_rebuild/ for the evaluation and the measurements of Sections 3 and 4, mapped to the paper's tables in its README
  • bash scripts/90_audit/release_gate.sh for the CPU release gate

Citation

If C3 or this repository helps your research, please cite the official arXiv paper below.

@misc{chen2026trace,
  title         = {The Trace Is the State: Exact Credit Assignment for {LLM} Agent Teams},
  author        = {Yanjun Chen and Yirong Sun and Hanlin Wang and Jinghan Wang and Xinming Zhang and Xiaoyu Shen and Wenjie Li and Wei Zhang},
  year          = {2026},
  eprint        = {2603.06859},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  doi           = {10.48550/arXiv.2603.06859},
  url           = {https://arxiv.org/abs/2603.06859}
}