On 8 October, Prime Intellect researcher Radhika Gaonkar posted TRACE, a preprint about a problem hidden inside agent benchmarks: a score can move even when the agent’s underlying work has not changed. The paper’s central point is narrower than “benchmarks are broken.” A score difference alone cannot tell a builder whether the agent changed, the task changed, or the evaluator changed.
That distinction matters when a benchmark is used to compare models, approve a release, or supply reward during training. If a tool is renamed and the agent receives a lower score despite making the same calls, the score records an interface mismatch as a capability loss. If the score is used as a training target, the same mismatch could teach the agent to satisfy the evaluator rather than complete the task.
A score shift can have three causes
The paper defines a verifier as a rule-based or learned evaluator that scores an answer, actions, or final task state. TRACE treats a score change as a question to diagnose, not a verdict about capability. Its sequence is straightforward: run paired tests, compare scores, inspect the trajectories to see whether behavior changed, then rescore the same trajectory after correcting a suspected scoring defect.
Those steps separate three questions that benchmark reports can blur. Did the full agent-task-evaluator system get a different score? Did the agent actually behave differently? Would the evaluator give a different verdict on an unchanged record? Only the last question directly tests evaluator sensitivity. The distinction is practical: a fresh run combines several moving parts, while rescoring the same trajectory holds the agent’s actions fixed.
TRACE uses paired runs and rescoring
Gaonkar first tested TRACE in a synthetic research-operations environment with 25 tasks and scripted agents whose behavior was known. One scripted agent performed the same operations after tool names changed, but its score fell by 0.250 because the verifier counted canonical tool names. Mapping the aliases back before rescoring the same trajectories removed that gap. A second scripted agent changed its behavior under the renamed tools; its remaining score loss was not fixed by correcting the evaluator.
The contrast is the useful result. The same perturbation exposed a scoring defect for one policy and a behavioral failure for another. A score delta by itself could not say which. The scripted agents were deterministic; uncertainty was estimated by resampling tasks, not by counting identical reruns as independent evidence.
The public benchmark study moved to τ²-bench, which tests conversational agents acting through tools in retail and airline settings. Four agents ran with the original interface, renamed tools, and reformatted observations. In an initial 30-task evaluation, seven of eight score-change intervals included zero. The one apparent improvement, for an internal Laguna model under renamed tools, did not hold in the paper’s follow-up.
The apparent gain did not replicate
The larger follow-up used 158 usable retail and airline tasks, including 88 not used in the initial study, with three runs per condition. Seven of eight agent-and-change comparisons were equivalent within the authors’ pre-set margin of 0.10; the eighth was inconclusive. A positive control, deliberately misleading tool names, lowered all four agents’ scores. That control matters: without it, a near-zero result could mean either the harmless changes made no difference or the test was too weak to detect one.
Repeated original runs also changed pass/fail outcomes 14.7% to 35.6% of the time, depending on the agent. That makes a single score comparison a noisy basis for declaring a regression or improvement. The early Laguna gain of 0.267 fell to -0.011 on the same 30 tasks after three runs per condition, and to -0.050 on an intermediate set of 40 new tasks. The authors say the initial result did not replicate.
A separate audit gave two language-model judges the same 118 saved trajectories. The judges disagreed with each other on 57% of records. One agreed with the benchmark’s outcome reward on 95% of verdicts; the other agreed on 38%. This does not establish which judge was right. The paper reports that one judge often objected to procedure even when the benchmark marked the task successful. Agreement with the benchmark is a comparison, not an independent ground truth.
The evidence is not a leaderboard audit
TRACE does not show that current agent rankings are generally wrong. The controlled defect used a designed task suite and scripted policies. The public studies cover four agents and two meaning-preserving interface changes, with a deliberately misleading-name condition used as a positive control, across retail and airline tasks; the author did not test every tool environment, verifier, model, or training setup. Nor does the paper estimate how often production evaluators contain these defects.
The paper is a preprint, submitted to arXiv on 8 October, not a peer-reviewed result. Its own limitations matter: the native benchmark reward is a reference rather than ground truth, the judge study measures agreement and repeatability rather than accuracy, and the fix for the synthetic tool-name problem was tested on one mapping. The study also leaves open whether benchmark builders can overfit to known mutation types when new tests are not held out by mutation family.
There is a second reproducibility wrinkle. The paper links a public code repository, but its README describes a synthetic stress-test suite and a saved 164-trajectory τ²-bench analysis for an August poster. It does not describe the paper’s October four-agent replication artifacts in the README. The paper says its frozen artifacts and analyses are available, but I have not independently rerun those experiments. Builders should inspect the exact snapshot and data accompanying the paper before treating the published table as independently reproduced.
Builders should test the measurement too
For teams using agent scores to choose a model or tune a system, the practical implication is to report more than a mean score. Keep paired trajectories; repeat runs; show absolute pass rates alongside score gaps; include a positive control; and, when a verifier defect is suspected, rescore the same records with only that rule changed. A score that stays stable because an agent fails in both conditions is not evidence of robustness.
For teams using verifier scores as training rewards, the exposure is sharper. An evaluator that rewards a proxy can make the optimization loop improve the proxy while task quality stays flat. TRACE’s synthetic “shortcut” policy demonstrates that possibility in a designed environment; it does not establish how common the failure is in deployed systems. The sensible builder response is not to discard automated evaluation, but to test whether the scoring rule remains aligned with the task under small, meaning-preserving changes.
The paper offers a diagnostic protocol, not a ready-made certification. Its most durable recommendation is methodological: when a score moves, hold the agent record fixed where possible, identify what changed, and show the evidence that links the score to task completion. Until that separation is made, a leaderboard number can tell a team that something changed without telling it what.
