AI & Tech

A New Paper Ships Two Billion Reasoning Traces to Fix Eval Chaos

Chain-of-thought extension, best-of-N voting and tree search all get reported as test-time scaling wins, but almost never at matched compute budgets — so the numbers are not comparable. Hariri and thirteen co-authors propose a fix: classify inference regimes structurally over the model’s implicit prefix tree, evaluate the entire inference system rather than the model alone, and report compute budgets and uncertainty together. They also separate exact-replay from distributional reproducibility and specify what artefacts each needs. The accompanying release is over two billion full reasoning traces with verifier and token-level annotations. Directly usable if you run an inference-time compute budget in production.

Read the original — via arXiv ↗

← All shorts