Giulio Zizzo, Ambrish Rawat, et al.
NeurIPS 2022
The field of Language Reasoning Models (LRMs) has advanced rapidly, enabling longer, and more accurate reasoning. However, a growing body of studies show that LRMs are still inefficient, over-generating verification and reflection steps. Additionally, the high-level role of each reasoning step and how different step types contribute to the generation of correct answers, is largely underexplored. To address this challenge, we introduce TRACES (Tagging of the Reasoning steps enabling Adaptive Cost-Efficient early- Stopping), a lightweight framework that tags reasoning steps in real-time, and enable adaptive, cost-efficient early stopping of large-language-model inferences. By monitoring reasoning behaviors during inferences, we find that LRMs tend to shift their reasoning behavior after reaching a correct answer. We demonstrate that the monitoring of the specific type of steps can produce effective interpretable early stopping criteria. Evaluated on three mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and two knowledge and reasoning benchmarks (MMLU and GPQA), TRACES achieve 20-50% token reduction while maintaining comparable accuracy to standard generation. On harder tasks (BeyondAIME, IMO AnswerBench), the accuracy-efficiency trade-off is steeper, requiring a more conservative early-stopping threshold. This work offers a novel way to study and control the generation’s behaviors of LRMs.
Giulio Zizzo, Ambrish Rawat, et al.
NeurIPS 2022
Sihem Amer-Yahia, Jasmina Bogojeska, et al.
EDBT 2025
Basel Shbita, Pengyuan Li, et al.
ESWC 2026
Junheng Hao, Chuan Lei, et al.
KDD 2021