Teaching reasoning models to rate their confidence makes them think up to 19% shorter
ConfSFT fine-tunes reasoning models only to predict their own confidence mid-thought, with no length penalty or stopping rule. Across four model families, output tokens fell 10 to 19% with accuracy unchanged, even on coding tasks the model never trained on.
Train a reasoning model to answer one question, âhow confident are you in your answer so far?â, and it starts thinking less. Researchers at the University of Maryland and Capital One fine-tuned four open reasoning models only on predicting their own mid-reasoning confidence. The loss never mentions length, stopping, or efficiency, and inference is unchanged. Yet average output tokens fell 10 to 19% with no significant change in accuracy, and the savings carried over from math to science and coding.
- Tokens down 19.2% on Qwen3-4B, 11.1% on Nemotron-Nano-8B and gpt-oss-20b, 10.3% on Gemma-4-E2B. Average accuracy moved by at most 1.1 points in either direction.
- Trained on 600 AIME problems, generalizes out of domain. Qwen3-4Bâs GPQA-Diamond output dropped 24.6% (9,075 to 6,838 tokens) while accuracy rose from 53.0% to 55.8%.
- Matches methods built to shorten reasoning. On Qwen, On-Policy SFT gets 17.2% at 69.3 accuracy; ConfSFT gets 19.2% at 69.5.
- The labels have to be real. Shuffling the confidence targets across examples wipes out the gain (-0.1% on Gemma, +4.0% tokens on Nemotron).
Models already know when theyâre done
The starting observation is that reasoning models keep going after they have the answer. The authors probe a model mid-thought: at each âWaitâ, they cut the trace, append The final answer is \boxed{, and greedily decode a trial answer. Confidence is the geometric mean of that answerâs token probabilities. No gold answer is needed.
Across 480 Nemotron traces on AIME 2024, higher confidence meant the trial answer was more likely to be correct and more likely to match the modelâs eventual final answer. The expected accuracy gain from reasoning further drops sharply once confidence is high. But after first hitting 95% confidence, Nemotron still generated a median of 1,766 more tokens, and many traces ran past 10K. The paperâs example: an AIME problem where the base model spends about 14K tokens and re-derives the same fact six times, while the fine-tuned model reaches the same answer (55) in about 3K.
How ConfSFT works
The training loop is short:
- Sample 8 normal reasoning rollouts per problem from the current model.
- At each âWaitâ (up to a capped number per trace), elicit a trial answer and compute its confidence. Round up to a 2% grid (50 levels).
- Build a training example: problem + reasoning prefix + âFrom 0% (very low) to 100% (very high), my confidence in the answer so far isâ + the label, e.g. â72%â.
- Apply cross-entropy only on the label tokens. The problem, reasoning, and prompt are all masked out.
- Repeat for up to 8 rounds on fresh problems, regenerating rollouts and labels from the updated model each time.
At inference nothing changes: no confidence prompt, no verifier, no early-exit rule. The fine-tuned models also donât start writing confidence percentages into their reasoning. They just reason for fewer tokens.
Average token reduction after ConfSFT
Averaged over AIME2025, GSM8K, GPQA-Diamond, LiveCodeBench, HumanEval. Average accuracy, base â ConfSFT, shown on the right.
The authors report that the token reductions are statistically significant and the accuracy changes are not. Per benchmark, the biggest cuts were Qwen3-4B on GSM8K (â25.4%) and GPQA-Diamond (â24.6%). The smallest was Gemma on GSM8K (â1.9%), where the base model was already brief.
Why not just stop early?
The obvious alternative is to use the same confidence signal at inference time and cut reasoning off once itâs high enough. Thatâs DEER, and it saves more tokens on paper, but it breaks badly outside math. On Nemotron, DEER cuts coding output by about 85% and accuracy with it:
Nemotron-Nano-8B coding accuracy
Early stopping at inference (DEER) vs. confidence training (ConfSFT). Bars scaled to 100%.
LiveCodeBench
HumanEval
Qwen3-4B is worse: DEER drops its average accuracy from 69.4 to 49.8 for a 55% token cut. A hard stopping rule treats every high-confidence moment as the end, even when the model is still halfway through writing code.
Compared with training-time methods that explicitly target length, ConfSFT holds up. A&Z (RL with a length penalty) cut Nemotronâs tokens 5.5% and Gemmaâs 1.8%, against ConfSFTâs 11.1% and 10.3%. On-Policy SFT matched ConfSFTâs token cut on Nemotron (11.0% vs. 11.1%) but lost 1.8 points of average accuracy, and on Gemma it made outputs slightly longer on average.
The shorter traces also look like the originals. Using a Schoenfeld-style taxonomy (Read, Analyze, Plan, Implement, Explore, Verify, Monitor, Answer), methods like L1-Max, ThinkPrune, and On-Policy SFT shift large shares of tokens between Implement, Explore, and Verify. ConfSFT barely changes the mix. It trims across the board rather than cutting one behavior, like verification, that you might want to keep.
Caveats
- Nobody knows why it works. The loss only touches a few label tokens appended after the reasoning, yet generation changes. The paper shows that it happens and that the labels must match the states (shuffled, position-only, and binary-correctness targets all do worse), but offers no mechanism.
- Some per-task accuracy losses. Gemma lost 2.9 points on AIME2025 (35.0% to 32.1%) and gpt-oss-20b lost 2.8 on LiveCodeBench (53.5% to 50.7%). The averages hide these, and the authors donât discuss them.
- Small models, modest gains. The largest model is 20B, and gpt-oss was tuned with LoRA rather than full fine-tuning. A 10 to 19% cut is useful but not dramatic next to methods that trade accuracy for 50%+.
- Confidence isnât calibrated. The authors say plainly that 80% predicted confidence does not mean 80% correct. Itâs a training signal, not a number you can use at runtime.
The practical upshot: if youâre fine-tuning an open reasoning model and paying for its long outputs, this is a cheap, self-supervised step to try. It needs no gold answers, no reward model, and no inference-time changes, and it didnât give up coding accuracy the way early stopping did.
Liked this? Engineer's Codex sends one deep dive and a link roundup every week.
No spam. Unsubscribe anytime.