AI Summary

Penalizing length makes reasoning models pad more. Filtering 'reasoning theater' works instead

ProFIL uses a frozen probe to spot when a reasoning model has already committed to its answer, then drops padded rollouts from RL training. On LiveCodeBench, post-commitment 'theater' fell 72%, while a matched length penalty made it 40% worse.

If you want a reasoning model to stop padding its chain of thought, don’t penalize length. In a new paper, a matched length penalty made LiveCodeBench chains shorter but raised the share of post-commitment ā€œtheaterā€ from 26.7% to 37.4%. ProFIL, a GRPO extension that uses a frozen probe to detect when the model has already committed to its answer and drops padded rollouts from the update, cut theater to 7.6%.

  • Theater down 11 to 100% across four domains: LiveCodeBench āˆ’72%, MMLU-Redux āˆ’50%, ToolUse āˆ’11%, GSM8K ~100% (from a 0.6% baseline).
  • Shorter chains in three of four domains: āˆ’19% on GSM8K, āˆ’18% on ToolUse, āˆ’4% on LiveCodeBench. MMLU-Redux got 5% longer.
  • Accuracy held or improved: GSM8K +5.2 points (77.8% to 83.0%), ToolUse +3 points, MMLU-Redux unchanged. LiveCodeBench depends on how you measure it (see caveats).
  • Outside judges agree: a GPT-4.1 judge that never saw probe scores rated 73% of ProFIL’s LiveCodeBench rollouts faithful, vs. 49% for baseline GRPO.

What ā€œreasoning theaterā€ is

The authors define it with a simple test. Cut the model’s reasoning at each step boundary, append ā€œStop reasoning now… give your final answer,ā€ and decode. The first prefix that yields the correct answer is the commitment point. Everything the model writes after that is theater: it looks like deliberation, but the answer was already there.

This is the same observation behind ConfSFT, which found Nemotron writing a median 1,766 more tokens after reaching 95% confidence. The framing is different. ConfSFT treats the extra tokens as a cost problem. This paper treats them as a faithfulness problem: if the visible chain keeps going after the decision is made, it isn’t an honest record of how the model got there.

The theater also has a recognizable look. On LiveCodeBench, 72% of high-theater baseline rollouts restart a finished solution under a heading like ### Approach, ### Solution, or ### Explanation, vs. 3% of faithful ones. And it isn’t a sign of struggle. High-theater rollouts were more accurate (about 33% vs. 9% for the low-theater third). The authors call it a ā€œconfidence ritualā€: the model re-explains problems it has already solved.

How ProFIL works

  1. Label offline. Run forced answering at every step of base-model rollouts. Steps after the commitment point get label 1, earlier steps 0. The labels come from a verifier (answer match, test execution, or tool-call match), not human annotators.
  2. Train a probe once. A small gated attention probe reads residual-stream activations from two layers of the base model and predicts whether each step is post-commitment. Held-out AUROC is 0.92 to 0.99. Then it’s frozen.
  3. Filter during GRPO. The policy samples 8 rollouts per prompt. A separate frozen copy of the base model reads each rollout’s text, and the probe scores it. If the mean score is above a threshold, that rollout’s reward and advantage are set to zero, even if it got the right answer.

That’s the whole change. The policy never sees the forced-answer prompt, the probe score, or any instruction to be brief. Because the probe reads a frozen model rather than the policy being trained, the policy can’t shift its own activations to fool it. The authors check this directly: the frozen probe still separates theatrical from faithful rollouts with AUROC 1.000 before, during, and after training.

LiveCodeBench: share of rollouts flagged as theater

DeepSeek-R1-Distill-Llama-8B, same data and settings. Lower is better. Bars scaled to 40%.

GRPO + length penalty37.4% Ā· 11,349 chars
GRPO baseline26.7% Ā· 12,382 chars
ProFIL7.6% Ā· 11,875 chars

Shorter isn’t the same as less padded

The length-penalty result is the most useful finding for anyone doing RL on reasoning models. The penalty produced the shortest chains of the three, yet it had the most theater, and ProFIL beat both baselines in every length bucket, including the longest 20% of responses. Shorter chains don’t explain the gain.

The paper’s showcase example makes the point. On one LiveCodeBench problem, the baseline revisits an already-solved edge case for 9,883 characters. The length-penalized model cuts that to 4,886 characters, but it’s still re-explaining and introduces an initialization bug. ProFIL writes the correct five-line program in 2,274 characters. A length penalty squeezes all reasoning, useful or not; the probe only targets what comes after the answer is settled.

A blind pairwise judge, shown two anonymized rollouts and asked which is more padded, picked the baseline over ProFIL in 71% of GSM8K pairs, 72% of ToolUse pairs, 58% of MMLU-Redux pairs, and 57% of LiveCodeBench pairs. The length-penalty model had no such edge over baseline (51%). On ToolUse, the number of tasks solved in a single thinking block rose from 1 of 68 to 8 of 68: the model picks the tool and acts instead of writing a second planning block to reconsider the same plan.

Self-correction survives by design. Because commitment is the first correct prefix, anything before it (wrong turns, fixes) is never flagged. In one GSM8K example, the model first answers 1,200, realizes it has to subtract the dragon’s reach, and corrects to 200. Only the restatement after that is counted as theater.

Caveats

  • LiveCodeBench accuracy depends on how you measure it. The headline 0.30 → 0.37 is the peak sampled score logged during training. At the final checkpoint, greedy accuracy on all 131 test problems fell, from 21.4% to 18.3% (the length-penalty model got 22.9%). The authors say this plainly: the theater reduction is robust, the accuracy claim isn’t.
  • One run per condition. No seed variance, so the confidence intervals don’t capture run-to-run noise in RL.
  • Two small models, not crossed with domains. Llama-8B did GSM8K and LiveCodeBench; Qwen-7B did ToolUse and MMLU-Redux. Architecture and domain effects can’t be separated.
  • Some headline numbers start near zero. GSM8K’s ā€œ~100%ā€ reduction is 0.6% to 0.0%, and MMLU-Redux’s āˆ’50% is 1.6% to 0.8%. The meaningful results are LiveCodeBench and ToolUse.
  • It needs a verifier and real compute. Labels require checkable answers, so open-ended tasks are out for now. The eight RL runs took about 480 H100-hours, and probe scoring adds roughly 60% of a policy update’s per-rollout cost.
  • ā€œFaithfulā€ here means temporal only. The test checks whether the model kept writing after the answer was recoverable, not whether its reasoning is logically valid or actually caused the answer.

The takeaway lines up with ConfSFT from a different direction: models already carry an internal signal for ā€œI’m done,ā€ and training against that signal beats training against token count. If you’re adding a length penalty to your GRPO reward to cut cost, this paper is a reason to check what the shorter outputs actually dropped.

#research #llms #reasoning #reinforcement-learning

Liked this? Engineer's Codex sends one deep dive and a link roundup every week.

No spam. Unsubscribe anytime.