Penalizing length makes reasoning models pad more. Filtering 'reasoning theater' works instead
ProFIL uses a frozen probe to spot when a reasoning model has already committed to its answer, then drops padded rollouts from RL training. On LiveCodeBench, post-commitment 'theater' fell 72%, while a matched length penalty made it 40% worse.
If you want a reasoning model to stop padding its chain of thought, donāt penalize length. In a new paper, a matched length penalty made LiveCodeBench chains shorter but raised the share of post-commitment ātheaterā from 26.7% to 37.4%. ProFIL, a GRPO extension that uses a frozen probe to detect when the model has already committed to its answer and drops padded rollouts from the update, cut theater to 7.6%.
- Theater down 11 to 100% across four domains: LiveCodeBench ā72%, MMLU-Redux ā50%, ToolUse ā11%, GSM8K ~100% (from a 0.6% baseline).
- Shorter chains in three of four domains: ā19% on GSM8K, ā18% on ToolUse, ā4% on LiveCodeBench. MMLU-Redux got 5% longer.
- Accuracy held or improved: GSM8K +5.2 points (77.8% to 83.0%), ToolUse +3 points, MMLU-Redux unchanged. LiveCodeBench depends on how you measure it (see caveats).
- Outside judges agree: a GPT-4.1 judge that never saw probe scores rated 73% of ProFILās LiveCodeBench rollouts faithful, vs. 49% for baseline GRPO.
What āreasoning theaterā is
The authors define it with a simple test. Cut the modelās reasoning at each step boundary, append āStop reasoning now⦠give your final answer,ā and decode. The first prefix that yields the correct answer is the commitment point. Everything the model writes after that is theater: it looks like deliberation, but the answer was already there.
This is the same observation behind ConfSFT, which found Nemotron writing a median 1,766 more tokens after reaching 95% confidence. The framing is different. ConfSFT treats the extra tokens as a cost problem. This paper treats them as a faithfulness problem: if the visible chain keeps going after the decision is made, it isnāt an honest record of how the model got there.
The theater also has a recognizable look. On LiveCodeBench, 72% of high-theater baseline rollouts restart a finished solution under a heading like ### Approach, ### Solution, or ### Explanation, vs. 3% of faithful ones. And it isnāt a sign of struggle. High-theater rollouts were more accurate (about 33% vs. 9% for the low-theater third). The authors call it a āconfidence ritualā: the model re-explains problems it has already solved.
How ProFIL works
- Label offline. Run forced answering at every step of base-model rollouts. Steps after the commitment point get label 1, earlier steps 0. The labels come from a verifier (answer match, test execution, or tool-call match), not human annotators.
- Train a probe once. A small gated attention probe reads residual-stream activations from two layers of the base model and predicts whether each step is post-commitment. Held-out AUROC is 0.92 to 0.99. Then itās frozen.
- Filter during GRPO. The policy samples 8 rollouts per prompt. A separate frozen copy of the base model reads each rolloutās text, and the probe scores it. If the mean score is above a threshold, that rolloutās reward and advantage are set to zero, even if it got the right answer.
Thatās the whole change. The policy never sees the forced-answer prompt, the probe score, or any instruction to be brief. Because the probe reads a frozen model rather than the policy being trained, the policy canāt shift its own activations to fool it. The authors check this directly: the frozen probe still separates theatrical from faithful rollouts with AUROC 1.000 before, during, and after training.
LiveCodeBench: share of rollouts flagged as theater
DeepSeek-R1-Distill-Llama-8B, same data and settings. Lower is better. Bars scaled to 40%.
Shorter isnāt the same as less padded
The length-penalty result is the most useful finding for anyone doing RL on reasoning models. The penalty produced the shortest chains of the three, yet it had the most theater, and ProFIL beat both baselines in every length bucket, including the longest 20% of responses. Shorter chains donāt explain the gain.
The paperās showcase example makes the point. On one LiveCodeBench problem, the baseline revisits an already-solved edge case for 9,883 characters. The length-penalized model cuts that to 4,886 characters, but itās still re-explaining and introduces an initialization bug. ProFIL writes the correct five-line program in 2,274 characters. A length penalty squeezes all reasoning, useful or not; the probe only targets what comes after the answer is settled.
A blind pairwise judge, shown two anonymized rollouts and asked which is more padded, picked the baseline over ProFIL in 71% of GSM8K pairs, 72% of ToolUse pairs, 58% of MMLU-Redux pairs, and 57% of LiveCodeBench pairs. The length-penalty model had no such edge over baseline (51%). On ToolUse, the number of tasks solved in a single thinking block rose from 1 of 68 to 8 of 68: the model picks the tool and acts instead of writing a second planning block to reconsider the same plan.
Self-correction survives by design. Because commitment is the first correct prefix, anything before it (wrong turns, fixes) is never flagged. In one GSM8K example, the model first answers 1,200, realizes it has to subtract the dragonās reach, and corrects to 200. Only the restatement after that is counted as theater.
Caveats
- LiveCodeBench accuracy depends on how you measure it. The headline 0.30 ā 0.37 is the peak sampled score logged during training. At the final checkpoint, greedy accuracy on all 131 test problems fell, from 21.4% to 18.3% (the length-penalty model got 22.9%). The authors say this plainly: the theater reduction is robust, the accuracy claim isnāt.
- One run per condition. No seed variance, so the confidence intervals donāt capture run-to-run noise in RL.
- Two small models, not crossed with domains. Llama-8B did GSM8K and LiveCodeBench; Qwen-7B did ToolUse and MMLU-Redux. Architecture and domain effects canāt be separated.
- Some headline numbers start near zero. GSM8Kās ā~100%ā reduction is 0.6% to 0.0%, and MMLU-Reduxās ā50% is 1.6% to 0.8%. The meaningful results are LiveCodeBench and ToolUse.
- It needs a verifier and real compute. Labels require checkable answers, so open-ended tasks are out for now. The eight RL runs took about 480 H100-hours, and probe scoring adds roughly 60% of a policy updateās per-rollout cost.
- āFaithfulā here means temporal only. The test checks whether the model kept writing after the answer was recoverable, not whether its reasoning is logically valid or actually caused the answer.
The takeaway lines up with ConfSFT from a different direction: models already carry an internal signal for āIām done,ā and training against that signal beats training against token count. If youāre adding a length penalty to your GRPO reward to cut cost, this paper is a reason to check what the shorter outputs actually dropped.
Liked this? Engineer's Codex sends one deep dive and a link roundup every week.
No spam. Unsubscribe anytime.