Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Abstract
Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/
Community
TL;DR: Asked to answer directly (no CoT), LLMs follow reference chains like K = apple; B = K; D = B; print(D) for only a few lines. A tiny rank-8 LoRA at one early layer lets the frozen middle layers carry the chain much further. The computation is there; it just stops too early.
Key findings
- 13 base models (0.6B–32B) reliably follow only 1.4–3.6 lines. Doubling depth (OLMo-3 7B → 32B) leaves reach at ~2.6 lines.
- Qwen3-8B + a task-trained rank-8 LoRA at layer 14 (65,537 params, all base weights frozen): exact accuracy on 24-line chains goes from 15.5% → 99%. A longer-trained version reaches 50 lines in a single forward pass.
- Mechanism: the LoRA acts on each token independently and moves no information between tokens. It starts a relay: program lines pass their chain identity along through frozen layers 16–22. Cutting each line's attention to its parent in layers 14–22 drops accuracy to chance; the same cut in layers 23–29 barely matters.
- Placement cliff: same recipe, LoRA at layer 20 → 20.5 lines; at layer 21 → 5.2 lines. A frozen-model measurement located this limit within a preregistered tolerance in 3 of 4 held-out models.
- Looped models: in Ouro-1.4B, a LoRA applied every loop reaches 60 lines after 4 loops and ≥160 after 8 (two-chain choice accuracy).
- Multi-hop QA: on MuSiQue (gold paragraphs), separately trained early-layer LoRAs add 9.4–17.9 EM across three standard models.
🎬 Website : https://lunamos.github.io/stop-thinking-too-early/
💻 Code: https://github.com/Lunamos/stop-thinking-too-early
Happy to answer questions!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory (2026)
- Shared Weights, Selected Computations: How Looped Transformers Route What Each Loop Does (2026)
- Loop Dropout: Regularizing Shared Updates in Looped Language Models (2026)
- Improving Test-Time Scaling with Adaptive Looped Transformers (2026)
- Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness (2026)
- Beyond the Training Horizon: Mechanisms and Limits of Length Generalization in Looped Transformers (2026)
- Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.36585 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper