Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment
Abstract
Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this work, we measure and improve the alignment between the reasoning described in an LLM's CoT and what it computes internally. We propose CoT-Interpretability Alignment (CIA), a metric that measures the agreement between a model's CoT traces and its internal reasoning strategies as detected by interpretability tools. We evaluate CIA on three tasks (two-hop question answering, hint intervention, and integer multiplication) across three LLMs, finding that LLMs exhibit limited alignment across all tasks (44.8-75.9%). We then experiment with improving CIA via post-training, setting both the task accuracy and parametric faithfulness signals as a reward. Experiments show that we can substantially improve CoT parametric faithfulness while maintaining or improving the task accuracy. We provide rich analysis, such as their generalization patterns. Our work provides both a framework for auditing CoT parametric faithfulness and a pathway toward making models' explicit reasoning more trustworthy. Code and data are available at https://github.com/yihuaihong/CIA-minimal-repro.
Community
Do LLMs actually reason the way their chain-of-thought says they do? We introduce CoT-Interpretability Alignment (CIA), a metric that checks whether the reasoning written in a model's CoT matches the internal strategy detected by interpretability tools. Across two-hop QA, hint intervention, and integer multiplication on three LLMs, alignment is limited (44.8–75.9%). We then show that post-training with a reward combining task accuracy and faithfulness substantially improves alignment while maintaining or improving accuracy. Code and data: https://github.com/yihuaihong/CIA-minimal-repro
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment (2026)
- From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness (2026)
- Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability (2026)
- Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models (2026)
- Detecting Hidden Chain-of-Thought in Large Language Models with Linguistic, Behavioral, and Mechanistic Indicators (2026)
- Doesn't Stop Reasoning: Analysis of Spurious CoT Termination (2026)
- Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.38972 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper